Snyk Codeを支える先進テクノロジーを探る
Frank Fischer
2021年10月20日
0 分で読めますSnyk Codeは、Snykが提供する静的アプリケーションセキュリティテスト(SAST)ソリューションです。SAST分野に革新的なテクノロジーをもたらします。その基盤となっているのは、ETH(チューリッヒ/スイス)からのスピンオフ企業であるDeepCodeが開発した研究成果とテクノロジーです。同社は2020年末にSnykに加わりました。この記事では、これらのテクノロジーに加え、Snykがオープンソースコミュニティに還元するだけでなく、静的プログラム解析分野の学術コミュニティを支援し、連携している取り組みをご紹介します。
Snyk Codeの処理の流れ
全体の流れを以下に示します。

左端の「Big Code」から始まります。これは、数十万件のオープンソースリポジトリとその変更履歴からなるトレーニングセットで、幅広いプログラミング言語を網羅しています。最初に、コードを解析して中間表現に変換します。Snyk Codeでは、抽象構文木(AST)とイベントグラフ(EG)を使用することで、データフローに配慮したコンテキスト認識型の解析を実現しています。詳しくは、この記事の後半で紹介する学術論文をご覧ください。EGは、さまざまなプログラミング言語を表現できます。
次に、中間表現と論理ルールを使って論理ソルバーを適用します。このソルバーには注目すべき点がいくつかあります。まず、独自開発されたもので、このタスクに最適化されています。標準的なDatalogソルバーで使用されるものより時間計算量の少ないアルゴリズムを採用しています。これは、業界をリードする実行速度にもつながっています。さらに、制約ベースのシステムとして、セマンティックな事実を生成し、人間がガイドする学習プロセスへの入力として利用できます。
Snyk Codeはオープンソースを活用し、世界中の開発者コミュニティの知見を掘り起こして、セキュリティの問題を特定し、対処します。トレーニングセットにはまだ含まれていないものの、潜在的な脅威となり得るソースとシンクの組み合わせも見つけ出します。また、トレーニングデータがまだ存在しない場合でも、ゼロデイエクスプロイトに対応するルールを追加できます。
学習プロセスでは、Snykのセキュリティエンジニアが機械学習アルゴリズムと連携し、ナレッジベースを生成・維持します。また、問題が検出された際に開発者が内容を理解し、対処できるよう、追加情報も提供します。
本番環境でもSnyk Codeはこのパイプラインに従いますが、コードを学習に使うのではなく、ナレッジベースを適用して、数分、場合によっては数秒で正確な結果を生成します。これが、Snyk Codeを強力なSASTソリューションにできた理由です。
Capture the Flagを始めよう
オンデマンドのバーチャル入門ワークショップを見て、Capture the Flagの課題の解き方を学びましょう。
学術論文や学会への参加
Snyk Codeの開発チームは、SASTツールの構築にとどまらず、長年にわたり、機械学習やプログラミング言語分野の主要な学会で論文を発表し、静的プログラム解析に関する学術ツールも複数公開してきました。
論文や研究システムの一部をご紹介します。
TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer — Berabi, B., He, J., Raychev, V. and Vechev, M., 2021年6月. TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer. Proceedings of the 38th International Conference on Machine Learning, PMLR 139(780–791ページ)所収。
Learning to find naming issues with big code and small supervision — He, J., Lee, C.C., Raychev, V. and Vechev, M., 2021年6月. Learning to find naming issues with big code and small supervision. Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation(296–311ページ)所収。
Unsupervised Learning of API Aliasing Specifications — Eberhardt, J., Steffen, S., Raychev, V. and Vechev, M., 2019年6月. Unsupervised learning of API aliasing specifications. Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design and Implementation(745–759ページ)所収。
Scalable Taint Specification Inference with Big Code — Chibotaru, V., Bichsel, B., Raychev, V. and Vechev, M., 2019年6月. Scalable taint specification inference with big code. Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design and Implementation(760–774ページ)所収。
Inferring Crypto API Rules from Code Changes — Paletov, R., Tsankov, P., Raychev, V. and Vechev, M., 2018年. Inferring crypto API rules from code changes. ACM SIGPLAN Notices, 53(4), 450–464ページ。
Learning a Static Analyzer from Data — Bielik, P., Raychev, V. and Vechev, M., 2017年7月. Learning a static analyzer from data. International Conference on Computer Aided Verification(233–253ページ)所収。Springer, Cham。
Probabilistic Model for Code with Decision Trees — Raychev, V., Bielik, P. and Vechev, M., 2016年. Probabilistic model for code with decision trees. ACM SIGPLAN Notices, 51(10), 731–747ページ。
PHOG: Probabilistic Model for Code — Bielik, P., Raychev, V. and Vechev, M., 2016年6月. PHOG: probabilistic model for code. International Conference on Machine Learning(2933–2942ページ)所収。PMLR。
Learning Programs from Noisy Data — Raychev, V., Bielik, P., Vechev, M. and Krause, A., 2016年. Learning programs from noisy data. ACM Sigplan Notices, 51(1), 761–774ページ。
Code Completion with Statistical Language Models — Raychev, V., Vechev, M. and Yahav, E., 2014年6月. Code completion with statistical language models. Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation(419–428ページ)所収。
Predicting Program Properties from "Big Code" — Raychev, V., Vechev, M. and Krause, A., 2015年. Predicting program properties from" big code". ACM SIGPLAN Notices, 50(1), 111–124ページ。
Statistical Deobfuscation of Android Applications — Bichsel, B., Raychev, V., Tsankov, P. and Vechev, M., 2016年10月. Statistical deobfuscation of android applications. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security(343–355ページ)所収。
DEBIN: Predicting Debug Information in Stripped Binaries — He, J., Ivanov, P., Tsankov, P., Raychev, V. and Vechev, M., 2018年10月. Debin: Predicting debug information in stripped binaries. Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security(1667–1680ページ)所収。
公開ツール
さらに、Snyk Codeを支えるテクノロジーの開発過程では、研究ツールも開発・公開されました。以下に、これらのツールをご紹介します。
注:Snykはこれらのツールに責任を負わず、保守やホスティングも行っていません。これらは、Snykのチームメンバーが参加した各研究プロジェクトの一部です。
JSNice
JSNiceは、機械学習を使ってJavaScriptプログラムの難読化を解除します。世界中の数万人のプログラマーに利用されています。

Nice2Predict
Nice2Predictは、構造化予測のための効率的でスケーラブルなオープンソースフレームワークです。新しい統計エンジンをより迅速に構築できます。

DeGuard
DeGuardは、Androidアプリのレイアウト難読化を解除します。セキュリティアナリストに日常的に利用されています。

まとめ
ご紹介したように、Snyk Codeチームは研究の最前線に立ち、学術論文の執筆や学生の受け入れ、その他の研究プロジェクトへの参加を通じて、研究成果を広く共有しています。ぜひ今すぐSnyk Codeに登録して、その研究成果を体験してください。
Capture the Flagを始めよう
オンデマンドのバーチャル入門ワークショップを見て、Capture the Flagの課題の解き方を学びましょう。
