techvit
← Projects
In DevelopmentResearch & implementation

LLM Evaluation Toolkit

LLMアプリケーションの品質評価に関する研究・実装。

LLMEvaluationRAG

LLMを使ったアプリケーションの品質を継続的に評価するためのツールキット。

対応領域

  • Golden Dataset の設計・構築
  • LLM-as-a-Judge による自動評価
  • RAG Evaluation(検索精度・回答精度の評価)
  • Hallucination Detection
  • Regression Testing(変更による品質劣化の検知)
  • Agent Evaluation

LLMアプリケーションが増えるほど「作った後にどう品質を担保するか」が課題になる。ここは今後さらに重要になる領域だと考えている。