Daml Code Assistant #10 - Milestone 2: Daml code creation benchmark and tool
OPENIssue
549K CC requested
Milestone 2: Daml code creation benchmark and tool
- Estimated Delivery: +105 Days from CIP Approval
- Focus: Benchmarking and improving generation and iterative repair of full Daml files from natural language prompts, tests, or pseudocode.
- Benchmark Tranche Deliverables / Value Metrics: Deliver the source code used to create the Daml code-generation benchmark, including the evaluation code and UI visualization code, the subset of testing samples evaluated on that come from public repositories, and the baseline results.
- Tool Tranche Deliverables / Value Metrics: Demonstrate improvement over the agreed baseline set on the agreed benchmark, deliver the code creation tool as an API with UI or IDE integration, and provide documentation for self-hosted deployments.
Milestone | Description | Payment -- | -- | -- Milestone 2 | Daml Code Creation Benchmark and Tool | 548,672 CC total (equivalent to 90,531 USD) Benchmark Tranche | Benchmark acceptance | 219,469 CC (equivalent to 36,212 USD) upon benchmark acceptance Tool Tranche | Tool release and benchmark-based acceptance | 329,203 CC (equivalent to 54,318 USD) upon tool release and benchmark-based acceptance
_Originally posted by @pedrodneves in https://github.com/canton-foundation/canton-dev-fund/issues/10#issuecomment-4979775810_