Evaluating LLMs for a P&C Actuarial Task Benchmark and Re-Evaluation Suite supports researchers building a scored benchmark and re-evaluation suite to assess language models on property and casualty actuarial tasks.
Funder: Casualty Actuarial Society
Due Dates: October 30, 2026 (proposals) | November 30, 2026 (notification) | December 30, 2026 (executive summary) | July 30, 2027 (final paper)
Funding Amounts: Typically $25,000–$65,000; maximum budget $75,000 USD
Summary: Develop an objectively scored benchmark and re-evaluation suite for large language models performing property and casualty actuarial tasks.
Key Information: Evaluation datasets must remain private to CAS, while the other benchmark materials are intended for its public GitHub repository.
The Casualty Actuarial Society (CAS) Artificial Intelligence Working Group invites researchers and subject matter experts to build a repeatable framework for evaluating large language models on property and casualty (P&C) actuarial work. The focus is on perception and classification tasks with defined answers—such as claims triage, underwriting judgments, fraud flagging, and reserving-related classification—rather than open-ended writing.
The selected team will design tasks, assemble or simulate data, develop objective scoring and model-evaluation protocols, and deliver a public platform comparing model performance. CAS intends to re-test the benchmark as models advance and to maintain the suite independently after delivery.