Longevity research is a good test for AI systems because a fluent answer can still be scientifically weak. The domain demands structured evidence, careful ontology handling, and evaluation that checks the reasoning path rather than only the final sentence.
The event sharpened the ideas behind my LongevityLLM Benchmark: provider-agnostic runs, seven task types, and scoring grounded in genes, alleles, phenotype identifiers, lifespan direction, and biological pathways.
