In one embodiment, a non-transitory computer-readable media stores instructions
executable by processors for generating a prompt configured for eliciting outputs from large language models (LLMs) based on information associated with a task, inputting the prompt to a first LLM configured to output a response based on
processing the prompt, determining
metrics for evaluating the first LLM based on the task, wherein each of the
metrics is associated with a scoring
guideline, generating metric prompts based on the respective
metrics and the scoring guidelines associated with the respective metrics, inputting the response and the metric prompts to second LLMs configured to output scores corresponding to the respective metrics based on
processing the response and the metric prompts, and generating an analysis report based on the metrics and their corresponding scores.