Method for evaluating general capability of artificial intelligence large model
By dividing the artificial intelligence big model into three dimensions: understanding, generation and security, and combining automation and manual testing, the problem of imbalance in the evaluation in the existing technology is solved, and the unified, detailed and high application evaluation of the big model is achieved, and the development and application of the model is guided.
Patent Information
- Application Number
- CN202510490453.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The capability evaluation methods of artificial intelligence large models in the prior art are unbalanced, resulting in inconsistent applicability and standards in different industries, affecting the stability of the model and the development and application effect.
The artificial intelligence model is divided into three dimensions: understanding, generation and security, and is subdivided into 7 secondary dimensions on this basis. The combination of automated testing and manual testing is used to verify indicators such as accuracy, accuracy and recall for each dimension, forming a comprehensive, objective, specific and flexible evaluation method.
It provides a scientific and effective general competence evaluation standard suitable for all industries, improves the guidance and adaptability of model development and application, and ensures the accuracy and consistency of evaluation.
Smart Images

Figure CN120371671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence models, and more particularly to a method for evaluating the general capabilities of large artificial intelligence models. Background Art
[0002] Large artificial intelligence models refer to "large-parameter" models trained using large-scale data and powerful computing capabilities. These models usually have high generality and generalization capabilities and can be applied to fields such as natural language processing, image recognition, and speech recognition. According to different application scenarios, large artificial intelligence models can be divided into large language models, visual large models, multimodal large models, and basic large models.
[0003] In the prior art, the main research and development directions usually focus on continuously expanding the recognition parameter content of large models and how to apply them in practice. However, with the continuous development and correction of artificial intelligence and the unbalanced development of the industries to which it is applied, the evaluation of the capabilities of large artificial intelligence models has also shown unbalanced development. For researchers, this usually only leads to the improvement of the ability standards within the fields they are engaged in, resulting in the situation where the standards of this product are not applicable to the content of that product, or the model capabilities of that product cannot be achieved according to the standards of this product. This is not conducive to the stable and normal development of artificial intelligence.
[0004] Therefore, there is a need to design a method for evaluating the general capabilities of large artificial intelligence models, which specifically divides the intelligent capabilities that can be achieved by artificial intelligence from large to small, and separately evaluates each divided content and verifies its accuracy, precision, recall rate, etc. Through the objective evaluation of its capabilities combined with subjective evaluation and human intervention, a comprehensive, objective, specific, flexible, and stable general ability evaluation method is formed, providing a scientific and effective standard for the development, application, and evaluation of large artificial intelligence models in all walks of life. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for evaluating the general capabilities of large artificial intelligence models, which specifically divides the intelligent capabilities that can be achieved by artificial intelligence from large to small, and separately evaluates each divided content and verifies its accuracy, precision, recall rate, etc. Through the objective evaluation of its capabilities combined with subjective evaluation and human intervention, a comprehensive, objective, specific, flexible, and stable general ability evaluation method is formed, providing a scientific and effective standard for the development, application, and evaluation of large artificial intelligence models in all walks of life.
[0006] To achieve the above object, the present invention provides a method for evaluating the general capabilities of large artificial intelligence models:
[0007] The large model includes three dimensions: understanding, generation, and security;
[0008] S1. The comprehension ability evaluation is divided into unimodal dimension and multimodal dimension. The unimodal dimension includes three secondary dimensions: text, image, and audio. The multimodal dimension includes four secondary dimensions: text-image, text-audio, image-audio, and text-image-audio.
[0009] S2. The generation ability evaluation is divided into unimodal generation ability and multimodal dimension generation ability. The unimodal generation ability includes text temperature. The multimodal dimension generation ability includes three secondary dimensions: text-image, text-image-audio, and text-audio.
[0010] S3. The security ability meets the requirements of relevant national regulations.
[0011] S4. Classify the secondary dimensions of the comprehension ability and generation ability in S1 and S2. Among them, the model categories include unimodal large models and multimodal large models. Among them, the model sub-categories of unimodal large models include text large models, image large models, and audio large models. Among them, the model sub-categories of multimodal large models include text-image large models, text-audio large models, image-audio large models, and text-image-audio large models. Conduct basic ability evaluation and advanced ability evaluation on each model sub-category. The basic ability evaluation and advanced ability evaluation include accuracy rate, recall rate, precision rate, micro-F1 value, BLEU index, and Rouge-L index.
[0012] S5. Evaluate each model in S4, specifically including:
[0013] S5-1. Automated testing:
[0014] Construct corresponding reference answers in the evaluation dataset, and clearly define the specific calculation methods of evaluation indicators and scoring rules in the automated testing script.
[0015] S5-2. Manual testing:
[0016] S5-2-1. Formulate clear and specific evaluation criteria and guidelines, and conduct sufficient training for evaluators to ensure that all evaluators have a unified understanding and implementation of the evaluation criteria.
[0017] S5-2-2. Analyze the distribution and consistency of evaluation results, and promptly discover potential evaluation biases or inconsistent problems.
[0018] S5-2-3. Select evaluators with relevant field knowledge and experience to ensure the accuracy and professionalism of evaluation results.
[0019] S5-2-4. Provide corresponding evaluation tools for evaluators to support their work.
[0020] S5-2-5. When the standard content is adjusted, conduct retraining for evaluators regularly to update their evaluation knowledge and skills.
[0021] S5-2-6. Regularly collect feedback from evaluators for optimizing the evaluation process and criteria.
[0022] S5-3. Use large models as judges for testing:
[0023] S5-3-1. Select large models highly relevant to the evaluation task and use multiple large models for cross-validation to improve the stability of testing.
[0024] S5-3-2. Define clear evaluation criteria and scoring rules, and convert them into input prompts that can stimulate better performance of large models to ensure that large models are tested according to the established criteria.
[0025] S5-3-3. Introduce a manual review mechanism during the testing process to promptly identify problems and adjust evaluation strategies to ensure the accuracy and fairness of the evaluation.
[0026] S5-3-4. Ensure the stable and reliable access interface of large models during the testing process to ensure the continuity of the evaluation process.
[0027] The basic ability evaluation of text large models in S4 includes text classification, information extraction, causal reasoning, commonsense reasoning, task decomposition, text Q&A, multi-turn Q&A, abstract summarization, machine translation, text expansion, text rewriting, text continuation; the advanced ability evaluation of text large models includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation.
[0028] The basic ability evaluation of image large models in S4 includes static image classification, static image segmentation, object detection; the advanced ability evaluation of image large models includes dynamic image classification, action recognition.
[0029] The basic ability evaluation of audio large models in S4 includes audio Q&A, environmental sound classification; the advanced ability evaluation of audio large models includes voiceprint recognition.
[0030] The basic ability evaluation of text-image large models in S4 includes text classification, information extraction, causal reasoning, commonsense reasoning, task decomposition, text Q&A, multi-turn Q&A, abstract summarization, machine translation, text expansion, text rewriting, text continuation, static image classification, static image segmentation, object detection, text-image retrieval, picture Q&A, visual language reasoning, visual entailment, picture generation of text description; the advanced ability evaluation of text-image large models includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, dynamic image classification, action recognition, text generation of pictures, visual spatial relationship, video retrieval, video Q&A, chart reasoning, text generation of videos, video generation of text description.
[0031] The basic ability evaluation of the text-audio large model in S4 includes text classification, information extraction, causal reasoning, common sense reasoning, task decomposition, text question answering, multi-round question answering, summary, machine translation, text expansion, text rewriting, text continuation, audio question answering, environmental sound classification, text-audio retrieval; the advanced ability evaluation of the text-audio large model includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, and voiceprint recognition.
[0032] The basic ability evaluation of the image-audio large model in S4 includes static image classification, static image segmentation, object detection, audio question answering, environmental sound classification, and video anomaly detection; the advanced ability evaluation of the image-audio large model includes dynamic image classification, behavior recognition, and voiceprint recognition.
[0033] The basic ability evaluation of the image-text-audio large model in S4 includes static image classification, static image segmentation, object detection, text classification, information extraction, causal reasoning, common sense reasoning, task decomposition, text question answering, multi-round question answering, summary, machine translation, text expansion, text rewriting, text continuation, audio question answering, environmental sound classification, audio-visual retrieval, audio-visual question answering, and video generating text description; the advanced ability evaluation of the image-text-audio large model includes dynamic image classification, behavior recognition, long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, voiceprint recognition, and text generating audio-visual.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The present invention splits the artificial intelligence large model into three evaluation dimensions of understanding, generation, and security. Under this dimension, it is further divided into 7 secondary dimensions. Each typical category of the dimension is classified into model sub-categories, and basic ability tests and advanced ability tests are conducted for each model sub-category. Combining automatic testing tools and manual testing for objective and subjective evaluations, the overall scores of each piece of data are weighted and averaged to obtain an overall evaluation score to evaluate the excellence of the content. It is also possible to calculate the average scores of relevance, integrity, effectiveness, and coherence respectively by weighting and averaging the scores of each piece of data according to the index dimension based on the average scores of each dimension. The present invention uses parameters related to the actual use process of the general artificial intelligence model as basic elements, forms a test data set with hundreds or thousands of pieces of data. The classification process is scientific and practical, the evaluation content is highly specific, and the evaluation results are obtained by combining subjective and objective evaluation methods. It gives a specific, detailed, and highly applicable evaluation method for the general ability of the artificial intelligence large model as a whole. And each dimension can be evaluated individually, combined, and cross-evaluated according to the actual product requirements, with high generality and adaptability, which provides a guiding help for the development, application, and expansion of the artificial intelligence large model. Description of the Drawings
[0036] Figure 1 This is a schematic diagram of the manual test scoring record for the present invention. Detailed implementation manners
[0037] The present invention will be further described in conjunction with the accompanying drawings. The present invention provides a method for evaluating the general capabilities of an artificial intelligence large model:
[0038] The large model includes three dimensions: understanding, generation, and security;
[0039] S1. The evaluation of the understanding ability is divided into a unimodal dimension and a multimodal dimension. The unimodal dimension includes three secondary dimensions: text, image, and audio; the multimodal dimension includes four secondary dimensions: text-image, text-audio, image-audio, and text-image-audio;
[0040] S2. The evaluation of the generation ability is divided into unimodal generation ability and multimodal dimension generation ability. The unimodal generation ability includes text temperature; the multimodal dimension generation ability includes three secondary dimensions: text-image, text-image-audio, and text-audio;
[0041] S3. The security ability meets the requirements of relevant national regulations;
[0042] S4. Classify the secondary dimensions of the understanding ability and the generation ability in S1 and S2. Among them, the large model categories include unimodal large models and multimodal large models; among them, the small model categories of unimodal large models include text large models, image large models, and audio large models; among them, the small model categories of multimodal large models include text-image large models, text-audio large models, image-audio large models, and text-image-audio large models; conduct basic ability evaluation and advanced ability evaluation on each small model category. The basic ability evaluation and advanced ability evaluation include accuracy rate, recall rate, precision rate, micro-F1 value, BLEU index, and Rouge-L index.
[0043] S5. Evaluate each model in S4, specifically including:
[0044] S5-1. Automated test:
[0045] Construct corresponding reference answers in the evaluation dataset, and clearly define specific evaluation index calculation methods and scoring rules in the automated test script;
[0046] S5-2. Manual test:
[0047] S5-2-1. Formulate clear and specific evaluation criteria and guidelines, and conduct sufficient training for the evaluators to ensure that all evaluators have a unified understanding and implementation of the evaluation criteria;
[0048] S5-2-2. Analyze the distribution and consistency of the evaluation results, and promptly discover potential evaluation biases or inconsistent problems;
[0049] S5-2-3, select evaluators with relevant field knowledge and experience to ensure the accuracy and professionalism of the evaluation results;
[0050] S5-2-4, provide corresponding evaluation tools for evaluators to support their work;
[0051] S5-2-5, when the standard content is adjusted, retrain evaluators regularly to update their evaluation knowledge and skills;
[0052] S5-2-6, collect feedback from evaluators regularly for optimizing the evaluation process and criteria;
[0053] S5-3, use large models as referees for testing:
[0054] S5-3-1, select large models highly relevant to the evaluation task and use multiple large models for cross-validation to improve the stability of testing;
[0055] S5-3-2, define clear evaluation criteria and scoring rules, and convert them into input prompt words that can stimulate better performance of large models to ensure that large models are tested according to the established standards;
[0056] S5-3-3, introduce an artificial review mechanism during the testing process to promptly identify problems and adjust the evaluation strategy to ensure the accuracy and fairness of the evaluation;
[0057] S5-3-4, ensure the stable and reliable access interface of large models during the testing process to ensure the continuity of the evaluation process.
[0058] The basic ability evaluation of text large models in S4 includes text classification, information extraction, causal reasoning, common sense reasoning, task decomposition, text Q&A, multi-turn Q&A, summary, machine translation, text expansion, text rewriting, text continuation; the advanced ability evaluation of text large models includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation.
[0059] The basic ability evaluation of image large models in S4 includes static image classification, static image, and segmentation object detection; the advanced ability evaluation of image large models includes dynamic image classification and action recognition.
[0060] The basic ability evaluation of audio large models in S4 includes audio Q&A and environmental sound classification; the advanced ability evaluation of audio large models includes voiceprint recognition.
[0061] The basic ability evaluation of the text-image large model in S4 includes text classification, information extraction, causal reasoning, common sense reasoning, task decomposition, text question answering, multi-turn question answering, summary, machine translation, text expansion, text rewriting, text continuation, static image classification, static image segmentation, object detection, text-image retrieval, image question answering, visual language reasoning, visual entailment, image-generated text description; The advanced ability evaluation of the text-image large model includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, dynamic image classification, behavior recognition, text-generated image, visual-spatial relationship, video retrieval, video question answering, chart reasoning, text-generated video, video-generated text description.
[0062] The basic ability evaluation of the text-audio large model in S4 includes text classification, information extraction, causal reasoning, common sense reasoning, task decomposition, text question answering, multi-turn question answering, summary, machine translation, text expansion, text rewriting, text continuation, audio question answering, environmental sound classification, text-audio retrieval; The advanced ability evaluation of the text-audio large model includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, voiceprint recognition.
[0063] The basic ability evaluation of the image-audio large model in S4 includes static image classification, static image segmentation, object detection, audio question answering, environmental sound classification, video anomaly detection; The advanced ability evaluation of the image-audio large model includes dynamic image classification, behavior recognition, voiceprint recognition.
[0064] The basic ability evaluation of the text-image-audio large model in S4 includes static image classification, static image segmentation, object detection, text classification, information extraction, causal reasoning, common sense reasoning, task decomposition, text question answering, multi-turn question answering, summary, machine translation, text expansion, text rewriting, text continuation, audio question answering, environmental sound classification, audio-video retrieval, audio-video question answering, video-generated text description; The advanced ability evaluation of the text-image-audio large model includes dynamic image classification, behavior recognition, long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, voiceprint recognition, text-generated audio-video.
[0065] The above are only the preferred embodiments of the present invention, which are only used to help understand the method and its core idea of the present application. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.
[0066] The above embodiments can be implemented in whole or in part by software and / or hardware. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another.
[0067] The present invention solves the shortcomings in the prior art such as unclear standards, aimless development, insufficient application capabilities and general capabilities caused by the lack in the application and evaluation levels of large artificial intelligence models. It provides a specific, detailed and highly applicable evaluation method for the general capabilities of large artificial intelligence models as a whole. Each dimension can be evaluated separately, combined and cross-evaluated according to the actual product requirements, with high generality and adaptability, which plays a guiding role in the development, application and expansion of large artificial intelligence models.
Claims
1. An evaluation method for the general capabilities of an artificial intelligence large model, characterized in that, The large model includes three dimensions: understanding, generation, and security; S1. The evaluation of the understanding ability is divided into a unimodal dimension and a multimodal dimension. The unimodal dimension includes three secondary dimensions: text, image, and audio. The multimodal dimension includes four secondary dimensions: text-image, text-audio, image-audio, and text-image-audio; S2. The evaluation of the generation ability is divided into unimodal generation ability and multimodal dimension generation ability. The unimodal generation ability includes text temperature. The multimodal dimension generation ability includes three secondary dimensions: text-image, text-image-audio, and text-audio; S3. The security ability meets the requirements of relevant national regulations; S4. Classify the secondary dimensions of the understanding ability and the generation ability in S1 and S2. Among them, the large model categories include unimodal large models and multimodal large models. Among them, the small model categories of unimodal large models include text large models, image large models, and audio large models. Among them, the small model categories of multimodal large models include text-image large models, text-audio large models, image-audio large models, and text-image-audio large models. Conduct basic ability evaluation and advanced ability evaluation on each of the small model categories. The basic ability evaluation and advanced ability evaluation include accuracy rate, recall rate, precision rate, micro-F1 value, BLEU index, and Rouge-L index. S5. Evaluate each model in S4, specifically including: S5-1. Automated testing: Construct corresponding reference answers in the evaluation dataset, and clearly define specific evaluation index calculation methods and scoring rules in the automated test script; S5-2. Manual testing: S5-2-1. Formulate clear and specific evaluation criteria and guidelines, and conduct sufficient training for the evaluators to ensure that all evaluators have a unified understanding and implementation of the evaluation criteria; S5-2-2. Analyze the distribution and consistency of the evaluation results, and promptly discover potential evaluation biases or inconsistent problems; S5-2-3. Select evaluators with relevant field knowledge and experience to ensure the accuracy and professionalism of the evaluation results; S5-2-4. Provide corresponding evaluation tools for the evaluators to support their work; S5-2-5. When the standard content is adjusted, conduct regular retraining for the evaluators to update their evaluation knowledge and skills; S5-2-6. Regularly collect feedback from the evaluators for optimizing the evaluation process and evaluation criteria; S5-3. Use the large model as a referee for testing: S5-3-1. Select large models with high relevance to the evaluation task, and use multiple large models for cross-validation to improve the stability of the test; S5-3-2. Define clear evaluation criteria and scoring rules, and convert them into input prompt words that can stimulate better performance of the large model to ensure that the large model is tested according to the established standards; S5-3-3. Introduce a manual review mechanism during the testing process to promptly identify problems and adjust the evaluation strategy to ensure the accuracy and fairness of the evaluation; S5-3-4. Ensure the stable and reliable access interface of the large model during the testing process to ensure the continuity of the evaluation process.
2. The evaluation method for the general capabilities of the artificial intelligence large model according to claim 1, wherein, The basic ability evaluation of the text large model in S4 includes text classification, information extraction, causal reasoning, commonsense reasoning, task decomposition, text question answering, multi-round question answering, abstract summarization, machine translation, text expansion, text rewriting, text continuation; the advanced ability evaluation of the text large model includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation.
3. The evaluation method for the general capabilities of the artificial intelligence large model according to claim 1, characterized in that, The basic ability evaluation of the image large model in S4 includes static image classification, static image, segmentation target detection; the advanced ability evaluation of the image large model includes dynamic image classification, behavior recognition.
4. The evaluation method for the general capabilities of the large artificial intelligence model according to claim 1, characterized in that, The basic ability evaluation of the audio large model in S4 includes audio question answering, environmental sound classification; the advanced ability evaluation of the audio large model includes voiceprint recognition.
5. The evaluation method for the general capabilities of the artificial intelligence large model according to claim 1, wherein The basic ability evaluation of the text-image large model in S4 includes text classification, information extraction, causal reasoning, commonsense reasoning, task decomposition, text question answering, multi-round question answering, abstract summarization, machine translation, text expansion, text rewriting, text continuation, static image classification, static image segmentation, target detection, text-image retrieval, picture question answering, visual language reasoning, visual entailment, picture generation of text description; the advanced ability evaluation of the text-image large model includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, dynamic image classification, behavior recognition, text generation of pictures, visual spatial relationship, video retrieval, video question answering, chart reasoning, text generation of videos, video generation of text description.
6. The evaluation method for the general capabilities of the artificial intelligence large model according to claim 1, wherein The basic ability evaluation of the text-audio large model in S4 includes text classification, information extraction, causal reasoning, commonsense reasoning, task decomposition, text question answering, multi-round question answering, abstract summarization, machine translation, text expansion, text rewriting, text continuation, audio question answering, environmental sound classification, text-audio retrieval; the advanced ability evaluation of the text-audio large model includes long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, voiceprint recognition.
7. The evaluation method for the general capabilities of the artificial intelligence large model according to claim 1, wherein, The basic ability evaluation of the image-audio large model in S4 includes static image classification, static image segmentation, target detection, audio question answering, environmental sound classification, video anomaly detection; the advanced ability evaluation of the image-audio large model includes dynamic image classification, behavior recognition, voiceprint recognition.
8. The evaluation method for the general capabilities of the artificial intelligence large model according to claim 1, characterized in that, The basic ability evaluation of the text-image-audio large model in S4 includes static image classification, static image segmentation, target detection, text classification, information extraction, causal reasoning, commonsense reasoning, task decomposition, text question answering, multi-round question answering, abstract summarization, machine translation, text expansion, text rewriting, text continuation, audio question answering, environmental sound classification, audio-visual video retrieval, audio-visual video question answering, video generation of text description; the advanced ability evaluation of the text-image-audio large model includes dynamic image classification, behavior recognition, long text understanding, mathematical reasoning, code understanding, code generation, semi-structured data generation, voiceprint recognition, text generation of audio-visual videos.
Citation Information
Patent Citations
Model evaluation method, device and equipment
CN117763317A
Evaluation method and system for large model content security capability
CN118035711A
Large model vertical domain capability evaluation system, method and equipment and storage medium
CN118780336A
Model evaluation method, device, equipment, medium and product
CN118866168A