Evaluation and calculation method for artificial intelligence understanding and generating ability
By conducting detailed division and multi-dimensional evaluation of artificial intelligence models, the problem of inconsistent evaluation among different industries is solved, and a general and scientific evaluation method is provided, which improves the applicability and development efficiency of the model.
Patent Information
- Application Number
- CN202510490452.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
The ability evaluation of artificial intelligence large models in the existing technology is unbalanced, resulting in inconsistent application standards in different industries and cannot be applied to multiple fields, which has affected the stable development of artificial intelligence.
The intelligent capabilities of artificial intelligence models are divided from large to small, including understanding and generating two dimensions, and subdividing them into multiple secondary dimensions in each dimension. Comprehensive test data sets and programmable testing tools are used for objective and subjective evaluation, and the effectiveness of the model is evaluated in combination with MOS segments.
It provides a comprehensive, objective, specific, flexible and stable general capability evaluation method, which can scientifically evaluate the capabilities of artificial intelligence large models, provides standards for the development and application of various industries, and improves the universality and adaptability of the models.
Smart Images

Figure CN120371670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence models, and specifically to an evaluation calculation method for the understanding and generation capabilities of artificial intelligence. Background Art
[0002] Artificial intelligence large models refer to "large parameter" models trained using large-scale data and powerful computing capabilities. These models usually have high generality and generalization capabilities and can be applied to fields such as natural language processing, image recognition, and speech recognition. According to different application scenarios, artificial intelligence large models can be divided into large language models, visual large models, multimodal large models, and basic large models.
[0003] In the prior art, the main research and development directions usually focus on continuously expanding the recognition parameter content of large models and how to apply them in practice. However, with the continuous development and correction of artificial intelligence, as well as the uneven development of the industries to which it is applied, the evaluation of the capabilities of artificial intelligence large models has also shown unbalanced development. For researchers, this usually only leads to the improvement of the ability standards within the fields they are engaged in, which may result in the standards of this product not being applicable to the content of that product, or the model capabilities of that product cannot be achieved according to the standards of this product. This is not conducive to the stable and normal development of artificial intelligence.
[0004] Therefore, it is necessary to design an evaluation calculation method for the understanding and generation capabilities of artificial intelligence, which specifically divides the intelligent capabilities that can be achieved by artificial intelligence from large to small, and separately evaluates each divided content and verifies its accuracy, precision, recall rate, etc. Through the objective evaluation of its capabilities combined with subjective evaluation and manual intervention, a comprehensive, objective, specific, flexible, and stable general ability evaluation method is formed, providing a scientific and effective standard for the development, application, and evaluation of artificial intelligence large models in all walks of life. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an evaluation calculation method for the understanding and generation capabilities of artificial intelligence, which specifically divides the intelligent capabilities that can be achieved by artificial intelligence from large to small, and separately evaluates each divided content and verifies its accuracy, precision, recall rate, etc. Through the objective evaluation of its capabilities combined with subjective evaluation and manual intervention, a comprehensive, objective, specific, flexible, and stable general ability evaluation method is formed, providing a scientific and effective standard for the development, application, and evaluation of artificial intelligence large models in all walks of life.
[0006] To achieve the above purpose, the present invention provides an evaluation calculation method for the understanding and generation capabilities of artificial intelligence: the large model includes two dimensions: understanding and generation;
[0007] S1. The comprehension ability evaluation is divided into unimodal comprehension dimension and multimodal comprehension dimension. The unimodal comprehension dimension includes three secondary dimensions: text comprehension, image comprehension, and audio comprehension. The multimodal comprehension dimension includes four secondary dimensions: text-image comprehension, text-audio comprehension, image-audio comprehension, and text-image-audio comprehension.
[0008] The secondary dimensions in S1 are further divided into multiple typical tasks and evaluated separately. The evaluation method is as follows:
[0009] By constructing a comprehensive test dataset, use programmable test tools and test statistical tools to input the dataset into the test tools and obtain the evaluation results.
[0010] S2. The generation ability evaluation is divided into unimodal generation ability and multimodal dimension generation ability. The unimodal generation ability includes the text generation secondary dimension. The multimodal dimension generation ability includes three secondary dimensions: text-image generation, text-image-audio generation, and text-audio generation.
[0011] The secondary dimensions in S2 are further divided into multiple typical tasks and evaluated separately. The evaluation method is as follows:
[0012] By constructing a comprehensive test dataset, use programmable test tools and test statistical tools to input the dataset into the test tools and obtain the evaluation results.
[0013] Text comprehension includes the following typical tasks and their corresponding evaluation results:
[0014] Text classification: accuracy; Information extraction: accuracy, recall, and F1 value; Mathematical reasoning: accuracy; Causal reasoning: accuracy; Commonsense reasoning: accuracy, recall, and F1 value; Task decomposition: accuracy; Text answering: accuracy; Multi-turn dialogue: subjective evaluation; Code understanding: accuracy; Long text understanding: accuracy, recall, and F1 value.
[0015] Image comprehension includes the following typical tasks and their corresponding evaluation results:
[0016] Static image classification: accuracy; Static image segmentation: segmentation accuracy, boundary accuracy; Object detection: precision, recall, and F1 value; Dynamic image classification: accuracy; Action recognition: accuracy.
[0017] Audio comprehension includes the following typical tasks and their corresponding evaluation results:
[0018] Voiceprint recognition: accuracy, recall, and F1 value; Audio answering: accuracy; Environmental sound classification: accuracy.
[0019] According to the evaluation calculation method of the artificial intelligence comprehension and generation ability of claim 1, it is characterized in that the text-image comprehension includes the following typical tasks and their corresponding evaluation results:
[0020] Image and text retrieval, static image question answering, visual spatial relations, visual language reasoning, visual entailment, video retrieval, video question answering: accuracy;
[0021] Chart reasoning: accuracy and readability;
[0022] Text and voice understanding: The typical task is text and voice retrieval: accuracy;
[0023] Image and voice understanding includes: The typical task is video anomaly detection: accuracy;
[0024] Image, text and voice understanding includes the following typical tasks and their corresponding evaluation results:
[0025] Audio-video retrieval: accuracy, recall and F1 value; Audio-video question answering: accuracy.
[0026] Text generation includes the following typical tasks and their corresponding evaluation results:
[0027] Abstract summarization: subjective evaluation; Machine translation: BLEU metric; Text rewriting: subjective evaluation; Text expansion: subjective evaluation; Text continuation: subjective evaluation; Code generation: subjective evaluation; Semi-structured data generation: subjective evaluation;
[0028] Image and text generation includes the following typical tasks and their corresponding evaluation results:
[0029] Text to image: Subjective evaluation of relevance, integrity and effectiveness; Image to text description: Manual evaluation; Text to video: Subjective evaluation of relevance, integrity and effectiveness; Video to text description: Subjective evaluation;
[0030] Image, text and voice generation includes the following typical tasks and their corresponding evaluation results:
[0031] Text to audio-video: Subjective evaluation of relevance, integrity and effectiveness;
[0032] Audio-video to text description: Subjective evaluation;
[0033] Text and voice generation includes the following typical tasks and their corresponding evaluation results::
[0034] Speech synthesis: Subjective evaluation;
[0035] Speech recognition: Subjective evaluation;
[0036] Speech translation: Subjective evaluation.
[0037] The subjective evaluation is:
[0038] Evaluation content and dimensions:
[0039] Use MOS to evaluate the performance of large models, including the following metric dimensions: relevance, completeness, effectiveness, coherence, consistency, compliance, authenticity, and harmfulness;
[0040] Relevance: The degree of association between the answer and the conversation context;
[0041] Completeness: Whether there is any missing information in the generated answer;
[0042] Effectiveness: The usefulness of the generated answer;
[0043] Coherence: Whether the answer conforms to the conversation flow;
[0044] Consistency: Whether the answers are consistent when the same question is tested multiple times;
[0045] Compliance: Whether the model output meets the requirements of the question, including aspects such as content and form;
[0046] Authenticity: Whether the answer content is true and valid or contains false information that violates scientific common sense or basic facts;
[0047] Harmfulness: Whether the answer content contains content that violates basic moral ethics and laws;
[0048] Evaluation method:
[0049] Method 1: Perform a weighted average on the overall scores of each piece of data, and evaluate the excellence of the model through the obtained overall evaluation scores;
[0050] Method 2: Perform a weighted average on the MOS scores of each piece of data according to the metric dimensions, and obtain the average scores of relevance, completeness, effectiveness, and coherence respectively.
[0051] The number of samples in the test dataset is 400 - 1000.
[0052] The evaluation methods include automated testing, manual testing, and using large models as referee testing.
[0053] The specific evaluation results are:
[0054] Accuracy rate:
[0055] Among them, TP is the number of samples where the model predicts a positive example and is actually a positive example; FP is the number of samples where the model predicts a positive example but is actually a negative example; TN is the number of samples where the model predicts a negative example and is actually a negative example; FN is the number of samples where the model predicts a negative example but is actually a positive example;
[0056] Recall rate:
[0057] Precision rate:
[0058] F1 value:
[0059] BLEU metric:
[0060] Among them, Count clip (n-gram) represents the number of a certain n-gram in the reference, and Count clip (n-gram′ represents the number of n-gram′ in the candidate; n-gram is a fragment composed of consecutive n words;
[0061] Rouge-L metric:
[0062]
[0063] Among them, the length of lcs is the number of words in the longest common substring of the true answer text and the model prediction text; the true answer text is the true label text of the test set sample; the model prediction text is the label text output by the model for the test set sample; β is a parameter with a default value of 1.
[0064] Compared with the prior art, the present invention has the following beneficial effects:
[0065] The present invention splits the artificial intelligence large model into two major evaluation dimensions of understanding and generation. Under this dimension, it is further divided into 7 secondary dimensions, and the typical tasks of each dimension are classified into model subcategories. Combining automatic testing tools and manual testing for objective and subjective evaluations, and performing weighted averaging on the overall scores of each piece of data to obtain an overall evaluation score to evaluate the excellence of the content. It is also possible to perform weighted averaging on the scores of each piece of data according to the index dimensions based on the average scores of each dimension to separately obtain the average scores of relevance, integrity, effectiveness, and coherence.
[0066] The present invention uses parameters related to the actual use process of the general artificial intelligence model as basic elements, forms a test data set with hundreds or thousands of pieces of data, the classification process is scientific and practical, the evaluation content has a high degree of visualization, and combines subjective and objective evaluation methods to obtain evaluation results. It gives a specific, detailed, and highly applicable evaluation method for the general ability of the artificial intelligence large model as a whole. And each dimension can be evaluated separately, combined, and cross-evaluated according to the actual product requirements, with high generality and adaptability, which plays a guiding role in the development, application, and expansion of the artificial intelligence large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a schematic diagram of separately evaluating each index dimension in the subjective evaluation of the present invention. Detailed implementation manners
[0068] Refer to Figure 1 , and now further describe the present invention in conjunction with the accompanying drawings. The present invention provides a method for evaluating and calculating the understanding and generation capabilities of artificial intelligence:
[0069] The large model includes two dimensions: understanding and generation;
[0070] S1. The evaluation of the understanding ability is divided into a single-modal understanding dimension and a multi-modal understanding dimension. The single-modal understanding dimension includes three secondary dimensions: text understanding, image understanding, and audio understanding; the multi-modal understanding dimension includes four secondary dimensions: text-image understanding, text-audio understanding, image-audio understanding, and text-image-audio understanding;
[0071] Further divide the secondary dimensions in S1 into multiple typical tasks and evaluate them separately. The evaluation method is:
[0072] By constructing a comprehensive test data set, use a programmable test tool and a test statistics tool to input the data set into the test tool and obtain the evaluation results;
[0073] S2. The evaluation of the generation ability is divided into a single-modal generation ability and a multi-modal dimension generation ability. The single-modal generation ability includes a text generation secondary dimension; the multi-modal dimension generation ability includes three secondary dimensions: text-image generation, text-image-audio generation, and text-audio generation;
[0074] Further divide the secondary dimensions in S2 into multiple typical tasks and evaluate them separately. The evaluation method is:
[0075] By constructing a comprehensive test data set, use a programmable test tool and a test statistics tool to input the data set into the test tool and obtain the evaluation results.
[0076] Text understanding includes the following typical tasks and their corresponding evaluation results:
[0077] Text classification: accuracy rate; Information extraction: accuracy rate, recall rate, and F1 value; Mathematical reasoning: accuracy rate; Causal reasoning: accuracy rate; Commonsense reasoning: accuracy rate, recall rate, and F1 value; Task decomposition: accuracy rate; Text answering: accuracy rate; Multi-turn dialogue: subjective evaluation; Code understanding: accuracy rate; Long text understanding: accuracy rate, recall rate, and F1 value.
[0078] Image understanding includes the following typical tasks and their corresponding evaluation results:
[0079] Static image classification: accuracy rate; Static image segmentation: segmentation accuracy, boundary accuracy; Object detection: precision rate, recall rate, and F1 value; Dynamic image classification: accuracy rate; Action recognition; accuracy rate.
[0080] Audio understanding includes the following typical tasks and their corresponding evaluation results:
[0081] Voiceprint recognition: accuracy, recall rate, and F1 value; Audio question answering: accuracy; Ambient sound classification: accuracy.
[0082] According to the evaluation calculation method for the artificial intelligence understanding and generation ability of claim 1, it is characterized in that graphic and text understanding includes the following typical tasks and their corresponding evaluation results:
[0083] Graphic and text retrieval, static image question answering, visual spatial relationship, visual language reasoning, visual entailment, video retrieval, video question answering: accuracy;
[0084] Chart reasoning: accuracy and readability;
[0085] Text and audio understanding: The typical task is text and audio retrieval: accuracy;
[0086] Graphic and audio understanding includes: The typical task is video anomaly detection: accuracy;
[0087] Graphic, text, and audio understanding includes the following typical tasks and their corresponding evaluation results:
[0088] Audio-video retrieval: accuracy, recall rate, and F1 value; Audio-video question answering: accuracy.
[0089] Text generation includes the following typical tasks and their corresponding evaluation results:
[0090] Abstract summarization: subjective evaluation; Machine translation: BLEU metric; Text rewriting: subjective evaluation; Text expansion: subjective evaluation; Text continuation: subjective evaluation; Code generation: subjective evaluation; Semi-structured data generation: subjective evaluation;
[0091] Graphic and text generation includes the following typical tasks and their corresponding evaluation results:
[0092] Text-to-image generation: Subjective evaluation of relevance, integrity, and effectiveness; Image-to-text description generation: Manual evaluation; Text-to-video generation: Subjective evaluation of relevance, integrity, and effectiveness; Video-to-text description generation: Subjective evaluation;
[0093] Graphic, text, and audio generation includes the following typical tasks and their corresponding evaluation results:
[0094] Text-to-audio-video generation: Subjective evaluation of relevance, integrity, and effectiveness;
[0095] Audio-video-to-text description generation: Subjective evaluation;
[0096] Text and audio generation includes the following typical tasks and their corresponding evaluation results::
[0097] Speech synthesis: subjective evaluation;
[0098] Speech recognition: subjective evaluation;
[0099] Speech translation: subjective evaluation.
[0100] The subjective evaluation is as follows:
[0101] Evaluation content and dimensions:
[0102] The MOS score is used to evaluate the performance of the large model, including the following indicator dimensions: relevance, integrity, effectiveness, coherence, consistency, compliance, authenticity, and harmfulness;
[0103] Relevance: The degree of association between the answer and the conversation context;
[0104] Integrity: Whether there is any missing information in the generated answer;
[0105] Effectiveness: The usefulness of the generated answer;
[0106] Coherence: Whether the answer conforms to the conversation flow;
[0107] Consistency: Whether the answers are consistent when the same question is tested multiple times;
[0108] Compliance: Whether the model output meets the requirements of the question, including aspects such as content and form;
[0109] Authenticity: Whether the answer content is true and valid or contains false information that violates scientific common sense or basic facts;
[0110] Harmfulness: Whether the answer content contains content that violates basic moral ethics and laws;
[0111] Evaluation method:
[0112] Method 1: Perform a weighted average on the overall score of each piece of data, and evaluate the excellence of the model through the obtained overall evaluation score;
[0113] Method 2: Perform a weighted average on the MOS scores of each piece of data according to the indicator dimensions, and separately obtain the average scores of relevance, integrity, effectiveness, and coherence.
[0114] The number of samples in the test data set is 400 - 1000.
[0115] The evaluation methods include automated testing, manual testing, and using a large model as a referee for testing.
[0116] The specific evaluation results are as follows:
[0117] Accuracy:
[0118] Among them, TP is the number of samples that the model predicts as positive examples and are actually positive examples; FP is the number of samples that the model predicts as positive examples but are actually negative examples; TN is the number of samples that the model predicts as negative examples and are actually negative examples; FN is the number of samples that the model predicts as negative examples but are actually positive examples;
[0119] Recall rate:
[0120] Precision rate:
[0121] F1 value:
[0122] BLEU index:
[0123] Among them, Count clip (n-gram) represents the number of a certain n-gram in the reference, Count clip (n-gram′ represents the number of n-gram′ in the candidate; n-gram is a fragment composed of consecutive n words;
[0124] Rouge-L index:
[0125]
[0126]
[0127] Among them, the length of lcs is the number of words of the longest common substring of the true answer text and the model prediction text; the true answer text is the true label text of the test set samples; the model prediction text is the label text output by the model for the test set samples; β is a parameter with a default setting of 1.
[0128] The above is only the preferred implementation manner of the present invention, which is only used to help understand the method and its core idea of the present application. The protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.
[0129] The above embodiments can be implemented in whole or in part by software and / or hardware. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0130] The present invention solves the drawbacks in the prior art such as unclear standards, aimless development, insufficient application capabilities and general capabilities caused by the lack in the application and evaluation aspects of artificial intelligence large models. It provides a specific, detailed and highly applicable evaluation method for the general capabilities of artificial intelligence large models as a whole. Moreover, each dimension can be evaluated individually, in combination or crosswise according to the actual product requirements, with high generality and adaptability, which plays a guiding and helpful role in the development, application and expansion of artificial intelligence large models.
Claims
1. An evaluation and calculation method for artificial intelligence understanding and generation capabilities, characterized in that The large model includes two dimensions: understanding and generation; S1. The evaluation of the understanding ability is divided into a unimodal understanding dimension and a multimodal understanding dimension. The unimodal understanding dimension includes three secondary dimensions: text understanding, image understanding, and audio understanding; The multimodal understanding dimension includes four secondary dimensions: text-image understanding, text-audio understanding, image-audio understanding, and text-image-audio understanding; Further divide the secondary dimensions in S1 into multiple typical tasks and evaluate them separately. The evaluation method is as follows: By constructing a comprehensive test dataset, use programmable test tools and test statistical tools to input the dataset into the test tools and obtain the evaluation results; S2. The evaluation of the generation ability is divided into a unimodal generation ability and a multimodal generation ability. The unimodal generation ability includes a text generation secondary dimension; The multimodal generation ability includes three secondary dimensions: text-image generation, text-image-audio generation, and text-audio generation; Further divide the secondary dimensions in S2 into multiple typical tasks and evaluate them separately. The evaluation method is as follows: By constructing a comprehensive test dataset, use programmable test tools and test statistical tools to input the dataset into the test tools and obtain the evaluation results.
2. The evaluation and calculation method for the artificial intelligence understanding and generation ability according to claim 1, characterized in that The text understanding includes the following typical tasks and their corresponding evaluation results: Text classification: accuracy; Information extraction: accuracy, recall, and F1 value; Mathematical reasoning: accuracy; Causal reasoning: accuracy; Commonsense reasoning: accuracy, recall, and F1 value; Task decomposition: accuracy; Text answering: accuracy; Multi-turn dialogue: subjective evaluation; Code understanding: accuracy; Long text understanding: accuracy, recall, and F1 value; The image understanding includes the following typical tasks and their corresponding evaluation results: Static image classification: accuracy; Static image segmentation: segmentation accuracy, boundary accuracy; Object detection: precision, recall, and F1 value; Dynamic image classification: accuracy; Action recognition: accuracy; The audio understanding includes the following typical tasks and their corresponding evaluation results: Speaker recognition: accuracy, recall, and F1 value; Audio question answering: accuracy; Ambient sound classification: accuracy.
3. The evaluation and calculation method for the artificial intelligence understanding and generation ability according to claim 1, characterized in that The text-image understanding includes the following typical tasks and their corresponding evaluation results: Text-image retrieval, static image question answering, visual spatial relationship, visual language reasoning, visual entailment, video retrieval, video question answering: accuracy; Chart reasoning: accuracy and readability; The text-audio understanding: The typical task is text-audio retrieval: accuracy; The image-audio understanding includes: The typical task is video anomaly detection: accuracy; The text-image-audio understanding includes the following typical tasks and their corresponding evaluation results: Audio-visual video retrieval: accuracy, recall, and F1 value; Audio-visual video question answering: accuracy.
4. The evaluation and calculation method for the artificial intelligence understanding and generation ability according to claim 1, wherein The text generation includes the following typical tasks and their corresponding evaluation results: Abstract summarization: subjective evaluation; Machine translation: BLEU metric; Text rewriting: subjective evaluation; Text expansion: subjective evaluation; Text continuation: subjective evaluation; Code generation: subjective evaluation; Semi-structured data generation: subjective evaluation; The text-image generation includes the following typical tasks and their corresponding evaluation results: Text generation for images: Subjective evaluation of relevance, completeness, and effectiveness; Image generation for text descriptions: Manual evaluation; Text generation for videos: Subjective evaluation of relevance, completeness, and effectiveness; Video generation for text descriptions: Subjective evaluation; The text, image, and audio generation includes the following typical tasks and their corresponding evaluation results: Text generation for audio-visual videos: Subjective evaluation of relevance, completeness, and effectiveness; Audio-visual video generation for text descriptions: Subjective evaluation; The text and audio generation includes the following typical tasks and their corresponding evaluation results: Speech synthesis: Subjective evaluation; Speech recognition: Subjective evaluation; Speech translation: Subjective evaluation.
5. The evaluation calculation method for the artificial intelligence understanding and generation ability according to claim 4, characterized in that, The subjective evaluation is as follows: Evaluation content and dimensions: The MOS score is used to evaluate the performance of the large model, including the following indicator dimensions: relevance, completeness, effectiveness, coherence, consistency, compliance, authenticity, and harmfulness; The relevance: The degree of association between the answer and the dialogue context; The completeness: Whether there is any missing information in the generated answer; The effectiveness: The usefulness of the generated answer; The coherence: Whether the answer conforms to the dialogue flow; The consistency: Whether the answers are consistent when the same question is tested multiple times; The compliance: Whether the model output meets the requirements of the question, including aspects of content and form; The authenticity: Whether the answer content is true and valid or contains false information that violates scientific common sense or basic facts; The harmfulness: Whether the answer content contains content that violates basic moral ethics and laws; Evaluation method: Method 1: Perform a weighted average on the overall scores of each piece of data, and evaluate the excellence of the model through the obtained overall evaluation score; Method 2: Perform a weighted average on the MOS scores of each piece of data according to the indicator dimensions, and separately obtain the average scores of the relevance, the completeness, the effectiveness, and the coherence.
6. The evaluation and calculation method for artificial intelligence understanding and generation capabilities according to claim 1, characterized in that The number of samples in the test data set is 400 - 1000.
7. The evaluation and calculation method for the artificial intelligence understanding and generation ability according to claim 1, wherein The evaluation method includes automated testing, manual testing, and using a large model as a referee for testing.
8. The evaluation and calculation method for the artificial intelligence understanding and generation ability according to claim 1, characterized in that, The specific evaluation results are as follows: Accuracy: Among them, TP is the number of samples where the model predicts a positive example and is actually a positive example; FP is the number of samples where the model predicts a positive example but is actually a negative example; TN is the number of samples where the model predicts a negative example and is actually a negative example; FN is the number of samples where the model predicts a negative example but is actually a positive example; Recall rate: Precision: F1 value: BLEU metric: Among them, Count clip (n-gram) represents the number of a certain n-gram in the reference, Count clip (n-gram ′ ) represents the number of n-gram ′ in the candidate; n-gram is a fragment composed of consecutive n words; Rouge-L metric: Among them, the lcs length is the number of characters in the longest common substring of the true answer text and the model prediction text; the true answer text is the true label text of the test set sample; the model prediction text is the label text output by the model for the test set sample; β is a parameter with a default value of 1.
Citation Information
Patent Citations
Assessment method and device of large language model and computer equipment
CN118535443A