Intelligent model accuracy metering method
By using a standard reference dataset and a multi-index calculation method, the problem of the lack of standards in the measurement of intelligent models was solved, and a reliable assessment of the accuracy of intelligent models was achieved, thereby improving the uniformity and reliability of the level of intelligence.
Patent Information
- Application Number
- CN202511066782.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-11
AI Technical Summary
The lack of a unified standard for measuring the accuracy of existing intelligent models has led to uneven levels of intelligence across various industries, necessitating a unified method for measuring the accuracy of intelligent models.
Using a standard reference dataset as the measurement standard, the accuracy of the intelligent model is measured by calculating indicators such as accuracy, F1 score, precision, overlap, receiver operating characteristic curve, AUC, and recall.
This paper presents a unified method for measuring the accuracy of intelligent models, which solves the problem of the lack of physical standard instruments for measuring intelligent models and ensures the reliability and accuracy of intelligent model quality assessment.
Smart Images

Figure CN120929787A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of metrology technology, specifically relating to an intelligent model accuracy measurement method. Background Technology
[0002] The level of intelligence in existing intelligent models directly affects their application effectiveness. To evaluate the level of intelligence, it is necessary to first evaluate the intelligent model itself.
[0003] Currently, various industries face uneven levels of intelligence, and there is an urgent need to measure the accuracy of their intelligent models.
[0004] However, there are many problems in the measurement and testing of the accuracy of intelligent models, including but not limited to the lack of unified standards for related expressions, nomenclature, terminology, and measurement indicators; and there is currently no industry-recognized or consensus-reaching method for testing the accuracy of intelligent models.
[0005] Among the various performance metrics of intelligent models, accuracy is a key indicator. The accuracy of an intelligent model determines its basic level of intelligence. Establishing accuracy testing methods for intelligent models based on metrological digitization and digital metrology technologies is a way to address this issue. Summary of the Invention
[0006] (a) Technical problems to be solved
[0007] The technical problem to be solved by the present invention is to provide a method for measuring the accuracy of an intelligent model, which proposes to use a standard reference dataset as input to the intelligent model and to calculate a set of indicators for measuring the accuracy of the model proposed by the method.
[0008] (II) Technical Solution
[0009] To address the aforementioned technical problems, this invention provides an intelligent model accuracy measurement method, comprising the following steps:
[0010] Step 1: Define the assessment task
[0011] First, clarify the evaluation object and application scenario, determine the user evaluation needs and evaluation object, and based on the analysis of the evaluation object and the user evaluation needs, clarify the evaluation task and select the evaluation method according to the evaluation task;
[0012] Step 2: Select evaluation indicators
[0013] Evaluation metrics: Seven evaluation metrics are used as the accuracy metrics for the intelligent model: accuracy, F1 score, precision, overlap, receiver operating characteristic curve, area under the receiver operating characteristic curve envelope, and recall.
[0014] Step 3: Prepare an assessment dataset that meets the following conditions.
[0015] a) The test dataset meets the requirements of the evaluation dataset;
[0016] b) The test dataset is matched with the intelligent model of the evaluation object;
[0017] c) The test dataset is diverse;
[0018] d) The test dataset does not contain data from the model training dataset;
[0019] e) The test dataset contains a certain amount of interference data and adversarial data;
[0020] Step 4: Prepare assessment resources and environmental conditions
[0021] To prepare for the assessment task and its implementation, relevant resources and environment should be allocated.
[0022] Step 5: Conduct the assessment
[0023] Start the evaluation in the configured evaluation resources and environment, input the prepared evaluation dataset into the intelligent model to be tested, monitor the evaluation process and obtain process data and result data;
[0024] Step 6: Data Acquisition
[0025] Collect and save the data generated during the evaluation process;
[0026] Step 7: Calculate the evaluation indicators
[0027] Based on the indicator calculation formula, substitute the evaluation dataset and the output results of the intelligent model under test to calculate the value of the evaluation indicator.
[0028] The present invention also provides a system for implementing the method.
[0029] (III) Beneficial Effects
[0030] This invention provides a method for measuring the accuracy of intelligent models. This method selects a reference dataset that conforms to metrological testing specifications and has metrological attributes as the standard metrological instrument for intelligent model metrology, thus solving the problem that there is no suitable physical metrological standard instrument in the field of intelligent model metrology. Attached Figure Description
[0031] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0032] To make the objectives, contents, and advantages of the present invention clearer, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples.
[0033] If the standard measuring instruments used to measure the accuracy of intelligent models are not tested, their quality cannot be guaranteed, and intelligent models lack physical measuring standards. Therefore, this invention proposes using the test dataset, i.e., the reference dataset, used to measure the accuracy of intelligent models as a measuring standard in the process of measuring the accuracy of intelligent models; and proposes a set of evaluation indicators for measuring the accuracy of intelligent models, including: accuracy, F1-measure (F1 score), precision, IoU (overlap), ROC (Receiving Operating Characteristic), AUC (Area of Envelope of Receiving Operating Characteristic), and recall.
[0034] This invention uses different reference datasets (test datasets that conform to metrological testing standards and have metrological attributes, and which are used as reference datasets after metrological evaluation) as inputs for different types of intelligent models. After the intelligent model performs calculations and inferences on the corresponding reference datasets, the following metrological evaluation indicators for the accuracy of the intelligent model are calculated: accuracy, F1-measure, precision, IoU (overlap), ROC (Receiving Operating Characteristic), AUC (Area of Envelope of Receiving Operating Characteristic), and recall. Users determine the accuracy level of the intelligent model based on the values of these indicators.
[0035] refer to Figure 1 The present invention specifically includes the following steps:
[0036] Step 1: Define the assessment task
[0037] First, clarify the evaluation object and application scenario, and understand the user's evaluation needs and evaluation object, such as the specific type of intelligent model: algorithm type, operating environment, development framework, algorithm language and structure, etc. Based on the analysis of the evaluation object and evaluation needs, clarify the evaluation task, and select the evaluation method according to the evaluation task.
[0038] Step 2: Select evaluation indicators
[0039] Evaluation metrics: Seven evaluation metrics are used as the accuracy metrics for the intelligent model: accuracy, F1-measure (F1 score), precision, IoU (overlap), ROC (Receiver Operating Characteristic curve), AUC (Area of envelope of Receiver Operating Characteristic curve), and recall.
[0040] Step 3: Prepare the assessment dataset (standard reference dataset)
[0041] a) The test dataset must meet the requirements of the evaluation dataset;
[0042] b) The test dataset must match the intelligent model of the evaluation object;
[0043] c) The test dataset should be diverse;
[0044] d) The test dataset does not contain data from the model training dataset;
[0045] e) The test dataset must contain a certain amount of interference data and adversarial data, such as OOD data (data outside the target domain), noisy data, rotated or transformed images, etc.
[0046] Step 4: Prepare assessment resources and environmental conditions
[0047] To prepare for the assessment task, relevant resources and environment are configured: assessment data, assessment indicators, assessment methods, assessment models, software and hardware operating environments, etc.
[0048] Step 5: Conduct the assessment
[0049] The evaluation begins in the configured evaluation resources and environment. The prepared evaluation dataset is input into the intelligent model to be tested. The evaluation process is monitored and the process data and result data are obtained.
[0050] Step 6: Data Acquisition
[0051] Collect and save the data generated during the implementation of the evaluation.
[0052] Step 7: Calculate the evaluation indicators
[0053] Based on the indicator calculation formula, substitute the evaluation dataset and the output results of the intelligent model under test to calculate the value of the evaluation indicator.
[0054] Step 8: Comprehensive Evaluation of Results
[0055] The final evaluation results are obtained by combining the indicator evaluation methods.
[0056] Step 9, Evaluation Report
[0057] Based on the comprehensive evaluation process and results, a corresponding evaluation report will be issued.
[0058] The following is a (standardized) assessment example:
[0059] A.1 File Format
[0060] When the interface functions are developed using C / Python, multithreading is supported, and they can be compiled into 32-bit or 64-bit versions.
[0061] A.2 Interface Functions
[0062] The evaluation interface functions are shown in Table A.1.
[0063] Table A.1
[0064]
[0065] A.3 Interface Function Description
[0066] Definitions of each test interface:
[0067] (1) Dataset import interface:
[0068] Number: 1
[0069] Interface name: Evaluation_inputdata
[0070] Interface functionality: Used to enable the integration of assessment datasets.
[0071] Interface input parameter symbol: inputdata
[0072] Interface input parameter meaning: Evaluation dataset
[0073] Interface input parameter data type: image data or other data types
[0074] Communication protocols: TCP / IP protocol, HTTP protocol
[0075] (2) Evaluation index selection and calculation interface:
[0076] Number: 2
[0077] Interface name: Evaluation_inputindex
[0078] Interface function: Select and calculate credibility evaluation indicators and implement the calculation process.
[0079] Interface input parameter symbol: cal_index
[0080] Interface input parameter meaning: Calculate evaluation indicators
[0081] Interface input parameter data type: None
[0082] Communication protocols: TCP / IP protocol, HTTP protocol
[0083] (3) Interface for selecting evaluation methods:
[0084] Number: 3
[0085] Interface name: Evaluation_inputmethod
[0086] Interface function: Used to import the selected evaluation method. Here, you can select an evaluation method for evaluation.
[0087] Interface input parameter symbol: cal_method
[0088] Meaning of interface input parameters: Select evaluation method
[0089] Interface input parameter data type: None
[0090] Communication protocols: TCP / IP protocol, HTTP protocol
[0091] (4) Interface for introducing assessment objects:
[0092] Number: 4
[0093] Interface name: Evaluation_inputmodel
[0094] Interface Functionality: Used to connect to the evaluation object: intelligent model
[0095] Interface input parameter symbol: input_model
[0096] Meaning of interface input parameters: Introducing the evaluation model
[0097] Interface input parameter data type: None
[0098] Communication protocols: TCP / IP protocol, HTTP protocol
[0099] (5) Output and display the results:
[0100] Number: 5
[0101] Interface name: Evaluation_outputresult
[0102] Interface function: Used to output and display the calculated index results.
[0103] Interface output parameter symbol: out_result
[0104] Interface output parameter meaning: Output and display metrics.
[0105] Interface output parameter data type: string
[0106] Communication protocols: TCP / IP protocol, HTTP protocol
[0107] (6) Result storage:
[0108] Number: 6
[0109] Interface name: Evaluation_storageresult
[0110] Interface function: Used to store data such as calculated evaluation indicators.
[0111] Interface output parameter symbol: store_result
[0112] Meaning of interface output parameters: Storage result
[0113] Interface output parameter data type: string
[0114] Communication protocols: TCP / IP protocol, HTTP protocol
[0115] (7) Results visualization:
[0116] Number: 7
[0117] Interface name: Evaluation_Visualizationresult
[0118] Interface Functionality: Used to input the calculation results of evaluation indicators into a visualization tool for visual display.
[0119] Interface output parameter symbol: vision_result
[0120] Meaning of interface output parameters: Evaluation results
[0121] Interface output parameter data type: Image data
[0122] Communication protocols: TCP / IP protocol, HTTP protocol. A.4 Uncertainty assessment A.4.1 Environmental conditions:
[0123] Temperature: 21.0℃, Relative humidity: 42%.
[0124] A.4.2 Assessment Dataset:
[0125] Temperature and humidity dial dataset.
[0126] A.4.3 The object under test:
[0127] Intelligent temperature and humidity meter dial reading recognition model.
[0128] A.4.4 Measurement Method:
[0129] The temperature and humidity dial dataset is used as the evaluation dataset and input into the smart temperature and humidity meter dial reading recognition model. The evaluation method selected is the time-of-test (TTA) method. Evaluation indicators are calculated, the indicator calculation results are collected, and the accuracy is evaluated by analyzing the indicator evaluation method.
[0130] A.4.5 Sources of uncertainty:
[0131] a) Uncertainty component V1 introduced by the evaluation method
[0132] b) The intelligent model has an uncertainty component V2
[0133] c) The evaluation dataset contains an uncertainty component V3
[0134] d) Uncertainty component V4 introduced by the index evaluation method
[0135] The sum of the uncertainties of the final results is u i .
[0136] A.4.6 Standard Uncertainty Assessment
[0137] The intelligent temperature and humidity meter dial reading recognition model itself has uncertain properties, and the temperature and humidity meter dial dataset has uncertainty. The choice of evaluation method will introduce uncertainty in the evaluation process, and different index evaluation methods will also introduce uncertainty.
[0138] A.4.7 Calculation of Combined Standard Uncertainty
[0139] A summary of the uncertainty components is shown in Table A.1.
[0140] Table A.1 Summary of Uncertainty Components
[0141]
[0142] The combined standard uncertainty is:
[0143]
[0144] A.4.8 Expanded Uncertainty A.4.8.1 Calculating Expanded Uncertainty
[0145] Take a confidence probability p = 95%, ν eff As we approach infinity, we find k = 2 by consulting the t-distribution table.
[0146] Expanded uncertainty U = ku c =2 × 0.32 = 0.64
[0147] The relative expanded uncertainty is:
[0148] A.4.8.2 Uncertainty of Other Intelligent Models
[0149] The uncertainty of the accuracy evaluation results of the intelligent model can be obtained by using the above method.
[0150] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for measuring the accuracy of an intelligent model, characterized in that, Includes the following steps: Step 1: Define the assessment task First, clarify the evaluation object and application scenario, determine the user evaluation needs and evaluation object, and based on the analysis of the evaluation object and user evaluation needs, clarify the evaluation task and select the evaluation method according to the evaluation task; Step 2: Select evaluation indicators Evaluation metrics: Seven evaluation metrics are used as the accuracy metrics for the intelligent model: accuracy, F1 score, precision, overlap, receiver operating characteristic curve, area under the receiver operating characteristic curve envelope, and recall. Step 3: Prepare an assessment dataset that meets the following conditions. a) The test dataset meets the requirements of the evaluation dataset; b) The test dataset is matched with the intelligent model of the evaluation object; c) The test dataset is diverse; d) The test dataset does not contain data from the model training dataset; e) The test dataset contains a certain amount of interference data and adversarial data; Step 4: Prepare assessment resources and environmental conditions To prepare for the assessment task and its implementation, relevant resources and environment should be allocated. Step 5: Conduct the assessment Start the evaluation in the configured evaluation resources and environment, input the prepared evaluation dataset into the intelligent model to be tested, monitor the evaluation process and obtain process data and result data; Step 6: Data Acquisition Collect and save the data generated during the evaluation process; Step 7: Calculate the evaluation indicators Based on the indicator calculation formula, substitute the evaluation dataset and the output results of the intelligent model under test to calculate the value of the evaluation indicator.
2. The method as described in claim 1, characterized in that, The method also includes step 8, comprehensive evaluation of results: combining the indicator evaluation method to obtain the final evaluation result.
3. The method as described in claim 1, characterized in that, The method also includes step 9, obtaining the evaluation report: Based on the comprehensive evaluation process and evaluation results, a corresponding evaluation report is obtained.
4. The method as described in claim 1, characterized in that, User evaluation requirements and evaluation objects include the specific types of intelligent models: algorithm type, operating environment, development framework, algorithm language, and structure.
5. The method as described in claim 1, characterized in that, The test dataset includes data outside the target domain, noisy data, and rotated or transformed images.
6. The method as described in claim 1, characterized in that, The relevant resources and environment configuration includes preparing evaluation data, evaluation indicators, evaluation methods, evaluation models, and software and hardware operating environments.
7. The method as described in claim 1, characterized in that, The test dataset consists of standard measuring instruments used for intelligent model measurement.
8. The method as described in claim 1, characterized in that, This method is applied in the field of metrology.
9. A system for implementing the method as described in any one of claims 1 to 8.
10. The system as described in claim 9, characterized in that, This system is applied in the field of metrology.