Model evaluation method, device, computer equipment and storage medium
By obtaining the feature information and category overlap parameters of the machine learning model training data, the problem of unstable evaluation accuracy in the prior art is solved, and the stability and accuracy of model evaluation are achieved.
Patent Information
- Application Number
- CN202110875966.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-07-30
AI Technical Summary
Existing machine learning model evaluation methods rely on model structure and parameters, resulting in unstable evaluation accuracy, especially when the distribution of validation sets and test sets is inconsistent.
By obtaining the feature information and data labels of the model training data, classifying and calculating category overlap parameters, and evaluating the category overlap of the model training data independently of the model structure and parameters to determine the model evaluation results.
The stability of machine learning model evaluation is achieved, the dependence on model structure and parameters is avoided, and the accuracy and consistency of evaluation is improved.
Smart Images

Figure CN113822326B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a model evaluation method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of computer technology, machine learning technology has emerged. Machine learning is the core of artificial intelligence technology and is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. A machine learning model refers to a model established based on machine learning technology, specifically including models such as decision trees and neural networks. In machine learning model building, the quality of training data labels affects the final model effect. At the same time, the test set often does not have real labels, and generally, a part of the validation set is divided from the training set to verify whether the effects of certain model optimizations such as model structure optimization become better. However, when the distributions of the validation set and the test set are inconsistent, the effect of the validation set often cannot reflect the effect of the final test set.
[0003] Currently, noise samples can be determined by monitoring the loss value of each sample in the training set. Finally, the label quality of the training data is determined according to the proportion of the noise samples, so as to evaluate the effect of the machine learning model. However, this evaluation method needs to evaluate the data quality based on the model training process and depends on the used model structure and model parameters, etc., resulting in unstable evaluation accuracy for different machine learning models. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a model evaluation method, apparatus, computer device, and storage medium that can improve the evaluation stability of machine learning models.
[0005] A model evaluation method, the method includes:
[0006] Obtain a model evaluation request and model training data corresponding to the model evaluation request, where the model training data includes data labels;
[0007] Extract feature information corresponding to the model training data;
[0008] Classify the model training data according to the data labels to obtain a classification result of the model training data;
[0009] Determine a class overlap parameter corresponding to the model training data according to the feature information and the classification result;
[0010] Obtain the training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data.
[0011] A model evaluation device, the device includes:
[0012] A request acquisition module, configured to obtain a model evaluation request and the model training data corresponding to the model evaluation request, where the model training data includes data labels;
[0013] A feature extraction module, configured to extract the feature information corresponding to the model training data;
[0014] A data classification module, configured to classify the model training data according to the data labels to obtain a model training data classification result;
[0015] An overlap parameter determination module, configured to determine the class overlap parameter corresponding to the model training data according to the feature information and the classification result;
[0016] A model evaluation module, configured to obtain the training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data.
[0017] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Obtain a model evaluation request and the model training data corresponding to the model evaluation request, where the model training data includes data labels;
[0019] Extract the feature information corresponding to the model training data;
[0020] Classify the model training data according to the data labels to obtain a model training data classification result;
[0021] Determine the class overlap parameter corresponding to the model training data according to the feature information and the classification result;
[0022] Obtain the training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data.
[0023] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0024] Obtain a model evaluation request and the model training data corresponding to the model evaluation request, where the model training data includes data labels;
[0025] Extract the feature information corresponding to the model training data;
[0026] Classify the model training data according to the data label to obtain the classification result of the model training data;
[0027] Determine the class overlap parameter corresponding to the model training data according to the feature information and the classification result;
[0028] Obtain the training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data.
[0029] The above model evaluation method, device, computer device and storage medium obtain a model evaluation request and the model training data corresponding to the model evaluation request, where the model training data includes data labels; extract the feature information corresponding to the model training data; classify the model training data according to the data label to obtain the classification result of the model training data; determine the class overlap parameter corresponding to the model training data according to the feature information and the classification result; obtain the training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data. In this application, after obtaining the model training data of the model, the class overlap parameter between the model training data is determined based on the label category and features of the model training data. When the overlap degree between the model training data of each class is higher, the model training is often more difficult, thus affecting the final model effect. Therefore, the training evaluation result corresponding to the model evaluation request can be obtained according to the class overlap parameter corresponding to the model training data. The model evaluation process of this application only needs to evaluate according to the model training data, without relying on the used model structure and model parameters, which can effectively ensure the stability of model evaluation. Brief Description of the Drawings
[0030] Figure 1 It is an application environment diagram of the model evaluation method in an embodiment;
[0031] Figure 2 It is a flowchart of the model evaluation method in an embodiment;
[0032] Figure 3 It is a schematic diagram of the class overlap degree in an embodiment;
[0033] Figure 4 It is a flowchart of the steps for obtaining the class overlap parameter in an embodiment;
[0034] Figure 5 It is a flowchart of the steps for determining the similarity between features in the same dimension according to the mean and variance in an embodiment;
[0035] Figure 6 It is a flowchart of the steps for obtaining the similarity between features in the same dimension in the model training data of each class according to the JS divergence in an embodiment;
[0036] Figure 7 Schematic flowchart of the step of obtaining the category overlap parameter according to the similarity in an embodiment;
[0037] Figure 8 Input page diagram of the model training and evaluation process in an embodiment;
[0038] Figure 9 Result diagram of the model training and evaluation process in an embodiment;
[0039] Figure 10 Input page diagram of the model optimization and evaluation process in an embodiment;
[0040] Figure 11 Result diagram of the model optimization and evaluation process in an embodiment;
[0041] Figure 12 Overall schematic flowchart of predicting the training effect of the model in an embodiment;
[0042] Figure 13 Structural block diagram of the model evaluation device in an embodiment;
[0043] Figure 14 Internal structure diagram of the computer device in an embodiment. Detailed implementation manners
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0045] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0046] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] The solution provided by the embodiments of this application relates to the field of machine learning (ML) in artificial intelligence. Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The solution of this application will be specifically described through the following embodiments:
[0048] The model evaluation method provided by this application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can send a model evaluation request and the model training data corresponding to the machine learning model to be evaluated to the server 104, so as to perform relevant machine learning model evaluations through the server 104 and obtain the evaluation results of the machine learning model. The server 104 then obtains the model evaluation request submitted by the terminal 102; obtains the original machine learning model evaluation results corresponding to the query terms of the model evaluation request; identifies the association relationship between the query terms and the original machine learning model evaluation results in the preset service knowledge graph; based on the association relationship, determines the target machine learning model evaluation results in the original machine learning model evaluation results. Then, the final machine learning model evaluation results are returned to the terminal 102. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers. In another embodiment, optionally, the machine learning model evaluation method of this application can also be applied to the terminal, and the user can directly execute this method on the terminal side. In one of the embodiments, multiple servers can form a blockchain, and the server is a node on the blockchain.
[0049] Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0050] The blockchain underlying platform can include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and the maintenance of the corresponding relationship between the real identity of the user and the blockchain address (permission management). And under authorized circumstances, it supervises and audits the transaction situations of certain real identities, and provides the rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices, used to verify the validity of business requests, and record them to the storage after reaching a consensus on the valid requests. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through the consensus algorithm (consensus management), transmits it to the shared ledger in a complete and consistent manner after encryption (network communication), and records and stores it; the smart contract module is responsible for the registration and issuance of contracts, as well as the triggering and execution of contracts. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), trigger the execution according to the logic of the contract terms by calling keys or other events, complete the contract logic, and at the same time provide the functions of contract upgrade and cancellation; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and the visual output of the real-time state during product operation, such as: alarm, monitoring network conditions, monitoring the health status of node devices, etc.
[0051] The platform product service layer provides the basic capabilities and implementation frameworks of typical applications. Developers can build on these basic capabilities and overlay the characteristics of the business to complete the blockchain implementation of the business logic. The application service layer provides application services based on the blockchain solution for business participants to use.
[0052] In one embodiment, as Figure 2 shown, a model evaluation method is provided. Taking the method applied to the Figure 1 server 104 as an example, it includes the following steps:
[0053] Step 201, obtain a model evaluation request and the model training data corresponding to the model evaluation request. The model training data includes data labels.
[0054] Among them, a model evaluation request refers to a request sent by the terminal 102 to the server 104 to request the server 104 to evaluate a specified machine learning model. Machine learning is a specialized study of how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. A machine learning model refers to a model established based on machine learning technology, specifically including models of types such as decision trees and neural networks. Model training data refers to the data used to train a machine learning model. A machine learning model can be understood as a function, and model training refers to using the existing data (i.e., model training data) to determine the parameters of the function through some methods (optimization or other methods). The function after the parameters are determined is the result of the training. Using the model is to substitute new data into the function for evaluation. The data label specifically refers to the thing that the machine learning model needs to predict. During the model training process, for supervised training, the thing to be predicted can be directly used as the label. For example, for a machine learning model used for classification, its label is specifically the classification result. In addition, the model training data also contains feature information. A feature is the variable data used for prediction when input into the machine learning model and can be regarded as the evidence for obtaining the final result. The main work of machine learning is to extract useful features from the original input data and then construct a mapping from the features to the labels based on the existing examples.
[0055] Specifically, when the operator on the terminal 102 side needs to evaluate a machine learning model that is about to be trained or has been trained, since the quality of the labels of the training data will affect the final model effect, the model evaluation request and the model training data corresponding to the model evaluation request can be directly submitted to the server 104. The server 104 can then evaluate the corresponding machine learning model based on the model training data corresponding to the model evaluation request, so as to estimate the final performance of the machine learning model.
[0056] Step 203, extract the feature information corresponding to the model training data.
[0057] Among them, as described above, feature information refers to the variable data used for prediction by the input machine learning model and can be regarded as the evidence for obtaining the final result. In a specific embodiment, the machine learning model evaluation method of the present application is specifically used for the evaluation of the machine learning model in the oral examination scenario. At this time, the model training data is audio and the corresponding manual labels. The features include text features and acoustic features. The text features mainly include semantic features, pragmatic features, keyword features, and text disfluency features. The keyword features mainly include extracting the keywords in the standard answer and the keywords in the answer content, and calculating the precision rate, recall rate, etc. The pragmatic features include the diversity of words in the answer content, the diversity of sentence patterns, and the grammatical accuracy of the answer content based on language model analysis. The semantic features include the topic features of the answer content, tf-idf features, etc. The acoustic features are mainly divided into pronunciation accuracy, pronunciation fluency, pronunciation rhythm, etc. The pronunciation accuracy refers to the pronunciation scores at the phoneme, word, sentence levels, etc. The pronunciation fluency includes the speech rate feature during the pronunciation process, features based on duration statistics such as the average duration of pronunciation segments, the average pause duration between pronunciation segments, etc. The pronunciation rhythm includes the evaluation of the pronunciation rhythm, the evaluation of the correctness of word stress in the sentence, the evaluation of the sentence boundary tone, etc.
[0058] Specifically, when performing machine learning model evaluation processing, the present application specifically designs an algorithm for measuring the overlap degree between different categories of data labels by analyzing the data labels and feature distributions of the model training data. Based on the calculated label category overlap parameter, the more the samples overlap between different categories, the worse the final model effect may be. Therefore, before evaluation, the feature information can be obtained first, and then the overlap degree calculation of the subsequent process can be performed based on the feature information.
[0059] Step 205: Classify the model training data according to the data labels to obtain the classification result of the model training data.
[0060] Among them, the data label specifically refers to the thing that the machine learning model needs to predict. During the model training process, for supervised training, the thing to be predicted can be directly used as the label. For example, for a machine learning model for classification, its label is specifically the classification result. Therefore, when performing machine learning model evaluation, the model training data can be classified according to the data labels. The model training data under the same category theoretically corresponds to the same prediction result (it may occur that the machine learning model classification is incorrect, resulting in different classification results for the model training data with the same label).
[0061] Specifically, this application designs an algorithm for measuring the overlap degree between different categories of data labels by analyzing the data labels and feature distributions of the model training data. Based on the calculated label category overlap parameter, the more overlapping the samples between different categories are, the worse the final model effect may be. Therefore, before calculating the label category overlap parameter, it is necessary to first classify the model training data according to the data labels, divide different types of model training data into different categories, and then calculate the category overlap parameters corresponding to the model training data of different categories.
[0062] Step 207: Determine the category overlap parameter corresponding to the model training data according to the feature information and the classification result.
[0063] Among them, the category overlap parameter is used to characterize the overlap degree between data labels of different categories, and the category overlap parameter can be specifically determined by the correlation of the feature information between the model training data of different categories.
[0064] Specifically, after classifying the model training data into different categories according to the classification result, the feature sets corresponding to the model training data of different categories can be established based on the feature information corresponding to the model training data, and then the category overlap parameter corresponding to the model training data can be determined according to the feature overlap degree between the feature sets, so as to evaluate the machine learning model.
[0065] Step 209: Obtain the training evaluation result corresponding to the model evaluation request according to the category overlap parameter corresponding to the model training data.
[0066] Among them, the training evaluation result specifically refers to the result obtained by training the machine learning model with the model training data. Generally speaking, the more overlapping the samples between different categories are, the worse the final model effect may be. Therefore, the higher the category overlap parameter is, the worse the training evaluation result corresponding to the model evaluation request is. Specifically, as Figure 3 shown, the data in the figure includes three categories 1, 2, and 3. The overlap degree of the three categories in the left figure is relatively high, and the overlap degree of the three categories in the right figure is relatively low. When the overlap degree between categories is higher, the model training is often more difficult, thus affecting the final model effect.
[0067] Specifically, after obtaining the category overlap parameter, the training evaluation result corresponding to the model evaluation request can be obtained according to the category overlap parameter corresponding to the model training data. In a specific embodiment, the threshold of the category overlap parameter can be set according to the actual model requirements. When the category overlap parameter is higher than or equal to the preset threshold, the training evaluation result corresponding to the model evaluation request is set to be of poor effect, and when the category overlap parameter is lower than the preset threshold, the training evaluation result corresponding to the model evaluation request is set to be of good effect.
[0068] The above model evaluation method includes: obtaining a model evaluation request and model training data corresponding to the model evaluation request, where the model training data includes data labels; extracting feature information corresponding to the model training data; classifying the model training data according to the data labels to obtain a classification result of the model training data; determining a class overlap parameter corresponding to the model training data according to the feature information and the classification result; and obtaining a training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data. In this application, after obtaining the model training data of a machine learning model, the class overlap parameter between the model training data is determined based on the label categories and features of the model training data. When the overlap degree between the model training data of each class is higher, the model training is often more difficult, thus affecting the final model effect. Therefore, the training evaluation result corresponding to the model evaluation request can be obtained according to the class overlap parameter corresponding to the model training data. The model evaluation process of this application only needs to evaluate according to the model training data, without relying on the used model structure and model parameters, which can effectively ensure the stability of the machine learning model evaluation.
[0069] In one embodiment, as Figure 4 shown, step 207 includes:
[0070] Step 401: Determine the classification category corresponding to the model training data according to the classification result.
[0071] Step 403: Obtain the feature representation of the model training data according to the feature information corresponding to the model training data.
[0072] Step 405: Obtain the mean and variance of each feature in the model training data of each class according to the feature representation.
[0073] Step 407: Determine the similarity between the same-dimensional features in the model training data of each class according to the mean and variance.
[0074] Step 409: Obtain the class overlap parameter corresponding to the model training data according to the similarity.
[0075] Among them, the classification category refers to the classification category information to which each model training data belongs. The model training data can be classified according to the classification result, so as to obtain the corresponding classification category of each model training data. The feature representation can specifically be a vector, representing the combination of all features within the model training data of a classification category. The similarity between features in the same dimension can specifically be represented by the distance parameter between each feature under each category and each feature in other categories. In one embodiment, the distance parameter can specifically be the Jensen-Shannon divergence. In other embodiments, the distance parameter can also be represented by the Kullback-Leibler divergence, etc.
[0076] Specifically, for calculating the category overlap parameter between model training data of different categories, it can specifically be carried out by calculating and determining the similarity between features in the same dimension in the model training data of each category. First, the model training data can be divided according to the classification result corresponding to the data label. Specifically, each classification category label contains m l_i samples, where l_i represents the label as i. Then, for the model training data, the model training data under this category can be represented by the corresponding feature information of the model training data. For example, for the model training data category F(l i ):
[0077] F(l i ) = [F li1 , F li2 ... F lim
[0078] F lij = [f1, f2, f3... fk]
[0079] Among them, F li1 represents that the j-th model training data under the data label i contains k features.
[0080] Then, calculate the mean and variance of each feature in each category. The specific formulas are as follows:
[0081] mean(F li ) = [mean(f1), mean(f2),.... mean(fk)]
[0082] std(F li ) = [std(f1), std(f2),.... std(fk)]
[0083] Among them, mean(F li ) represents the mean set of each feature in the data label i, and std(Fli ) represents the set of variances of each feature in data tag i.
[0084] Then, based on the calculated mean and variance, determine the similarity between the features of the same dimension in the model training data of each category; and obtain the category overlap parameter corresponding to the model training data according to the similarity. In this embodiment, first obtain the feature representation of the model training data, and then calculate the mean and variance of each feature based on the feature representation, so as to effectively calculate the similarity between the features of the same dimension in the model training data of each category, and determine the final category overlap parameter, which can effectively ensure the accuracy of the calculation of the category overlap parameter.
[0085] In one of the embodiments, as Figure 5 shown, step 407 includes:
[0086] Step 502, based on the mean and variance, perform random sampling on the model training data of each category according to the Gaussian distribution, and obtain the feature distribution of each feature in the model training data of each category.
[0087] Step 504, determine the similarity between the features of the same dimension in the model training data of each category according to the feature distribution.
[0088] Among them, the Gaussian distribution is also called the normal distribution or the normal state distribution. The Gaussian distribution has extremely extensive practical backgrounds, and the probability distributions of many random variables in production and scientific experiments can be approximately described by the Gaussian distribution. In this application, sampling is mainly performed through the Gaussian distribution to determine the feature distribution, so as to calculate the similarity between the features of the same dimension based on the feature distribution, and reduce the complexity of the calculation process and improve the efficiency of machine learning model evaluation.
[0089] Specifically, when determining the similarity hypothesis between the features of the same dimension in the model training data of each category, it can be assumed that each feature conforms to the Gaussian distribution. Therefore, based on the mean and variance of each feature under each category obtained in step 305, perform random sampling according to the Gaussian distribution to sample the distribution of this feature under this category. Then calculate the distance parameter between each feature under each category and each feature of other categories. In this embodiment, random sampling is performed through the Gaussian distribution to obtain the feature distribution of the feature, and the similarity between the features of the same dimension in the model training data of each category is determined based on the feature distribution, which can effectively ensure the accuracy of the similarity calculation and improve the calculation efficiency of the calculation process at the same time.
[0090] In one of the embodiments, the feature distribution includes the probability density function of each feature in the model training data of each category under the feature distribution, as Figure 6 shown, step 504 includes:
[0091] Step 601: Determine the Jensen-Shannon (JS) divergence between the features of the same dimension in the model training data of each category according to the probability density function.
[0092] Step 603: Obtain the similarity between the features of the same dimension in the model training data of each category according to the JS divergence.
[0093] Among them, the JS divergence measures the similarity between two probability distributions and is a variant based on the KL divergence, which solves the problem of the asymmetry of the KL divergence. Generally, the JS divergence is symmetric and its value ranges from 0 to 1. The specific definition is as follows:
[0094]
[0095] Where p(x) is the probability density function of the features under one category, and q(x) is the probability density function of the features under another category. Finally, the distance vector of the features of category i and category j can be obtained.
[0096] Specifically, the JS divergence between the features of the same dimension in the model training data of each category can be determined according to the probability density function, and then the similarity between the features of the same dimension in the model training data of each category can be obtained according to the JS divergence. After calculating the JS divergence, the distance vector of the features of category i and category j can be obtained according to the following formula:
[0097] D ij =[d ij1 , d ij2 ... d ijk
[0098] d ijk =JSD(p ik ||p jk )
[0099] Where d ijk represents the distance between the model training data of category i and category j in the k-th dimension feature, and also their similarity in the k-th dimension feature. p ik and p jk respectively represent the probability density functions of the model training data of category i and category j in the k-th dimension feature. In this embodiment, the similarity between the features of the same dimension in the model training data of each category is determined by the JS divergence, which can effectively ensure the accuracy of the similarity calculation.
[0100] In one embodiment, as Figure 7 shown, step 409 includes:
[0101] Step 702: Obtain the discrimination degree corresponding to the model training data of each category according to the similarity.
[0102] Step 704: Obtain the average discrimination degree corresponding to the model training data of all categories to obtain the category overlap parameter corresponding to the model training data.
[0103] Among them, the discrimination degree is used to represent the overlap degree between the model training data of the current category and the model training data of other categories. The higher the discrimination degree, the smaller the overlap degree.
[0104] Specifically, when calculating the category overlap parameter, the discrimination degree corresponding to the model training data of each category can be calculated in turn, and then the final category overlap parameter can be obtained based on the average discrimination degree value of all categories. First, when calculating the discrimination degree, select the similarity of the feature with the largest distance between each category and other categories as the discrimination degree between this category and other categories, that is, Max(D ij ). Finally, the discrimination degree values between each category and other categories are obtained, and the average is calculated to obtain the average discrimination degree corresponding to the model training data of category i.
[0105] Di = mean([max( i1 ), max( i2 )…max(D ik )])
[0106] Then calculate the average discrimination degree value of all categories as the final category overlap parameter class_overlap.
[0107] class_overlap = mean([D1, D2, … Dk])
[0108] In this embodiment, by using the similarity to obtain the discrimination degree corresponding to the model training data of each category, and then estimating the category overlap parameter based on the average discrimination degree, the accuracy of calculating the category overlap parameter can be effectively guaranteed.
[0109] In one of the embodiments, after step 209, it further includes: testing the machine learning model corresponding to the model evaluation request; obtaining the model test data corresponding to the model evaluation request and the classification result data corresponding to the model test data; using the model test data as the model training data, using the classification result data as the training data classification result, and obtaining the test evaluation result corresponding to the model evaluation request.
[0110] Among them, the structure of the model test data is similar to that of the model training data. The difference is that the model training data is used to train the machine learning model, while the model test data is used to test the trained machine learning model. The model training data includes data labels, while the model test data does not include labels. The availability of the trained machine learning model can be judged through the model test data. After the model test data is input into the machine learning model, corresponding classification and recognition results can be obtained. Therefore, the machine learning model evaluation method of the present application can also be used to evaluate the test process of the machine learning model. During the evaluation, the model test data can be used as the model training data, and the classification result data can be used as the classification result of the training data, where the classification result data is equivalent to the posterior result obtained after inputting the model training data into the machine learning model, and the classification result of the training data is the prior classification result obtained according to the prior model label. Here, the posterior result is used to represent the prior result, so that label data does not need to be used during the evaluation process of model testing. Then, steps similar to those of the independent claim are executed. First, the feature information corresponding to the model training data is extracted to obtain the pre-optimization class overlap parameter corresponding to these model test data; the pre-optimization test evaluation result corresponding to the model evaluation request is obtained according to the pre-optimization class overlap parameter.
[0111] Specifically, the machine learning model evaluation method of the present application can also evaluate the test process of the machine learning model. Specifically, this process can refer to the above evaluation process of the model training process. First, the feature information corresponding to the model test data is extracted; the model test data is classified according to the classification result data corresponding to the model test data to obtain the classification result of the model test data; the class overlap parameter corresponding to the model test data is determined according to the feature information and the classification result of the model test data; the pre-optimization test evaluation result corresponding to the model evaluation request is obtained according to the class overlap parameter corresponding to the model test data. The specific steps of the above process can refer to the corresponding embodiments in the above content. In this embodiment, for the test process of the machine learning model, the pre-optimization class overlap parameter corresponding to the model test data is calculated through the model test data and the corresponding classification result data, so as to estimate the pre-optimization test evaluation result, and the test process of the model can be effectively evaluated.
[0112] In one embodiment, after obtaining the test evaluation result corresponding to the model evaluation request, it further includes: optimizing the machine learning model; testing the optimized machine learning model to obtain the classification result optimization data corresponding to the model evaluation request; using the model test data as the model training data and the classification result optimization data as the classification result of the training data to obtain the post-optimization test evaluation result corresponding to the model evaluation request; obtaining the model optimization evaluation result according to the test evaluation result and the post-optimization test evaluation result.
[0113] Specifically, the model evaluation device in this application can also be applied to the evaluation of the model optimization effect. By comparing the optimization evaluation results corresponding to the machine learning model before and after optimization, it is determined whether the model optimization has achieved the optimization effect. When evaluating the model optimization effect, first obtain the test evaluation result before optimization. Then, input the model test data into the optimized machine learning model to obtain the classification result optimization data. Then, similar to the above evaluation process for model training, use the model test data as the model training data and the classification result optimization data as the training data classification result, and re-execute the step of extracting the feature information corresponding to the model training data to obtain the optimized category overlap parameter corresponding to the model test data. By comparing the category overlap parameter before optimization and the category overlap parameter after optimization, the effect before and after model optimization is judged. If the category overlap parameter before optimization is higher than the category overlap parameter after optimization, it means that after the model is optimized, the category overlap parameter has decreased and the effect of the model has been improved. If the category overlap parameter before optimization is lower than the category overlap parameter after optimization, it means that after the model is optimized, the category overlap parameter has increased and the effect of the model has decreased. In this embodiment, based on the prediction results of the test set before model optimization and the prediction results of the test set after model optimization, calculate the change in the category overlap degree before and after optimization, and use this to measure whether the effect before and after model optimization has been improved, which can effectively identify the optimization effect of the model.
[0114] In one embodiment, the model evaluation method of this application is used in the field of machine learning model evaluation in the oral exam scenario, and is mainly applied to the oral English exam training engine. The oral exam includes objective question types such as reading aloud question types and open question types such as talking about pictures. In application, it can be specifically applied to the evaluation of the model training process and the evaluation of the model optimization process. For the model training process, specifically refer to Figure 8 and Figure 9 , as Figure 8 shown, the user can click the button to upload data, and then upload the corresponding training data by selection. As Figure 9 shown, after the machine learning model evaluation, the data quality can be directly output. Through the data quality, the model effect prediction corresponding to the machine learning model can be directly obtained. For the model optimization process, it can refer to Figure 10 and Figure 11 . By sequentially inputting the model training data, classification result data, and classification result optimization data, and then clicking the Figure 10 model optimization button, as Figure 11 shown, the corresponding model optimization result can be directly obtained.
[0115] The specific evaluation process can refer to Figure 12, first, obtain the training data and test data corresponding to the machine learning model, then extract the features of the training data and test data through the feature extraction module. Then, combine the labels and features of the model training data to evaluate the data quality corresponding to the model training process, and obtain the training data quality and the evaluation result of the model training process. For the model optimization process, it is necessary to evaluate the optimization effect of the model by comparing the test results and features of the model before optimization with the test results and features of the model after optimization.
[0116] In a specific embodiment, the present application is used to test the base models of different question types in an oral examination, specifically including scenario question types, rapid response question types, oral compositions, and semi-open question types. Each question type includes at least two questions, and each question includes 250 training samples and 1400 test samples. The evaluation indicators are the Pearson correlation coefficient and the agreement rate (i.e., the probability that the label and the model prediction value are less than a certain threshold). Calculate the correlation degree and agreement rate between the predicted values of the test set of the base model and the manual labels as the model effect of the base model. Tune the parameters of the model and change the model structure. Look at the correlation degree and agreement rate between the predicted values of the test set of the optimized model and the manual labels as the model effect of the optimized model. If the correlation degree and agreement rate improve before and after, it means that the model effect has improved. Use this value as the label for the improvement of the model effect, rising to 1 if it rises and remaining 0 if it does not. Use the category overlap degree analysis method of the present invention to analyze the change in the category overlap degree before and after the test set. If the category overlap degree improves and the distance between categories becomes smaller, the prediction model effect decreases, that is, the predicted value is 0, otherwise it is 1. Finally, calculate the precision and recall rate. The final precision rate reaches 80%, and the recall rate reaches 100%. Based on the category overlap degree in the present application, it is possible to estimate whether the model effect deteriorates after optimization.
[0117] It should be understood that although Figures 2 - 7 the steps in the flowchart of Figures 2 - 7 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0118] In an embodiment, as Figure 13As shown, a model evaluation device is provided. This device can be implemented as a software module, a hardware module, or a combination of both, and becomes part of a computer device. Specifically, the device includes: a request acquisition module 1302, a feature extraction module 1304, a data classification module 1306, an overlap parameter determination module 1308, and a model evaluation module 1310, where:
[0119] The request acquisition module 1302 is configured to obtain a model evaluation request and model training data corresponding to the model evaluation request. The model training data includes data labels.
[0120] The feature extraction module 1304 is configured to extract feature information corresponding to the model training data.
[0121] The data classification module 1306 is configured to classify the model training data according to the data labels to obtain a classification result of the model training data.
[0122] The overlap parameter determination module 1308 is configured to determine a class overlap parameter corresponding to the model training data according to the feature information and the classification result.
[0123] The model evaluation module 1310 is configured to obtain a training evaluation result corresponding to the model evaluation request according to the class overlap parameter corresponding to the model training data.
[0124] In one embodiment, the overlap parameter determination module 1308 is specifically configured to: determine the classification classes corresponding to the model training data according to the classification result; obtain the feature representation of the model training data according to the feature information corresponding to the model training data; obtain the mean and variance of each feature in the model training data of each class according to the feature representation; determine the similarity between the same-dimensional features in the model training data of each class according to the mean and variance; and obtain the class overlap parameter corresponding to the model training data according to the similarity.
[0125] In one embodiment, the overlap parameter determination module 1308 is further configured to: randomly sample the model training data of each class based on the Gaussian distribution according to the mean and variance to obtain the feature distribution of each feature in the model training data of each class; and determine the similarity between the same-dimensional features in the model training data of each class according to the feature distribution.
[0126] In one embodiment, the overlap parameter determination module 1308 is further configured to: determine the JS divergence between the same-dimensional features in the model training data of each class according to the probability density function; and obtain the similarity between the same-dimensional features in the model training data of each class according to the JS divergence.
[0127] In one embodiment, the overlapping parameter determination module 1308 is further configured to: obtain the discrimination degree corresponding to the model training data of each category according to the similarity; obtain the average discrimination degree corresponding to the model training data of all categories, and obtain the category overlapping parameter corresponding to the model training data.
[0128] In one embodiment, it further includes a model optimization evaluation module, configured to: test the machine learning model corresponding to the model evaluation request; obtain the model test data corresponding to the model evaluation request and the classification result data corresponding to the model test data; use the model test data as the model training data, use the classification result data as the data label, and obtain the test evaluation result corresponding to the model evaluation request.
[0129] In one embodiment, the model optimization evaluation module is further configured to: optimize the machine learning model; test the optimized machine learning model to obtain the classification result optimization data corresponding to the model evaluation request; use the model test data as the model training data, use the classification result optimization data as the training data classification result, and obtain the optimized test evaluation result corresponding to the model evaluation request; obtain the model optimization evaluation result according to the test evaluation result and the optimized test evaluation result.
[0130] For the specific definition of the model evaluation device, reference may be made to the definition of the model evaluation method in the foregoing text, which will not be elaborated herein. Each module in the foregoing model evaluation device may be implemented in whole or in part by software, hardware, and their combination. The foregoing modules may be embedded in or independent of the processor in the computer device in the form of hardware, or may be stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the foregoing modules.
[0131] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 14 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store model evaluation data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a model evaluation method.
[0132] Those skilled in the art can understand, Figure 14The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0133] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0134] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0135] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the steps in the above method embodiments.
[0136] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0138] The above embodiments only illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A model evaluation method, characterized in that, The method is applied to the evaluation of machine learning models in the scenario of oral examinations, including: Obtaining a model evaluation request and model training data corresponding to the model evaluation request, where the model training data includes data labels, the model training data includes audio, and the data labels include manual labels corresponding to the audio; Extracting feature information corresponding to the model training data, where the feature information includes text features and acoustic features; Classifying the model training data according to the data labels to obtain a classification result of the model training data; Determining a classification category corresponding to the model training data according to the classification result; obtaining a feature representation of the model training data according to the feature information corresponding to the model training data; obtaining the mean and variance of each feature in the model training data of each category according to the feature representation; determining the similarity between the same-dimensional features in the model training data of each category according to the mean and the variance; obtaining a category overlap parameter corresponding to the model training data according to the similarity; Obtaining a training evaluation result corresponding to the model evaluation request according to the category overlap parameter corresponding to the model training data.
2. The method according to claim 1, characterized in that, The determining the similarity between the same-dimensional features in the model training data of each category according to the mean and the variance includes: Randomly sampling the model training data of each category based on the Gaussian distribution according to the mean and the variance to obtain a feature distribution of each feature in the model training data of each category; Determining the similarity between the same-dimensional features in the model training data of each category according to the feature distribution.
3. The method according to claim 2, wherein The feature distribution includes the probability density function of each feature in the model training data of each category under the feature distribution; The determining the similarity between the same-dimensional features in the model training data of each category according to the feature distribution includes: Determining the Jensen-Shannon divergence between the same-dimensional features in the model training data of each category according to the probability density function; Obtaining the similarity between the same-dimensional features in the model training data of each category according to the Jensen-Shannon divergence.
4. The method according to claim 1, wherein The obtaining the category overlap parameter corresponding to the model training data according to the similarity includes: Obtaining the discrimination degree corresponding to the model training data of each category according to the similarity; Obtaining the average discrimination degree corresponding to the model training data of all categories to obtain the category overlap parameter corresponding to the model training data.
5. The method according to claim 1, characterized in that After obtaining the training evaluation result corresponding to the model evaluation request according to the category overlap parameter corresponding to the model training data, it further includes: Testing the machine learning model corresponding to the model evaluation request; Obtaining the model test data corresponding to the model evaluation request and the classification result data corresponding to the model test data; Taking the model test data as the model training data and the classification result data as the training data classification result to obtain the test evaluation result corresponding to the model evaluation request.
6. The method according to claim 5, wherein After obtaining the test evaluation result corresponding to the model evaluation request, it further includes: Optimizing the machine learning model; Test the optimized machine learning model to obtain the classification result optimization data corresponding to the model evaluation request; Use the model test data as model training data, and use the classification result optimization data as the training data classification result to obtain the optimized test evaluation result corresponding to the model evaluation request; Obtain the model optimization evaluation result according to the test evaluation result and the optimized test evaluation result.
7. A model evaluation device, characterized in that, The device is applied to the evaluation of a machine learning model in an oral examination scenario, including: A request acquisition module, configured to obtain a model evaluation request and the model training data corresponding to the model evaluation request. The model training data includes data labels, the model training data includes audio, and the data labels include manual labels corresponding to the audio; A feature extraction module, configured to extract the feature information corresponding to the model training data. The feature information includes text features and acoustic features; A data classification module, configured to classify the model training data according to the data labels to obtain a model training data classification result; An overlapping parameter determination module, configured to determine the classification category corresponding to the model training data according to the classification result; obtain the feature representation of the model training data according to the feature information corresponding to the model training data; obtain the mean and variance of each feature in the model training data of each category according to the feature representation; determine the similarity between the same-dimensional features in the model training data of each category according to the mean and the variance; obtain the category overlapping parameter corresponding to the model training data according to the similarity; A model evaluation module, configured to obtain the training evaluation result corresponding to the model evaluation request according to the category overlapping parameter corresponding to the model training data.
8. The device according to claim 7, characterized in that The overlapping parameter determination module is specifically configured to: perform random sampling on the model training data of each category based on the Gaussian distribution according to the mean and the variance to obtain the feature distribution of each feature in the model training data of each category; determine the similarity between the same-dimensional features in the model training data of each category according to the feature distribution.
9. The device according to claim 8, characterized in that, The feature distribution includes the probability density function of each feature in the model training data of each category under the feature distribution; the overlapping parameter determination module is further configured to: determine the JS divergence between the same-dimensional features in the model training data of each category according to the probability density function; obtain the similarity between the same-dimensional features in the model training data of each category according to the JS divergence.
10. The device according to claim 7, characterized in that, The overlapping parameter determination module is specifically configured to: obtain the discrimination degree corresponding to the model training data of each category according to the similarity; obtain the average discrimination degree corresponding to the model training data of all categories to obtain the category overlapping parameter corresponding to the model training data.
11. The device according to claim 7, characterized in that It further includes a model optimization evaluation module, configured to: test the machine learning model corresponding to the model evaluation request; obtain the model test data corresponding to the model evaluation request and the classification result data corresponding to the model test data; Use the model test data as model training data, use the classification result data as the classification result of the training data, and obtain the test evaluation result corresponding to the model evaluation request.
12. The device according to claim 11, wherein, The model optimization evaluation module is further configured to: optimize the machine learning model; test the optimized machine learning model to obtain classification result optimization data corresponding to the model evaluation request; Use the model test data as model training data, use the classification result optimization data as the classification result of the training data, and obtain the optimized test evaluation result corresponding to the model evaluation request; Obtain a model optimization evaluation result according to the test evaluation result and the optimized test evaluation result.
13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Vehicle abnormal behavior detection method based on Hidden Markov Model
CN103235933A
Model multi-terminal collaborative training method and medical risk prediction method and device
CN110797124A