A method, related device, equipment and storage medium for model training

By sorting and reorganizing the activation values ​​of predicted value vectors in the training of classification network models, the problem of gradient attenuation and category aggregation in traditional model training is solved, and the training effect and recognition ability of the model are improved.

CN113392868BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110049736.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-14
Publication Date
2025-05-30
Estimated Expiration
2041-01-14

AI Technical Summary

Technical Problem

In traditional classification network model training, the output value of the softmax function depends on the difference size of elements in the predicted value vector, resulting in the model's aggregation in some categories, resulting in poor training effect.

Method used

During the model training process, the activation values ​​of label elements and non-label elements in the predicted value vector set are sorted and reorganized to generate the recombinant predicted value vector set, thereby increasing the backpropagation gradient and improving the model training effect.

Benefits of technology

Through the sorting and reorganization of activation values, the effectiveness and recognition ability of model training are improved, ensuring that the model obtains a larger backpropagation gradient during the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113392868B_ABST
    Figure CN113392868B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method implemented based on artificial intelligence technology, including: obtaining a set of data to be trained; obtaining a set of predicted value vectors through a classification network model based on the set of data to be trained, and the absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among the M non-label elements is a first value; sorting the label element and the M non-label elements of each predicted value vector to obtain a set of reorganized predicted value vectors, and the absolute value of the difference between the activation value of the label element in the target reorganized predicted value vector and the maximum activation value among the M non-label elements is a second value, and the second value is less than the first value; training the classification network model according to the set of reorganized predicted value vectors. In the present application, the activation values of the label elements and the activation values of the non-label elements in the set of predicted value vectors are sorted and reorganized to obtain predicted value vectors with closer activation values, that is, the greater the backpropagation gradient, the better the training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of machine learning, and in particular, to a method for model training, related devices, equipment, and storage media. Background Art

[0002] Neural networks have broad and attractive prospects in fields such as system identification, pattern recognition, and intelligent control. A neural network is a machine learning technology that simulates the working principle of the human brain to achieve artificial intelligence similar to humans. It can process various types of data such as images, texts, voices, and sequences, and can achieve classification, regression, prediction, etc. Therefore, how to train a high-quality network model has become an issue worthy of attention.

[0003] In the training of traditional classification network models, the softmax layer is often used as the mapping function of the network activation layer. The softmax function maps the activation values in the predicted value vector into a probability distribution with a sum of 1. Using this probability distribution and the true class label corresponding to the input data, a loss function can be constructed to train the classification network model.

[0004] The output value of the softmax function depends on the difference size of the elements in the predicted value vector. If the activation value of the j-th element in the predicted value vector is greater than the activation values of other elements, the probability value corresponding to the j-th element is close to 1, while the probability values corresponding to other elements are close to 0. If the true class label corresponding to the j-th element is 1 at this time, the gradient that can be backpropagated by the loss function is also relatively small, and the classification network model tends to be stable. However, in this case, the fitting of the classification network model to the data is not optimal. For example, there will be problems such as the aggregation of some categories not being tight enough, resulting in poor model training effects. Summary of the Invention

[0005] Embodiments of this application provide a method for model training, related devices, equipment, and storage media. During the model training process, sorting and reorganizing the activation values of the label elements and non-label elements in the predicted value vector set can obtain a predicted value vector with more similar activation values. Therefore, the greater the backpropagated gradient, to achieve a better model training effect, and then improve the recognition ability of the model.

[0006] In view of this, on the one hand, this application provides a method for model training, including:

[0007] Obtain a set of data to be trained, where the set of data to be trained includes at least two pieces of data to be trained, and each piece of data to be trained has the same true class label;

[0008] Based on the set of data to be trained, a set of predicted value vectors is obtained through a classification network model. Among them, the set of predicted value vectors includes at least two predicted value vectors. Each predicted value vector corresponds to a piece of data to be trained. Each predicted value vector includes a label element and M non-label elements. The set of predicted value vectors includes a target predicted value vector. The absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among the M non-label elements is a first value, and M is an integer greater than or equal to 1;

[0009] Sort the label element and the M non-label elements of each predicted value vector in the set of predicted value vectors to obtain a set of reorganized predicted value vectors. Among them, the set of reorganized predicted value vectors includes a target reorganized predicted value vector. The absolute value of the difference between the activation value of the label element in the target reorganized predicted value vector and the maximum activation value among the M non-label elements is a second value, and the second value is less than the first value;

[0010] Train the classification network model according to the set of reorganized predicted value vectors.

[0011] On the other hand, the present application provides a model training device, including:

[0012] An acquisition module, configured to acquire a set of data to be trained. Among them, the set of data to be trained includes at least two pieces of data to be trained, and each piece of data to be trained has the same true class label;

[0013] The acquisition module is further configured to, based on the set of data to be trained, obtain a set of predicted value vectors through a classification network model. Among them, the set of predicted value vectors includes at least two predicted value vectors. Each predicted value vector corresponds to a piece of data to be trained. Each predicted value vector includes a label element and M non-label elements. The set of predicted value vectors includes a target predicted value vector. The absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among the M non-label elements is a first value, and M is an integer greater than or equal to 1;

[0014] A sorting module, configured to sort the label element and the M non-label elements of each predicted value vector in the set of predicted value vectors to obtain a set of reorganized predicted value vectors. Among them, the set of reorganized predicted value vectors includes a target reorganized predicted value vector. The absolute value of the difference between the activation value of the label element in the target reorganized predicted value vector and the maximum activation value among the M non-label elements is a second value, and the second value is less than the first value;

[0015] A training module, configured to train the classification network model according to the set of reorganized predicted value vectors.

[0016] In a possible design, in another implementation manner of the other aspect of the embodiments of the present application,

[0017] An acquisition module, specifically configured to acquire an initial training data set, where the initial training data set includes Q training data to be trained, each of the Q training data to be trained has a true class label, and Q is an integer greater than or equal to 2;

[0018] From the initial training data set, at least two training data to be trained with the same true class label are acquired as the training data set to be trained.

[0019] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0020] A sorting module, specifically configured to, for each label element of each prediction value vector in the prediction value vector set, arrange a column of activation values corresponding to the label elements in the prediction value vector set in descending order to obtain a first sequence;

[0021] For each of the M non-label elements of each prediction value vector in the prediction value vector set, arrange the M columns of activation values corresponding to the M non-label elements in the prediction value vector set in ascending order to obtain M second sequences;

[0022] According to the first sequence and the M second sequences, a recombined prediction value vector set is generated.

[0023] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0024] A sorting module, specifically configured to obtain a first sequence according to the label elements of each prediction value vector in the prediction value vector set;

[0025] For each of the M non-label elements of each prediction value vector in the prediction value vector set, arrange the M columns of activation values corresponding to the M non-label elements in the prediction value vector set in ascending order to obtain M second sequences;

[0026] According to the first sequence and the M second sequences, a recombined prediction value vector set is generated.

[0027] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0028] A sorting module, specifically configured to, for each label element of each prediction value vector in the prediction value vector set, arrange a column of activation values corresponding to the label elements in the prediction value vector set in descending order to obtain a first sequence;

[0029] For each of the M non-label elements of each prediction value vector in the prediction value vector set, obtain the maximum activation value of the M non-label elements in each prediction value vector;

[0030] For each predicted value vector in the set of predicted value vectors, replace the activation values corresponding to the M non-label elements with the maximum activation value of the M non-label elements to obtain the updated predicted value vector corresponding to each predicted value vector;

[0031] Arrange the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0032] Generate a set of recombined predicted value vectors according to the first sequence and the M second sequences.

[0033] In a possible design, in another implementation manner of another aspect of the embodiment of the present application,

[0034] The sorting module is specifically configured to arrange, in descending order, a column of activation values corresponding to the label elements in each predicted value vector in the set of predicted value vectors to obtain a first sequence;

[0035] For the M non-label elements of each predicted value vector in the set of predicted value vectors, obtain the K maximum activation values of the M non-label elements in each predicted value vector, where the K maximum activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M;

[0036] For each predicted value vector in the set of predicted value vectors, replace the K minimum activation values among the M non-label elements with the K maximum activation values to obtain the updated predicted value vector corresponding to each predicted value vector, where the K minimum activation values represent the first K activation values after arranging the M non-label elements in ascending order;

[0037] Arrange the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0038] Generate a set of recombined predicted value vectors according to the first sequence and the M second sequences.

[0039] In a possible design, in another implementation manner of another aspect of the embodiment of the present application,

[0040] The training module is specifically configured to obtain N recombined predicted value vectors from the set of recombined predicted value vectors, where the N recombined predicted value vectors are the last 1 to N recombined predicted value vectors in the set of recombined predicted value vectors, and N is an integer greater than or equal to 1;

[0041] Update the network parameters of the classification network model by using a loss function according to the N recombined predicted value vectors.

[0042] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0043] The sorting module is specifically configured to sort the label elements of each prediction value vector in the prediction value vector set in ascending order to obtain a first sequence;

[0044] For the M non-label elements of each prediction value vector in the prediction value vector set, sort them in descending order to obtain M second sequences;

[0045] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0046] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0047] The sorting module is specifically configured to obtain a first sequence according to the label elements of each prediction value vector in the prediction value vector set;

[0048] For the M non-label elements of each prediction value vector in the prediction value vector set, sort them in descending order to obtain M second sequences;

[0049] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0050] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0051] The sorting module is specifically configured to, for the label elements of each prediction value vector in the prediction value vector set, sort a column of activation values corresponding to the label elements in the prediction value vector set in ascending order to obtain a first sequence;

[0052] For the M non-label elements of each prediction value vector in the prediction value vector set, obtain the maximum activation value of the M non-label elements in each prediction value vector;

[0053] For each prediction value vector in the prediction value vector set, replace the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain an updated prediction value vector corresponding to each prediction value vector;

[0054] Sort the M columns of activation values corresponding to the updated prediction value vectors corresponding to each prediction value vector in ascending order to obtain M second sequences;

[0055] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0056] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0057] A sorting module, specifically configured to, for each label element of the prediction value vectors in the prediction value vector set, sort a column of activation values corresponding to the label elements in the prediction value vector set in ascending order to obtain a first sequence;

[0058] For each of the M non-label elements of the prediction value vectors in the prediction value vector set, obtain the K maximum activation values of the M non-label elements in each prediction value vector, where the K maximum activation values represent the first K activation values after sorting the M non-label elements in descending order, and K is an integer greater than 1 and less than M;

[0059] For each prediction value vector in the prediction value vector set, replace the K minimum activation values among the M non-label elements with the K maximum activation values to obtain an updated prediction value vector corresponding to each prediction value vector, where the K minimum activation values represent the first K activation values after sorting the M non-label elements in ascending order;

[0060] Sort the M columns of activation values corresponding to the updated prediction value vector corresponding to each prediction value vector in ascending order to obtain M second sequences;

[0061] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0062] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0063] A training module, specifically configured to obtain N recombined prediction value vectors from the recombined prediction value vector set, where the N recombined prediction value vectors are the first to N recombined prediction value vectors in the recombined prediction value vector set, and N is an integer greater than or equal to 1;

[0064] Update the network parameters of the classification network model by using a loss function according to the N recombined prediction value vectors.

[0065] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;

[0066] Wherein, the memory is used to store a program;

[0067] The processor is used to execute the program in the memory, and the processor is used to execute the methods in the above aspects according to the instructions in the program code;

[0068] The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.

[0069] Another aspect of the present application provides a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute the methods of the above aspects.

[0070] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the methods provided in the above aspects.

[0071] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0072] In an embodiment of the present application, a method for model training is provided. First, a set of data to be trained is obtained, and then, based on the set of data to be trained, a set of predicted value vectors is obtained through a classification network model. The set of predicted value vectors includes a target predicted value vector, and the absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among M non-label elements is a first value. Then, the label element and M non-label elements of each predicted value vector in the set of predicted value vectors are sorted to obtain a set of reorganized predicted value vectors. The set of reorganized predicted value vectors includes a target reorganized predicted value vector, and the absolute value of the difference between the activation value of the label element in the target reorganized predicted value vector and the maximum activation value among M non-label elements is a second value. Finally, the classification network model is trained according to the set of reorganized predicted value vectors. Through the above method, during the model training process, the activation values of the label elements and non-label elements in the set of predicted value vectors are sorted and reorganized to obtain a set of reorganized predicted value vectors. And in the set of reorganized predicted value vectors obtained after reorganization, there is at least one target reorganized predicted value vector, and the second value corresponding to the target reorganized predicted value vector is smaller than the first value of a certain predicted value vector before sorting. That is to say, for the target reorganized predicted value vector after the activation value rearrangement, the included activation values are closer to each other. Therefore, the greater the backpropagation gradient, the fact that the backpropagation gradient increases means that the model has not tended to be stable and still needs to be continuously trained to achieve a better model training effect, thereby improving the recognition ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a schematic diagram for updating network parameters based on a classification network model in an embodiment of the present application;

[0074] Figure 2 It is a schematic diagram of an application environment of a model training system in an embodiment of the present application;

[0075] Figure 3Schematic diagram of an embodiment of the model training method in the present application embodiment;

[0076] Figure 4 Schematic diagram of rearranging activation values based on a binary classification network model in the present application embodiment;

[0077] Figure 5 Schematic diagram of rearranging activation values based on a multi - classification network model in the present application embodiment;

[0078] Figure 6 Schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0079] Figure 7 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0080] Figure 8 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0081] Figure 9 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0082] Figure 10 Schematic diagram of selecting N recombined predicted value vectors for model training in the present application embodiment;

[0083] Figure 11 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0084] Figure 12 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0085] Figure 13 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0086] Figure 14 Another schematic diagram of constructing a set of recombined predicted value vectors in the present application embodiment;

[0087] Figure 15 Another schematic diagram of selecting N recombined predicted value vectors for model training in the present application embodiment;

[0088] Figure 16 Schematic diagram of an embodiment of the model training device in the present application embodiment;

[0089] Figure 17 Schematic diagram of a structure of the server in the present application embodiment. Detailed implementation manners

[0090] Embodiments of the present application provide a method for model training, related devices, equipment, and storage media. During the process of model training, by sorting and reorganizing the activation values of the label elements and the non-label elements in the predicted value vector set, a predicted value vector with more similar activation values can be obtained. Therefore, the greater the gradient of backpropagation, the better the model training effect can be achieved, thereby improving the recognition ability of the model.

[0091] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0092] Data classification methods based on Machine Learning (ML) are a current research hotspot. For example, in the field of Computer Vision (CV), image classification methods are studied; in the field of Natural Language Processing (NLP), text classification methods are studied; in the field of Speech Technology, sound classification methods are studied, etc.

[0093] Machine learning is a technology implemented based on Artificial Intelligence (AI). Among them, machine learning is an interdisciplinary subject involving multiple fields, such as probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0094] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0095] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0096] In the process of data classification, a classification network model is often required. Therefore, how to train a classification network model with better classification effect has become a problem worthy of attention. In the process of training the classification network model, the softmax layer is used as the mapping function of the network activation layer. The softmax function maps the activation values in the predicted value vector into a probability distribution with a sum of 1. For ease of understanding, please refer to Figure 1 , Figure 1 This is a schematic diagram of updating network parameters based on the classification network model in the embodiments of this application. As shown in the figure, x1, x2, and x3 are input data. After passing through the classification network model, activation values z1, z2, and z3 are output. The activation values z1, z2, and z3 form a predicted value vector (z1, z2, z3). Next, the predicted value vector (z1, z2, z3) will be mapped into a probability distribution (a1, a2, a3) with a sum of 1 through the softmax layer (or softmax function).

[0097] The softmax layer (or softmax function) is used for binary classification or multi-classification and is often used as a multi-classifier in the last layer of a neural network. There are as many neurons in the last layer as there are classes. The class label of each sample is encoded in one-hot format, that is, encoded as a string of 0s and 1s, and the length of this string is the number of classes. If it belongs to the jth classification, a 1 is marked at the position of the jth element. For example, if there are three classes and the output y value is 3, its encoding is 001.

[0098] The function of softmax is expressed as:

[0099]

[0100] Among them, σ(z) j represents the probability value corresponding to the position of the j-th element, z = (z 1 , z 2 ,..., z K ) represents the predicted value vector output by the classification network model, k represents the k-th category, and K represents the total number of all categories.

[0101] For easy understanding, take a specific example for illustration. Suppose z1 is equal to 5, z2 is equal to 1, and z3 is equal to -1, then the probability values of these three are respectively:

[0102]

[0103]

[0104]

[0105] The calculation result is that σ(z) 1 = 0.9817, σ(z) 2 = 0.0180, σ(z) 3 = 0.0003, that is, the corresponding probability distribution is (0.9817, 0.0180, 0.0003), and the sum of the probability values in the probability distribution is 1.

[0106] It can be seen from the above formula that the output value of the softmax function depends on the difference in the activation values corresponding to each element in the predicted value vector. If the activation value of the j-th element in the predicted value vector is greater than the activation values of other elements, then the probability value corresponding to the j-th element is close to 1, while the probability values corresponding to other elements are close to 0. If the true class label corresponding to the j-th element is 1 at this time, the gradient that the loss function can backpropagate is also relatively small, and the classification network model tends to be stable. On the contrary, if the activation values of all elements in the predicted value vector are about the same, the backpropagated gradient is larger. In the case where the sample is correctly classified, a large difference in the activation values of elements means that the loss function will enter the saturation region, resulting in a relatively small backpropagated gradient. However, correct classification does not mean that the features are well learned. If the largest activation value is still small, it means that the model is not optimal.

[0107] For example, assume that the value range of zj is from -5 to 5, the predicted value vector a is (0, -5, -5), and the predicted value vector b is (5, 0, 0). The values of these two predicted value vectors after being mapped by the softmax layer are the same. If the true class labels corresponding to these two predicted value vectors are both 1, then their loss functions are the same in these two cases, and the gradients of their backpropagation both tend to 0. And the gradients of their backpropagation both tend to 0. However, when the true class label is 1, it is expected that the activation value z1 is as large as possible, indicating that the feature vector of the sample has a high consistency with the feature vector of the class in the classifier. In the above example, the activation value z1 of the predicted value vector b is better than that of the predicted value vector a. Although during the model training process, the value of the activation value z1 will continuously increase until it is greater than the activation value z2 and the activation value z3, when the difference reaches a certain magnitude, the model tends to be stable, the loss function enters the saturation region, and the magnitude of the activation value z1 will no longer change. Therefore, the predicted value vector a and the predicted value vector b are two stable states. During the model training process, these two states cannot be distinguished. For this reason, the present application proposes a model training method based on activation value rearrangement, which can solve the problem of gradient decay in the traditional model training process, thereby accelerating the training speed of the model, enabling the model to obtain a more appropriate backpropagation gradient, and improving the recognition accuracy and recall rate of the model.

[0108] The model training method proposed by the present application can be used in a model training system such as Figure 2 shown. Please refer to Figure 2 , Figure 2 which is a schematic diagram of an application environment of the model training system in an embodiment of the present application. As shown in the figure, the model training system may include at least one of a server and a terminal device. Taking the model training system including a server and a terminal device as an example, the terminal device can collect a large number of initial training data sets, and then transmit the initial training data sets to the server. The server classifies the initial training data sets to obtain two or more sets of data sets to be trained. Then, the server uses the method provided by the present application to process the data sets to be trained, and finally obtains a set of recombined predicted value vectors for model training.

[0109] The server involved in this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a handheld computer, a personal computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. The number of servers and terminal devices is also not restricted.

[0110] Combined with the above introduction, the method for model training in this application will be introduced below. The solution provided in the embodiments of this application involves technologies such as machine learning in artificial intelligence. Please refer to Figure 3 , an embodiment of the model training method in the embodiments of this application includes:

[0111] 101. Obtain a set of data to be trained, where the set of data to be trained includes at least two data to be trained, and each data to be trained has the same true class label;

[0112] In this embodiment, the model training device obtains a set of data to be trained. The set of data to be trained includes at least two data to be trained, and all the data to be trained in the set of data to be trained have corresponding true class labels. Among them, the true class label is usually a label obtained after manual annotation.

[0113] It should be noted that the model training device provided in this application can be deployed on a server, or on a terminal device, or on a model training system jointly composed of a server and a terminal device.

[0114] 102. Based on the set of data to be trained, obtain a set of predicted value vectors through a classification network model, where the set of predicted value vectors includes at least two predicted value vectors, each predicted value vector corresponds to a data to be trained, each predicted value vector includes a label element and M non-label elements, the set of predicted value vectors includes a target predicted value vector, and the absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among the M non-label elements is a first value, and M is an integer greater than or equal to 1;

[0115] In this embodiment, the model training device inputs a set of data to be trained into a classification network model, and outputs a predicted value vector corresponding to each piece of data to be trained through the classification network model, thereby obtaining a set of predicted value vectors. Suppose 300 pieces of data to be trained are input in one iteration of training, that is, taking a batch equal to 300 as an example, after inputting 300 pieces of data to be trained into the classification network model, 300 predicted value vectors are obtained. Among them, each predicted value vector includes a label element and M non-label elements, where M is an integer greater than or equal to 1.

[0116] Specifically, in one example, if M equals 1, it belongs to a binary classification scenario. Suppose the predicted value vector is (5, -5), and the true class label is "Classification A", that is, the element in the first position of the predicted value vector is the label element, and the activation value corresponding to this label element is "5", while the element in the second position is the non-label element, and the activation value corresponding to this non-label element is "-5". Suppose the predicted value vector is (-3, 2), and the true class label is "Classification B", that is, the element in the second position of the predicted value vector is the label element, and the activation value corresponding to this label element is "2", while the element in the first position is the non-label element, and the activation value corresponding to this non-label element is "-3".

[0117] In another example, if M equals 2, it belongs to a ternary classification scenario. Suppose the predicted value vector is (5, 0, -5), and the true class label is "Classification A", that is, the element in the first position of the predicted value vector is the label element, and the activation value corresponding to this label element is "5", while the elements in the second and third positions are non-label elements, and the activation values corresponding to the non-label elements are "5" and "-5" respectively. Suppose the predicted value vector is (-1, 3, -5), and the true class label is "Classification B", that is, the element in the second position of the predicted value vector is the label element, and the activation value corresponding to this label element is "3", while the elements in the first and third positions are non-label elements, and the activation values corresponding to the non-label elements are "-1" and "-5" respectively. Suppose the predicted value vector is (-2, -3, 5), and the true class label is "Classification C", that is, the element in the third position of the predicted value vector is the label element, and the activation value corresponding to this label element is "5", while the elements in the first and second positions are non-label elements, and the activation values corresponding to the non-label elements are "-2" and "-3" respectively.

[0118] It should be noted that in actual use, M can also take other values greater than 1. Here, M equals 1 or 2 is taken as an example for illustration. The processing methods for other values are similar to the above examples, so they are not enumerated here.

[0119] After obtaining the set of predicted value vectors, at least one target predicted value vector can be taken from it. Taking one target predicted value vector as an example, the absolute value of the difference between the activation value of the labeled element in the target predicted value vector and the maximum activation value among the M unlabeled elements is the first value. For example, if the target predicted value vector is (5, 0, 0), assuming the element in the first position is the labeled element with an activation value of "5", and the remaining M unlabeled elements are "0" and "0" with a corresponding maximum activation value of "0", then the first value is obtained as 5.

[0120] 103. Sort the labeled elements and the M unlabeled elements of each predicted value vector in the set of predicted value vectors to obtain a set of reorganized predicted value vectors. Among them, the set of reorganized predicted value vectors includes a target reorganized predicted value vector, and the absolute value of the difference between the activation value of the labeled element in the target reorganized predicted value vector and the maximum activation value among the M unlabeled elements is the second value, and the second value is less than the first value;

[0121] In this embodiment, the model training device needs to sort the labeled elements and the M unlabeled elements of each predicted value vector in the set of predicted value vectors. Taking 300 predicted value vectors as an example, then it is necessary to arrange the activation values corresponding to each labeled element in the 300 predicted value vectors, or arrange the activation values corresponding to each unlabeled element in the 300 predicted value vectors, or arrange both the activation values corresponding to each labeled element and the activation values corresponding to each unlabeled element in the 300 predicted value vectors. The arrangement methods include but are not limited to ascending order, descending order, and random arrangement, etc.

[0122] It should be noted that after rearranging the activation values of each predicted value vector in the set of predicted value vectors, a set of reorganized predicted value vectors is obtained, and there is at least one target reorganized predicted value vector in the set of reorganized predicted value vectors. Taking one target reorganized predicted value vector as an example, the absolute value of the difference between the activation value of the labeled element in the target reorganized predicted value vector and the maximum activation value among the M unlabeled elements is the second value. For example, if the target reorganized predicted value vector is (0.1, 0, 0), assuming the element in the first position is the labeled element with an activation value of "0.1", and the remaining M unlabeled elements are "0" and "0" with a corresponding maximum activation value of "0", then the second value is obtained as 0.1.

[0123] Since the probability values output by the softmax layer only depend on the magnitude of the differences within the input vector, therefore, when the second value is less than the first value, it means that the target reorganized predicted value vector has a larger backpropagation gradient compared to the target predicted value vector, enabling the model to learn better.

[0124] Specifically, in one example, taking the binary classification scenario as an example, and in one iteration of training, 2 data to be trained are input. For the convenience of introduction, please refer to Figure 4 , Figure 4 FIG. Figure 4 is a schematic diagram of rearranging activation values based on a binary classification network model in an embodiment of the present application. As shown in the figure, assuming that the true class label is "Classification A", that is, the element at the first position in the predicted value vector is the label element. Among them, the predicted value vector a is (0.1, -5), and the predicted value vector b is (5, 0). The label element in the predicted value vector a is rearranged with the label element in the predicted value vector b, that is, by swapping the activation value z1 in the two predicted value vectors, two new predicted value vectors are obtained, that is, the recombined predicted value vector a is (5, -5), and the recombined predicted value vector b is (0.1, 0). Considering that the difference values between the activation values corresponding to the elements in the recombined predicted value vector b are relatively small, therefore, a backpropagation gradient with an appropriate magnitude is provided during the training process, so that the model can learn more efficiently.

[0125] Taking the target predicted value vector as the predicted value vector a as an example, its corresponding first value is |0.1 - (-5)|, that is, equal to 5.1. There is at least one target recombined predicted value vector among the recombined predicted value vector a and the recombined predicted value vector b. Taking the target recombined predicted value vector as the recombined predicted value vector b as an example, its corresponding second value is less than the first value. Here, taking the recombined predicted value vector b as the target recombined predicted value vector, its corresponding second value is |0.1 - 0|, that is, equal to 0.1. Therefore, the predicted value vector a can be replaced with the recombined predicted value vector b for model training.

[0126] In another example, taking the multi-classification scenario as an example, for the convenience of introduction, please refer to Figure 5 , Figure 5 FIG. Figure 5 is a schematic diagram of rearranging activation values based on a multi-classification network model in an embodiment of the present application. As shown in the figure, assuming that the true class label is "Classification A", that is, the element at the first position in the predicted value vector is the label element. Among them, the predicted value vector a is (0.1, -5, -5), and the predicted value vector b is (5, 0, 0). The label element in the predicted value vector a is rearranged with the label element in the predicted value vector b, that is, by swapping the activation value z1 in the two predicted value vectors, two new predicted value vectors are obtained, that is, the recombined predicted value vector a is (5, -5, -5), and the recombined predicted value vector b is (0.1, 0, 0). Considering that the difference values between the activation values corresponding to the elements in the recombined predicted value vector b are relatively small, therefore, a backpropagation gradient with an appropriate magnitude is provided during the training process, so that the model can learn more efficiently.

[0127] Taking the target prediction value vector as the prediction value vector a as an example, the corresponding first value is |0.1 - (-5)|, which is equal to 5.1. There is at least one target reorganized prediction value vector in the reorganized prediction value vector a and the reorganized prediction value vector b. Taking the target reorganized prediction value vector as the reorganized prediction value vector b as an example, its corresponding second value is less than the first value. Here, taking the reorganized prediction value vector b as the target reorganized prediction value vector, its corresponding second value is |0.1 - 0|, which is equal to 0.1. Therefore, the prediction value vector a can be replaced with the reorganized prediction value vector b for model training.

[0128] 104. Train the classification network model according to the set of reorganized prediction value vectors.

[0129] In this embodiment, after obtaining the set of reorganized prediction value vectors, the model training device uses the set of reorganized prediction value vectors to replace the set of prediction value vectors to train the classification network model, that is, constructs a classification loss function using the set of reorganized prediction value vectors to train the classification network model.

[0130] Specifically, the backpropagation algorithm can be used to train the network parameters of the classification network model. Among them, the expression ability of the neural network for the model depends on the optimization algorithm. Optimization is a process of continuously calculating gradients and adjusting learnable parameters. In the training stage, after the forward propagation of the classification network model, there is a gap between the predicted label and the true class label obtained. The loss function can be used to reflect this gap. The role of the loss function can be understood as that when the predicted label obtained by forward propagation is close to the true class label, a smaller value is taken, otherwise the value increases. In the backpropagation process, the network parameters are continuously adjusted in a gradient descent manner. Gradient descent is a method for finding the minimum value of the loss function.

[0131] It can be understood that the classification network model can be an image classification model, or a text classification model or an audio classification model, etc., and can also be a model applied to image segmentation and image detection. The present application does not make a limitation.

[0132] In the embodiment of the present application, a method for model training is provided. Through the above method, in the process of model training, the activation values of the label elements and the non-label elements in the set of prediction value vectors are sorted and reorganized to obtain a set of reorganized prediction value vectors. And in the set of reorganized prediction value vectors obtained after reorganization, there is at least one target reorganized prediction value vector. The second value corresponding to this target reorganized prediction value vector is less than the first value of a certain prediction value vector before sorting. That is to say, for the target reorganized prediction value vector after the activation value rearrangement, the included activation values are closer to each other. Therefore, the greater the backpropagation gradient, the increase in the backpropagation gradient means that the model has not tended to be stable and still needs to be continuously trained to achieve a better model training effect, thereby improving the recognition ability of the model.

[0133] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, obtaining a set of data to be trained specifically includes:

[0134] Obtaining an initial training data set, where the initial training data set includes Q data to be trained, and each data to be trained in the Q data to be trained has a true class label, and Q is an integer greater than or equal to 2;

[0135] From the initial training data set, at least two data to be trained with the same true class label are obtained as the data set to be trained.

[0136] In this embodiment, a method of grouping the data set to be trained according to the true class label is introduced. The model training device obtains the initial training data set, and then classifies the initial training data set according to the true class label of each data to be trained in the initial training data set, and extracts at least two data to be trained with the same true class label as the data set to be trained.

[0137] Specifically, taking the three-classification scenario based on images as an example, first, all the data to be trained in the initial training data set are manually labeled. Please refer to Table 1, which is a schematic representation of the true class labels corresponding to the data to be trained in the initial training data set.

[0138] Table 1

[0139] Data to be trained True class label Data to be trained 1 Class A Data to be trained 2 Class A Data to be trained 3 Class B Data to be trained 4 Class B Data to be trained 5 Class C Data to be trained 6 Class B Data to be trained 7 Class A Data to be trained 8 Class C Data to be trained 9 Class A Data to be trained 10 Class B Data to be trained 11 Class C

[0140] As can be seen from Table 1, assuming that the initial training data set includes 11 data to be trained (i.e., Q is equal to 11), the initial training data set is grouped, and the data to be trained belonging to the same true class label are added to the same group of data to be trained. For example, the true class labels of data to be trained 1, data to be trained 2, data to be trained 7, and data to be trained 9 are all "class A", so these 4 data to be trained are used as the data set to be trained. Another example is that the true class labels of data to be trained 3, data to be trained 4, data to be trained 6, and data to be trained 10 are all "class B", so these 4 data to be trained are used as the data set to be trained. Another example is that the true class labels of data to be trained 5, data to be trained 8, and data to be trained 11 are all "class C", so these 3 data to be trained are used as the data set to be trained.

[0141] Secondly, in the embodiments of the present application, a method for grouping the training data set according to the true class labels is provided. Through the above method, the consistency of the true class labels can be ensured during the process of reordering the activation values. Therefore, the classification network model can be trained using the recombined prediction value vector set. During the training process, the true class labels corresponding to the prediction value vector set are used as the true class labels of the recombined prediction value vector set, that is, there is no need to relabel each recombined prediction value vector, thereby improving the training efficiency.

[0142] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, the label elements and M non-label elements of each prediction value vector in the prediction value vector set are sorted to obtain a recombined prediction value vector set, which specifically includes:

[0143] For the label elements of each prediction value vector in the prediction value vector set, a column of activation values corresponding to the label elements in the prediction value vector set is arranged in descending order to obtain a first sequence;

[0144] For the M non-label elements of each prediction value vector in the prediction value vector set, M columns of activation values corresponding to the M non-label elements in the prediction value vector set are arranged in ascending order to obtain M second sequences;

[0145] According to the first sequence and the M second sequences, a recombined prediction value vector set is generated.

[0146] In this embodiment, a recombination method of arranging the label elements in descending order and the non-label elements in ascending order is introduced. For each training data in each training data set, the activation values corresponding to the label elements in each training data are extracted, thereby obtaining a column of activation values, and then the column of activation values is arranged in descending order to obtain a first sequence. Similarly, the activation values corresponding to the other M non-label elements are extracted, thereby obtaining M columns of activation values, and then each column of activation values in the M columns of activation values is arranged in ascending order to obtain M second sequences. Finally, the first sequence and the M second sequences are combined, and the activation values under the same index are combined into a recombined prediction value vector until a recombined prediction value vector set is generated.

[0147] It should be noted that each prediction value vector has an index, and each recombined prediction value vector after rearrangement also has an index. The index of the prediction value vector and the index of the recombined prediction value vector have a corresponding relationship.

[0148] Specifically, for the sake of easy understanding, taking the three-classification scenario as an example (i.e., taking M equal to 2), and 3 training data are input in one iteration of training, please refer to Figure 6 ,Figure 6 This is a schematic diagram for constructing a set of recombinant predicted value vectors in an embodiment of the present application. As shown in the figure, assume that the true class label is "Classification A", that is, the element at the first position in the predicted value vector is the label element. Among them, the predicted value vector a is (0.2, -1, -1), the predicted value vector b is (0.1, -5, -5), and the predicted value vector c is (5, 0, 0). Rearrange the label elements in the predicted value vector a, the label elements in the predicted value vector b, and the label elements in the predicted value vector c, that is, rearrange the activation values z1 in the three predicted value vectors in descending order, and thus obtain the first sequence (5, 0.2, 0.1). Rearrange the non-label elements in the predicted value vector a, the non-label elements in the predicted value vector b, and the non-label elements in the predicted value vector c, that is, rearrange the activation values z2 and z3 in the three predicted value vectors in ascending order, and thus obtain two second sequences, namely the second sequence (-5, -1, 0) and the second sequence (-5, -1, 0).

[0149] Based on the corresponding positions of the activation value z1, the activation value z2, and the activation value z3, combine the first sequence and the two second sequences. Among them, the index value of the predicted value vector a is "a", and the index value of the corresponding recombinant predicted value vector a is also "a", and the recombinant predicted value vector a is (5, -5, -5). The index value of the predicted value vector b is "b", and the index value of the corresponding recombinant predicted value vector b is also "b", and the recombinant predicted value vector b is (0.2, -1, -1). The index value of the predicted value vector c is "c", and the index value of the corresponding recombinant predicted value vector c is also "c", and the recombinant predicted value vector c is (0.1, 0, 0).

[0150] Taking the target predicted value vector as the predicted value vector a as an example, its corresponding first value is |0.2 - (-1)|, that is, equal to 1.2. There is at least one target recombinant predicted value vector among the recombinant predicted value vector a, the recombinant predicted value vector b, and the recombinant predicted value vector c. Taking the target recombinant predicted value vector as the recombinant predicted value vector c as an example, its corresponding second value is |0.1 - 0|, that is, equal to 0.1, which satisfies the condition that the second value is less than the first value. Therefore, the predicted value vector a can be replaced with the recombinant predicted value vector c for model training.

[0151] Secondly, in the embodiment of the present application, a recombination method for arranging label elements in descending order and non-label elements in ascending order is provided. Through the above method, recombinant predicted value vectors with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further improving the recognition ability of the classification network model.

[0152] Optionally, in the aboveFigure 3 Based on the corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, the label elements and M non-label elements of each prediction value vector in the prediction value vector set are sorted to obtain a recombined prediction value vector set, which specifically includes:

[0153] Obtain a first sequence according to the label elements of each prediction value vector in the prediction value vector set;

[0154] For the M non-label elements of each prediction value vector in the prediction value vector set, arrange the M columns of activation values corresponding to the M non-label elements in the prediction value vector set in ascending order to obtain M second sequences;

[0155] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0156] In this embodiment, a recombination method of arranging the label elements in the original order and arranging the non-label elements in ascending order is introduced. For each training data in each training data set, extract the activation values corresponding to the label elements in each training data, thereby obtaining a column of activation values, and arrange the column of activation values in the original order to obtain a first sequence. Similarly, extract the activation values corresponding to the other M non-label elements, thereby obtaining M columns of activation values, and then arrange each column of activation values in the M columns of activation values in ascending order to obtain M second sequences. Finally, combine the first sequence and the M second sequences, and combine the activation values under the same index into a recombined prediction value vector until a recombined prediction value vector set is generated.

[0157] It should be noted that each prediction value vector has an index, and each recombined prediction value vector after rearrangement also has an index, and there is a corresponding relationship between the index of the prediction value vector and the index of the recombined prediction value vector.

[0158] Specifically, for the sake of easy understanding, taking the three-classification scenario as an example (i.e., taking M equal to 2), and 3 training data are input in one iteration training, please refer to Figure 7 , Figure 7Another schematic diagram for constructing a set of recombinant prediction value vectors in the embodiments of the present application is shown in the figure. Assume that the true class label is "Classification A", that is, the element at the first position in the prediction value vector is the label element. Among them, the prediction value vector a is (0.2, -1, -1), the prediction value vector b is (0.1, -5, -5), and the prediction value vector c is (5, 0, 0). Arrange the label elements in the prediction value vector a, the label elements in the prediction value vector b, and the label elements in the prediction value vector c, that is, arrange the activation values z1 in the three prediction value vectors in the original order, and thus obtain the first sequence (0.2, 0.1, 5). Rearrange the non-label elements in the prediction value vector a, the non-label elements in the prediction value vector b, and the non-label elements in the prediction value vector c, that is, rearrange the activation values z2 and z3 in the three prediction value vectors in ascending order, and thus obtain two second sequences, which are the second sequence (-5, -1, 0) and the second sequence (-5, -1, 0) respectively.

[0159] Based on the corresponding positions of the activation value z1, the activation value z2, and the activation value z3, combine the first sequence and the two second sequences. Among them, the index value of the prediction value vector a is "a", and the index value of the corresponding recombinant prediction value vector a is also "a", and the recombinant prediction value vector a is (0.2, -5, -5). The index value of the prediction value vector b is "b", and the index value of the corresponding recombinant prediction value vector b is also "b", and the recombinant prediction value vector b is (0.1, -1, -1). The index value of the prediction value vector c is "c", and the index value of the corresponding recombinant prediction value vector c is also "c", and the recombinant prediction value vector c is (5, 0, 0).

[0160] Taking the target prediction value vector as the prediction value vector a as an example, the corresponding first value is |0.2 - (-1)|, that is, equal to 1.2. There is at least one target recombinant prediction value vector among the recombinant prediction value vector a, the recombinant prediction value vector b, and the recombinant prediction value vector c. Taking the target recombinant prediction value vector as the recombinant prediction value vector b as an example, the corresponding second value is |0.1 - (-1)|, that is, equal to 1.1, which satisfies the condition that the second value is less than the first value. Therefore, the prediction value vector a can be replaced with the recombinant prediction value vector b for model training.

[0161] Secondly, in the embodiments of the present application, a recombination method for arranging the label elements in the original order and arranging the non-label elements in ascending order is provided. Through the above method, recombinant prediction value vectors with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further improving the recognition ability of the classification network model.

[0162] Optionally, in the above Figure 3Based on the corresponding embodiments, in another alternative embodiment provided by the embodiments of the present application, sorting the label elements and M non-label elements of each predicted value vector in the set of predicted value vectors to obtain a set of recombined predicted value vectors may include:

[0163] For the label element of each predicted value vector in the set of predicted value vectors, arranging a column of activation values corresponding to the label elements in the set of predicted value vectors in descending order to obtain a first sequence;

[0164] For the M non-label elements of each predicted value vector in the set of predicted value vectors, obtaining the maximum activation value of the M non-label elements in each predicted value vector;

[0165] For each predicted value vector in the set of predicted value vectors, replacing the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain an updated predicted value vector corresponding to each predicted value vector;

[0166] Arranging the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0167] Generating a set of recombined predicted value vectors according to the first sequence and the M second sequences.

[0168] In this embodiment, a recombination method of arranging label elements in descending order and arranging non-label elements by replicating the maximum value is introduced. For each piece of training data in each set of training data to be trained, extracting the activation values corresponding to the label elements in each piece of training data to be trained, thereby obtaining a column of activation values, and then arranging this column of activation values in descending order to obtain a first sequence. For the M non-label elements of each predicted value vector, extracting the maximum value of the activation values among the M non-label elements, that is, obtaining the maximum activation value, and then replacing the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain an updated predicted value vector corresponding to each predicted value vector, where the updated predicted value vector does not include the activation values corresponding to the label elements. Based on this, then arranging each column of activation values in the updated predicted value vector in ascending order to obtain M second sequences. Finally, combining the first sequence and the M second sequences, and combining the activation values under the same index into a recombined predicted value vector until a set of recombined predicted value vectors is generated.

[0169] It should be noted that each predicted value vector has an index, and each recombined predicted value vector after rearrangement also has an index, and there is a corresponding relationship between the index of the predicted value vector and the index of the recombined predicted value vector.

[0170] Specifically, for the sake of easy understanding, take the three-classification scenario as an example (i.e., M = 2), and 3 training data to be trained are input in one iteration of training. Please refer to Figure 8 , Figure 8 is another schematic diagram for constructing a set of recombined predicted value vectors in an embodiment of the present application. As shown in the figure, assume that the true class label is "Classification A", that is, the element at the first position in the predicted value vector is the label element. Among them, the predicted value vector a is (0.2, -1, 0.1), the predicted value vector b is (0.1, -5, -3), and the predicted value vector c is (5, 2, 0). Based on the M non-label elements in the predicted value vector a, select the maximum activation value corresponding to the M non-label elements, that is, the maximum activation value in the predicted value vector a is 0.1. Based on the M non-label elements in the predicted value vector b, select the maximum activation value corresponding to the M non-label elements, that is, the maximum activation value in the predicted value vector b is -3. Based on the M non-label elements in the predicted value vector c, select the maximum activation value corresponding to the M non-label elements, that is, the maximum activation value in the predicted value vector c is 2.

[0171] Therefore, for each predicted value vector, replace the activation values corresponding to the M non-label elements with the maximum activation value to obtain the updated predicted value vector corresponding to each predicted value vector. For example, replace the activation value under each non-label element in the predicted value vector a with the maximum activation value 0.1, that is, the updated predicted value vector a is (0.1, 0.1). Replace the activation value under each non-label element in the predicted value vector b with the maximum activation value -3, that is, the updated predicted value vector a is (-3, -3). Replace the activation value under each non-label element in the predicted value vector c with the maximum activation value 2, that is, the updated predicted value vector c is (2, 2).

[0172] Based on this, rearrange the label elements in the predicted value vector a, the label elements in the predicted value vector b, and the label elements in the predicted value vector c, that is, rearrange the activation values z1 in the three predicted value vectors in descending order, and thus obtain the first sequence (5, 0.2, 0.1). Rearrange the non-label elements in the updated predicted value vector a, the non-label elements in the updated predicted value vector b, and the non-label elements in the updated predicted value vector c, that is, rearrange the activation values z2 and activation values z3 in the three updated predicted value vectors in ascending order, and thus obtain two second sequences, namely the second sequence (-3, 0.1, 2) and the second sequence (-3, 0.1, 2).

[0173] Based on the corresponding positions of the activation values z1, z2, and z3, the first sequence and two second sequences are combined. Among them, the index value of the prediction value vector a is "a", and the index value of the corresponding recombined prediction value vector a is also "a", and the recombined prediction value vector a is (5, -3, -3). The index value of the prediction value vector b is "b", and the index value of the corresponding recombined prediction value vector b is also "b", and the recombined prediction value vector b is (0.2, 0.1, 0.1). The index value of the prediction value vector c is "c", and the index value of the corresponding recombined prediction value vector c is also "c", and the recombined prediction value vector c is (0.1, 2, 2).

[0174] Taking the target prediction value vector as the prediction value vector c as an example, its corresponding first value is |5 - 2|, which is equal to 3. There is at least one target recombined prediction value vector among the recombined prediction value vector a, the recombined prediction value vector b, and the recombined prediction value vector c. Taking the target recombined prediction value vector as the recombined prediction value vector b as an example, its corresponding second value is |0.2 - 0.1|, which is equal to 0.1, satisfying the condition that the second value is less than the first value. Therefore, the prediction value vector c can be replaced with the recombined prediction value vector b for model training.

[0175] Secondly, in the embodiments of the present application, a recombination method is provided for arranging label elements in descending order and arranging non-label elements according to the copied maximum value. Through the above method, a recombined prediction value vector with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further enhancing the recognition ability of the classification network model.

[0176] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding embodiments, sorting the label elements and M non-label elements of each prediction value vector in the prediction value vector set to obtain a set of recombined prediction value vectors may include:

[0177] For the label elements of each prediction value vector in the prediction value vector set, arranging a column of activation values corresponding to the label elements in the prediction value vector set in descending order to obtain a first sequence;

[0178] For the M non-label elements of each prediction value vector in the prediction value vector set, obtaining K maximum activation values of the M non-label elements in each prediction value vector, where the K maximum activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M;

[0179] For each predicted value vector in the set of predicted value vectors, replace the K smallest activation values among the M non-label elements with the K largest activation values to obtain the updated predicted value vector corresponding to each predicted value vector, where the K smallest activation values represent the first K activation values after arranging the M non-label elements in ascending order;

[0180] Arrange the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0181] Generate a set of recombined predicted value vectors based on the first sequence and the M second sequences.

[0182] In this embodiment, a recombination method of arranging label elements in descending order and arranging non-label elements in a semi-duplication manner is introduced. For each training data in each set of training data to be trained, extract the activation value corresponding to the label element in each training data to obtain a column of activation values, and then arrange this column of activation values in descending order to obtain the first sequence. For the M non-label elements of each predicted value vector, extract the first K larger activation values among the M non-label elements, that is, obtain the K largest activation values, and then replace the K smallest activation values corresponding to the M non-label elements with the K largest activation values to obtain the updated predicted value vector corresponding to each predicted value vector, where the updated predicted value vector does not include the activation value corresponding to the label element. Based on this, arrange each column of activation values in the updated predicted value vector in ascending order to obtain M second sequences. Finally, combine the first sequence and the M second sequences, and combine the activation values under the same index into a recombined predicted value vector until a set of recombined predicted value vectors is generated. It can be understood that the value of K is an integer greater than 1 and less than M, and usually K can take the value of K = 2 / M, or K = (M + 1) / 2, or K = (M - 1) / 2, but it is necessary to ensure that K is an integer greater than 1.

[0183] It should be noted that each predicted value vector has an index, and each recombined predicted value vector after rearrangement also has an index, and there is a corresponding relationship between the index of the predicted value vector and the index of the recombined predicted value vector.

[0184] Specifically, for the sake of easy understanding, take the five-classification scenario as an example (that is, take M equal to 4), and 3 training data to be trained are input in one iterative training. Please refer to Figure 9 , Figure 9Another schematic diagram for constructing a set of recombinant predicted value vectors in the embodiments of this application is shown in the figure. Assume that the true class label is "Classification A", that is, the element in the first position of the predicted value vector is the label element. Among them, the predicted value vector a is (0.2, -1, 0.1, 1, 0.5), the predicted value vector b is (0.1, -5, -3, 2, 1), and the predicted value vector c is (5, 2, 0, 0.8, 0). Based on the M non-label elements in the predicted value vector a, assuming K is 2, select the K largest activation values corresponding to the M non-label elements, that is, the K largest activation values in the predicted value vector a are 1 and 0.5. Based on the M non-label elements in the predicted value vector b, select the K largest activation values corresponding to the M non-label elements, that is, the K largest activation values in the predicted value vector b are 2 and 1. Based on the M non-label elements in the predicted value vector c, select the K largest activation values corresponding to the M non-label elements, that is, the K largest activation values in the predicted value vector c are 2 and 0.8.

[0185] Therefore, for each predicted value vector, replace the K smallest activation values corresponding to the M non-label elements with the largest activation values to obtain the updated predicted value vector corresponding to each predicted value vector. For example, replace the K smallest activation values under each non-label element in the predicted value vector a with the K largest activation values, that is, the updated predicted value vector a is (1, 0.5, 1, 0.5). Replace the K smallest activation values under each non-label element in the predicted value vector b with the K largest activation values, that is, the updated predicted value vector b is (2, 1, 2, 1). Replace the K smallest activation values under each non-label element in the predicted value vector c with the K largest activation values, that is, the updated predicted value vector c is (2, 0.8, 2, 0.8).

[0186] Based on this, rearrange the label elements in the predicted value vector a, the label elements in the predicted value vector b, and the label elements in the predicted value vector c, that is, rearrange the activation values z1 in the three predicted value vectors in descending order, and thus obtain the first sequence (5, 0.2, 0.1). Rearrange the non-label elements in the updated predicted value vector a, the non-label elements in the updated predicted value vector b, and the non-label elements in the updated predicted value vector c, that is, rearrange the activation values z2, activation value z3, activation value z4, and activation value z5 in the three updated predicted value vectors in ascending order, and thus obtain four second sequences, namely the second sequence (1, 2, 2), the second sequence (0.5, 0.8, 1), the second sequence (1, 2, 2), and the second sequence (0.5, 0.8, 1).

[0187] Based on the corresponding positions of activation values z1, z2, z3, z4, and z5, combine the first sequence and four second sequences. Among them, the index value of the prediction value vector a is "a", and the index value of the corresponding recombined prediction value vector a is also "a", and the recombined prediction value vector a is (5, 1, 0.5, 1, 0.5). The index value of the prediction value vector b is "b", and the index value of the corresponding recombined prediction value vector b is also "b", and the recombined prediction value vector b is (0.2, 2, 0.8, 2, 0.8). The index value of the prediction value vector c is "c", and the index value of the corresponding recombined prediction value vector c is also "c", and the recombined prediction value vector c is (0.1, 2, 1, 2, 1).

[0188] Taking the target prediction value vector as the prediction value vector c as an example, its corresponding first value is |5 - 2|, which is equal to 3. There is at least one target recombined prediction value vector among the recombined prediction value vector a, the recombined prediction value vector b, and the recombined prediction value vector c. Taking the target recombined prediction value vector as the recombined prediction value vector b as an example, its corresponding second value is |0.2 - 2|, which is equal to 1.8, satisfying the condition that the second value is less than the first value. Therefore, the prediction value vector c can be replaced with the recombined prediction value vector b for model training.

[0189] Secondly, in the embodiments of the present application, a recombination method is provided for arranging label elements in descending order and arranging non-label elements in a semi-duplication manner. Through the above method, a recombined prediction value vector with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further enhancing the recognition ability of the classification network model.

[0190] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding embodiment, training the classification network model according to the set of recombined prediction value vectors may include:

[0191] Obtain N recombined prediction value vectors from the set of recombined prediction value vectors, where the N recombined prediction value vectors are the last 1 to N recombined prediction value vectors in the set of recombined prediction value vectors, and N is an integer greater than or equal to 1;

[0192] Update the network parameters of the classification network model using the loss function according to the N recombined prediction value vectors.

[0193] In this embodiment, a training method based on arranging tag elements in descending order and non-tag elements in ascending order is introduced. After obtaining the set of recombined predicted value vectors, considering that in some recombined predicted value vectors, the activation value gaps may still be relatively large, therefore, a part of the recombined predicted value vectors can be selected for training, and the part with relatively large activation value gaps can be discarded. In the mode of arranging tag elements in descending order and non-tag elements in ascending order, for the middle-segment recombined predicted value vectors in the set of recombined predicted value vectors, the activation value gaps corresponding to each element will be relatively small. At the same time, considering that the activation values under tag elements are usually greater than or equal to 0, therefore, for the latter-segment recombined predicted value vectors in the set of recombined predicted value vectors, the activation value gaps corresponding to each element are smaller than those corresponding to each element in the former-segment recombined predicted value vectors. Therefore, N recombined predicted value vectors can be selected from the set of recombined predicted value vectors in the order from back to front, and then the loss function can be constructed using these N recombined predicted value vectors to train the classification network model.

[0194] Specifically, for the convenience of introduction, please refer to Figure 10 , Figure 10 FIG. is a schematic diagram for selecting N recombined predicted value vectors for model training in the embodiment of the present application. As shown in the figure, it is assumed that the set of recombined predicted value vectors includes recombined predicted value vector a, recombined predicted value vector b, recombined predicted value vector c, recombined predicted value vector d, recombined predicted value vector e, and recombined predicted value vector f. It is assumed that the true class label is "Classification A", that is, the element in the first position of the predicted value vector is the tag element. Based on this, please refer to Table 2, which is a schematic diagram of the gap between the recombined predicted value vector and the activation value.

[0195] Table 2

[0196] Vector identifier Recombinant predicted value vector Activation value gap Recombinant predicted value vector a (5,-5,-5) 10 Recombinant predicted value vector b (4,-3,-2) 6 Recombinant predicted value vector c (1,-2,0) 1 Recombinant predicted value vector d (0.8,-1,1) 0.2 Recombinant predicted value vector e (0.2,-1,3) 2.8 Recombinant predicted value vector f (0.1,0,4) 3.9

[0197] As can be seen from Table 2, the gap between the activation values in the recombined predicted value vector d is the smallest. Therefore, the recombined predicted value vectors after the recombined predicted value vector d can be directly used as the recombined predicted value vectors for constructing the loss function, and thus N recombined predicted value vectors are obtained. For example, the recombined predicted value vector d, the recombined predicted value vector e, and the recombined predicted value vector f are used as 3 recombined predicted value vectors for constructing the loss function, and the classification network model is trained with this.

[0198] It should be noted that in practical applications, the following method can also be used to select N recombined predicted value vectors.

[0199] Method 1: Only select the middle-segment recombined predicted value vectors;

[0200] Taking Figure 10For example, the value of N can be preset. Assuming N is 2, the middle two recombinant prediction value vectors are preferentially selected for training the classification network model, that is, the recombinant prediction value vector c and the recombinant prediction value vector d are selected.

[0201] Method 2: Select the recombinant prediction value vector according to a threshold;

[0202] Taking Figure 10 as an example, an activation value difference threshold can be preset. Assuming the activation value difference threshold is 3, the recombinant prediction value vectors less than or equal to the activation value difference threshold are selected, that is, the recombinant prediction value vector c, the recombinant prediction value vector d, and the recombinant prediction value vector e are selected.

[0203] Again, in the embodiment of the present application, a training method based on descendingly arranging label elements and ascendingly arranging non-label elements is provided. Through the above method, when the label elements are arranged in descending order and the non-label elements are arranged in ascending order, the recombinant prediction value vectors in the second half of the recombinant prediction value vector set can be directly selected, which is more in line with the characteristic that the activation values are close, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, and enabling the classification network model to learn more efficiently. Furthermore, the recognition ability of the classification network model is improved.

[0204] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiment of the present application, sorting the label elements and M non-label elements of each prediction value vector in the prediction value vector set to obtain a recombinant prediction value vector set may include:

[0205] For the label elements of each prediction value vector in the prediction value vector set, arrange them in ascending order to obtain a first sequence;

[0206] For the M non-label elements of each prediction value vector in the prediction value vector set, arrange them in descending order to obtain M second sequences;

[0207] Generate a recombinant prediction value vector set according to the first sequence and the M second sequences.

[0208] In this embodiment, a recombination method is introduced in which label elements are arranged in ascending order and non-label elements are arranged in descending order. For each training data in each training data set, the activation values corresponding to the label elements in each training data are extracted, thereby obtaining a column of activation values. Then, the column of activation values is arranged in ascending order to obtain a first sequence. Similarly, the activation values corresponding to the other M non-label elements are extracted, thereby obtaining M columns of activation values. Then, each column of the M columns of activation values is arranged in descending order to obtain M second sequences. Finally, the first sequence and the M second sequences are combined, and the activation values at the same index are combined into a recombined prediction value vector until a set of recombined prediction value vectors is generated.

[0209] It should be noted that each prediction value vector has an index, and each recombined prediction value vector after rearrangement also has an index. The index of the prediction value vector and the index of the recombined prediction value vector have a corresponding relationship.

[0210] Specifically, for the sake of easy understanding, a three-classification scenario is taken as an example (i.e., M equals 2), and 3 training data are input in one iteration of training. Please refer to Figure 11 , Figure 11 is another schematic diagram for constructing a set of recombined prediction value vectors in the embodiment of the present application. As shown in the figure, it is assumed that the true class label is "Classification A", that is, the element in the first position of the prediction value vector is a label element. Among them, the prediction value vector a is (0.2, -1, -1), the prediction value vector b is (0.1, -5, -5), and the prediction value vector c is (5, 0, 0). The label elements in the prediction value vector a, the label elements in the prediction value vector b, and the label elements in the prediction value vector c are rearranged, that is, the activation values z1 in the three prediction value vectors are rearranged in ascending order, thereby obtaining a first sequence (0.1, 0.2, 5). The non-label elements in the prediction value vector a, the non-label elements in the prediction value vector b, and the non-label elements in the prediction value vector c are rearranged, that is, the activation values z2 and z3 in the three prediction value vectors are rearranged in descending order, thereby obtaining two second sequences, namely the second sequence (0, -1, -5) and the second sequence (0, -1, -5).

[0211] Based on the corresponding positions of the activation values z1, z2, and z3, the first sequence and two second sequences are combined. Among them, the index value of the prediction value vector a is "a", and the index value of the corresponding recombined prediction value vector a is also "a", and the recombined prediction value vector a is (0.1, 0, 0). The index value of the prediction value vector b is "b", and the index value of the corresponding recombined prediction value vector b is also "b", and the recombined prediction value vector b is (0.2, -1, -1). The index value of the prediction value vector c is "c", and the index value of the corresponding recombined prediction value vector c is also "c", and the recombined prediction value vector c is (5, -5, -5).

[0212] Taking the target prediction value vector as the prediction value vector a as an example, the corresponding first value is |0.2 - (-1)|, which is equal to 1.2. There is at least one target recombined prediction value vector among the recombined prediction value vector a, the recombined prediction value vector b, and the recombined prediction value vector c. Taking the target recombined prediction value vector as the recombined prediction value vector a as an example, the corresponding second value is |0.1 - 0|, which is equal to 0.1, satisfying the condition that the second value is less than the first value. Therefore, the prediction value vector a can be replaced with the recombined prediction value vector a for model training.

[0213] Secondly, in the embodiments of the present application, a recombination method is provided in which the label elements are arranged in ascending order and the non-label elements are arranged in descending order. Through the above method, a recombined prediction value vector with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further improving the recognition ability of the classification network model.

[0214] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the above embodiment, sorting the label elements and M non-label elements of each prediction value vector in the prediction value vector set to obtain a recombined prediction value vector set may include:

[0215] Obtain a first sequence according to the label elements of each prediction value vector in the prediction value vector set;

[0216] For the M non-label elements of each prediction value vector in the prediction value vector set, arrange them in descending order to obtain M second sequences;

[0217] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0218] In this embodiment, a recombination method is introduced in which the label elements are arranged in the original order and the non-label elements are arranged in descending order. For each training data in each training data set to be trained, the activation values corresponding to the label elements in each training data to be trained are extracted, and thus a column of activation values is obtained. The column of activation values is arranged in the original order to obtain a first sequence. Similarly, the activation values corresponding to the other M non-label elements are extracted, and thus M columns of activation values are obtained. Then, each column of the M columns of activation values is arranged in descending order to obtain M second sequences. Finally, the first sequence and the M second sequences are combined, and the activation values under the same index are combined into a recombined prediction value vector until a set of recombined prediction value vectors is generated.

[0219] It should be noted that each prediction value vector has an index, and each recombined prediction value vector after rearrangement also has an index. The index of the prediction value vector and the index of the recombined prediction value vector have a corresponding relationship.

[0220] Specifically, for the sake of easy understanding, taking the three-classification scenario as an example (i.e., taking M equal to 2), and 3 training data to be trained are input in one iteration of training, please refer to Figure 12 , Figure 12 is another schematic diagram for constructing a set of recombined prediction value vectors in the embodiment of the present application. As shown in the figure, it is assumed that the true class label is "Classification A", that is, the element in the first position of the prediction value vector is a label element. Among them, the prediction value vector a is (0.2, -1, -1), the prediction value vector b is (0.1, -5, -5), and the prediction value vector c is (5, 0, 0). The label elements in the prediction value vector a, the label elements in the prediction value vector b, and the label elements in the prediction value vector c are arranged, that is, the activation values z1 in the three prediction value vectors are arranged in the original order, and thus a first sequence (0.2, 0.1, 5) is obtained. The non-label elements in the prediction value vector a, the non-label elements in the prediction value vector b, and the non-label elements in the prediction value vector c are rearranged, that is, the activation values z2 and z3 in the three prediction value vectors are rearranged in descending order, and thus two second sequences are obtained, namely the second sequence (0, -1, -5) and the second sequence (0, -1, -5).

[0221] Based on the corresponding positions of activation values z1, z2, and z3, the first sequence and two second sequences are combined. Among them, the index value of the prediction value vector a is "a", the index value of the corresponding recombined prediction value vector a is also "a", and the recombined prediction value vector a is (0.2, 0, 0). The index value of the prediction value vector b is "b", the index value of the corresponding recombined prediction value vector b is also "b", and the recombined prediction value vector b is (0.1, -1, -1). The index value of the prediction value vector c is "c", the index value of the corresponding recombined prediction value vector c is also "c", and the recombined prediction value vector c is (5, -5, -5).

[0222] Taking the target prediction value vector as the prediction value vector a as an example, the corresponding first value is |0.2 - (-1)|, which is equal to 1.2. There is at least one target recombined prediction value vector among the recombined prediction value vector a, the recombined prediction value vector b, and the recombined prediction value vector c. Taking the target recombined prediction value vector as the recombined prediction value vector a as an example, the corresponding second value is |0.2 - 0|, which is equal to 0.2, satisfying the condition that the second value is less than the first value. Therefore, the prediction value vector a can be replaced with the recombined prediction value vector a for model training.

[0223] Secondly, in the embodiments of the present application, a recombination method is provided for arranging label elements in the original order and arranging non-label elements in descending order. Through the above method, recombined prediction value vectors with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further enhancing the recognition ability of the classification network model.

[0224] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding embodiments, sorting the label elements and M non-label elements of each prediction value vector in the prediction value vector set to obtain a recombined prediction value vector set may include:

[0225] For the label elements of each prediction value vector in the prediction value vector set, arranging the column of activation values corresponding to the label elements in the prediction value vector set in ascending order to obtain the first sequence;

[0226] For the M non-label elements of each prediction value vector in the prediction value vector set, obtaining the maximum activation value of the M non-label elements in each prediction value vector;

[0227] For each prediction value vector in the prediction value vector set, replacing the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain the updated prediction value vector corresponding to each prediction value vector;

[0228] Arrange the M column activation values corresponding to the updated prediction value vectors corresponding to each prediction value vector in ascending order to obtain M second sequences;

[0229] Generate a set of recombined prediction value vectors according to the first sequence and the M second sequences.

[0230] In this embodiment, a recombination method of arranging label elements in ascending order and arranging non-label elements by copying the maximum value is introduced. For each training data in each training data set to be trained, extract the activation values corresponding to the label elements in each training data to obtain a column of activation values, and then arrange the column of activation values in descending order to obtain the first sequence. For the M non-label elements of each prediction value vector, extract the maximum activation value among the activation values of the M non-label elements, that is, obtain the maximum activation value, and then replace the activation values corresponding to the M non-label elements with the maximum activation value of the M non-label elements to obtain the updated prediction value vector corresponding to each prediction value vector, where the updated prediction value vector does not include the activation values corresponding to the label elements. Based on this, arrange each column of activation values in the updated prediction value vector in descending order to obtain M second sequences. Finally, combine the first sequence and the M second sequences, and combine the activation values under the same index into a recombined prediction value vector until a set of recombined prediction value vectors is generated.

[0231] It should be noted that each prediction value vector has an index, and each recombined prediction value vector after rearrangement also has an index, and there is a corresponding relationship between the index of the prediction value vector and the index of the recombined prediction value vector.

[0232] Specifically, for the sake of easy understanding, take the three-classification scenario as an example (that is, take M equal to 2), and in one iteration training, 3 training data to be trained are input. Please refer to Figure 13 , Figure 13 is another schematic diagram for constructing a set of recombined prediction value vectors in the embodiment of the present application. As shown in the figure, assume that the true class label is "Classification A", that is, the element in the first position of the prediction value vector is the label element. Among them, the prediction value vector a is (0.2, -1, 0.1), the prediction value vector b is (0.1, -5, -3), and the prediction value vector c is (5, 2, 0). Based on the M non-label elements in the prediction value vector a, select the maximum activation value corresponding to the activation values of the M non-label elements, that is, the maximum activation value in the prediction value vector a is 0.1. Based on the M non-label elements in the prediction value vector b, select the maximum activation value corresponding to the activation values of the M non-label elements, that is, the maximum activation value in the prediction value vector b is -3. Based on the M non-label elements in the prediction value vector c, select the maximum activation value corresponding to the activation values of the M non-label elements, that is, the maximum activation value in the prediction value vector c is 2.

[0233] Then, for each predicted value vector, the activation values corresponding to the M non-label elements are replaced with the maximum activation value to obtain the updated predicted value vector corresponding to each predicted value vector. For example, by replacing the activation value under each non-label element in the predicted value vector a with the maximum activation value 0.1, the updated predicted value vector a is obtained as (0.1, 0.1). By replacing the activation value under each non-label element in the predicted value vector b with the maximum activation value -3, the updated predicted value vector a is obtained as (-3, -3). By replacing the activation value under each non-label element in the predicted value vector c with the maximum activation value 2, the updated predicted value vector c is obtained as (2, 2).

[0234] Based on this, the label elements in the predicted value vector a, the label elements in the predicted value vector b, and the label elements in the predicted value vector c are rearranged, that is, the activation values z1 in the three predicted value vectors are rearranged in ascending order, thereby obtaining the first sequence (0.1, 0.2, 5). The non-label elements in the updated predicted value vector a, the non-label elements in the updated predicted value vector b, and the non-label elements in the updated predicted value vector c are rearranged, that is, the activation values z2 and z3 in the three updated predicted value vectors are rearranged in descending order, thereby obtaining two second sequences, namely the second sequence (2, 0.1, -3) and the second sequence (2, 0.1, -3).

[0235] Based on the corresponding positions of the activation value z1, the activation value z2, and the activation value z3, the first sequence and the two second sequences are combined. Among them, the index value of the predicted value vector a is "a", and the index value of the corresponding recombined predicted value vector a is also "a", and the recombined predicted value vector a is (0.1, 2, 2). The index value of the predicted value vector b is "b", and the index value of the corresponding recombined predicted value vector b is also "b", and the recombined predicted value vector b is (0.2, 0.1, 0.1). The index value of the predicted value vector c is "c", and the index value of the corresponding recombined predicted value vector c is also "c", and the recombined predicted value vector c is (5, -3, -3).

[0236] Taking the target predicted value vector as the predicted value vector c as an example, the corresponding first value is |5 - 2|, which is equal to 3. There is at least one target recombined predicted value vector among the recombined predicted value vector a, the recombined predicted value vector b, and the recombined predicted value vector c. Taking the target recombined predicted value vector as the recombined predicted value vector b as an example, the corresponding second value is |0.2 - 0.1|, which is equal to 0.1, satisfying the condition that the second value is less than the first value. Therefore, the predicted value vector c can be replaced with the recombined predicted value vector b for model training.

[0237] Secondly, in the embodiments of the present application, a recombination method is provided, which arranges label elements in ascending order and arranges non-label elements according to the maximum copied value. Through the above method, a recombined predicted value vector with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further enhancing the recognition ability of the classification network model.

[0238] Optionally, based on the corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present application, sorting the label elements and M non-label elements of each predicted value vector in the predicted value vector set to obtain a recombined predicted value vector set may include: Figure 3 For the label elements of each predicted value vector in the predicted value vector set, arranging a column of activation values corresponding to the label elements in the predicted value vector set in ascending order to obtain a first sequence;

[0239] For the M non-label elements of each predicted value vector in the predicted value vector set, obtaining K maximum activation values of the M non-label elements in each predicted value vector, where the K maximum activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M;

[0240] For each predicted value vector in the predicted value vector set, replacing the K minimum activation values among the M non-label elements with the K maximum activation values to obtain an updated predicted value vector corresponding to each predicted value vector, where the K minimum activation values represent the first K activation values after arranging the M non-label elements in ascending order;

[0241] Arranging the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0242] Generating a recombined predicted value vector set according to the first sequence and the M second sequences.

[0243] According to the first sequence and the M second sequences, generate a recombined predicted value vector set.

[0244] In this embodiment, a recombination method is introduced in which the label elements are arranged in ascending order and the non-label elements are arranged in a semi-duplication manner. For each training data in each training data set, the activation values corresponding to the label elements in each training data are extracted, thus obtaining a column of activation values. Then, the column of activation values is arranged in ascending order to obtain the first sequence. For the M non-label elements of each prediction value vector, the top K larger activation values among the M non-label elements are extracted, that is, K maximum activation values are obtained. Then, the K minimum activation values corresponding to the M non-label elements are replaced with the K maximum activation values to obtain the updated prediction value vector corresponding to each prediction value vector, where the updated prediction value vector does not include the activation values corresponding to the label elements. Based on this, each column of activation values in the updated prediction value vector is arranged in descending order to obtain M second sequences. Finally, the first sequence and the M second sequences are combined, and the activation values under the same index are combined into a recombined prediction value vector until a set of recombined prediction value vectors is generated. It can be understood that the value of K is an integer greater than 1 and less than M, and usually K can take the value of K = 2 / M, or K = (M + 1) / 2, or K = (M - 1) / 2, but it is necessary to ensure that K is an integer greater than 1.

[0245] It should be noted that each prediction value vector has an index, and each recombined prediction value vector after rearrangement also has an index. The index of the prediction value vector and the index of the recombined prediction value vector have a corresponding relationship.

[0246] Specifically, for the sake of easy understanding, taking the five-classification scenario as an example (that is, taking M equal to 4), and in one iteration training, 3 training data are input. Please refer to Figure 14 , Figure 14 is another schematic diagram for constructing a set of recombined prediction value vectors in the embodiment of the present application. As shown in the figure, it is assumed that the true class label is "Classification A", that is, the element in the first position of the prediction value vector is the label element. Among them, the prediction value vector a is (0.2, -1, 0.1, 1, 0.5), the prediction value vector b is (0.1, -5, -3, 2, 1), and the prediction value vector c is (5, 2, 0, 0.8, 0). Based on the M non-label elements in the prediction value vector a, assuming K is 2, the K maximum activation values corresponding to the M non-label elements are selected, that is, the K maximum activation values in the prediction value vector a are 1 and 0.5. Based on the M non-label elements in the prediction value vector b, the K maximum activation values corresponding to the M non-label elements are selected, that is, the K maximum activation values in the prediction value vector b are 2 and 1. Based on the M non-label elements in the prediction value vector c, the K maximum activation values corresponding to the M non-label elements are selected, that is, the K maximum activation values in the prediction value vector c are 2 and 0.8.

[0247] Then, for each predicted value vector, the K minimum activation values corresponding to the M non-label elements are replaced with the maximum activation values to obtain the updated predicted value vector corresponding to each predicted value vector. For example, by replacing the K minimum activation values under each non-label element in the predicted value vector a with the K maximum activation values, the updated predicted value vector a is obtained as (1, 0.5, 1, 0.5). By replacing the K minimum activation values under each non-label element in the predicted value vector b with the K maximum activation values, the updated predicted value vector b is obtained as (2, 1, 2, 1). By replacing the K minimum activation values under each non-label element in the predicted value vector c with the K maximum activation values, the updated predicted value vector c is obtained as (2, 0.8, 2, 0.8).

[0248] Based on this, the label elements in the predicted value vector a, the label elements in the predicted value vector b, and the label elements in the predicted value vector c are rearranged, that is, the activation values z1 in the three predicted value vectors are rearranged in ascending order, and thus the first sequence (0.1, 0.2, 5) is obtained. The non-label elements in the updated predicted value vector a, the non-label elements in the updated predicted value vector b, and the non-label elements in the updated predicted value vector c are rearranged, that is, the activation values z2, z3, z4, and z5 in the three updated predicted value vectors are rearranged in descending order, and thus four second sequences are obtained, namely the second sequence (2, 2, 1), the second sequence (1, 0.8, 0.5), the second sequence (2, 2, 1), and the second sequence (1, 0.8, 0.5).

[0249] Based on the corresponding positions of the activation values z1, z2, z3, z4, and z5, the first sequence and the four second sequences are combined. Among them, the index value of the predicted value vector a is "a", and the index value of the corresponding recombined predicted value vector a is also "a", and the recombined predicted value vector a is (0.1, 2, 1, 2, 1). The index value of the predicted value vector b is "b", and the index value of the corresponding recombined predicted value vector b is also "b", and the recombined predicted value vector b is (0.2, 2, 0.8, 2, 0.8). The index value of the predicted value vector c is "c", and the index value of the corresponding recombined predicted value vector c is also "c", and the recombined predicted value vector c is (5, 1, 0.5, 1, 0.5).

[0250] Taking the target prediction value vector as the prediction value vector c as an example, the corresponding first value is |5 - 2|, which is equal to 3. Among the recombined prediction value vector a, the recombined prediction value vector b, and the recombined prediction value vector c, there is at least one target recombined prediction value vector. Taking the target recombined prediction value vector as the recombined prediction value vector b as an example, the corresponding second value is |0.2 - 2|, which is equal to 1.8, satisfying the condition that the second value is less than the first value. Therefore, the prediction value vector c can be replaced with the recombined prediction value vector b for model training.

[0251] Secondly, in the embodiments of the present application, a recombination method for arranging label elements in ascending order and non - label elements in semi - replication arrangement is provided. Through the above method, recombined prediction value vectors with more similar activation values can be obtained, thereby increasing the gradient of backpropagation, improving the training effect of the classification network model, enabling the classification network model to learn more efficiently, and further enhancing the recognition ability of the classification network model.

[0252] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, training the classification network model according to the set of recombined prediction value vectors includes:

[0253] Obtain N recombined prediction value vectors from the set of recombined prediction value vectors, where the N recombined prediction value vectors are the 1st to Nth recombined prediction value vectors in the set of recombined prediction value vectors, and N is an integer greater than or equal to 1;

[0254] Update the network parameters of the classification network model using a loss function according to the N recombined prediction value vectors.

[0255] In this embodiment, a training method based on arranging label elements in ascending order and non - label elements in descending order is introduced. After obtaining the set of recombined prediction value vectors, considering that in some recombined prediction value vectors, the gap between activation values may still be large, some recombined prediction value vectors can be selected for training, and the part with a large gap between activation values can be discarded. In the mode of arranging label elements in ascending order and non - label elements in descending order, for the middle - segment recombined prediction value vectors in the set of recombined prediction value vectors, the gap between the activation values corresponding to each element is relatively small. At the same time, considering that the activation values under label elements are usually greater than or equal to 0, for the first - half - segment recombined prediction value vectors in the set of recombined prediction value vectors, the gap between the activation values corresponding to each element is smaller than that of the second - half - segment recombined prediction value vectors. Therefore, N recombined prediction value vectors can be selected from the set of recombined prediction value vectors in the order from front to back, and then a loss function is constructed using these N recombined prediction value vectors to train the classification network model.

[0256] Specifically, for the convenience of introduction, please refer to Figure 15 , Figure 15 which is another schematic diagram for selecting N recombinant predicted value vectors for model training in the embodiments of the present application. As shown in the figure, it is assumed that the set of recombinant predicted value vectors includes recombinant predicted value vector a, recombinant predicted value vector b, recombinant predicted value vector c, recombinant predicted value vector d, recombinant predicted value vector e, and recombinant predicted value vector f. It is assumed that the true class label is "Classification A", that is, the element in the first position of the predicted value vector is the label element. Based on this, please refer to Table 3, which is a schematic diagram of the gap between the recombinant predicted value vector and the activation value.

[0257] Table 3

[0258] Vector identifier Recombinant predicted value vector Activation value gap Recombinant predicted value vector a (0.1,0,4) 3.9 Recombinant predicted value vector b (0.2,-1,3) 2.8 Recombinant predicted value vector c (0.8,-1,1) 0.2 Recombinant predicted value vector d (1,-2,0) 1 Recombinant predicted value vector e (4,-3,-2) 6 Recombinant predicted value vector f (5,-5,-5) 10

[0259] As can be seen from Table 3, the gap between the activation values in the recombinant predicted value vector c is the smallest. Therefore, the recombinant predicted value vectors before the recombinant predicted value vector c can be directly used as the recombinant predicted value vectors for constructing the loss function, and thus N recombinant predicted value vectors are obtained. For example, the recombinant predicted value vector a, the recombinant predicted value vector b, and the recombinant predicted value vector c are used as 3 recombinant predicted value vectors for constructing the loss function, and the classification network model is trained accordingly.

[0260] It should be noted that in practical applications, the following methods can also be used to select N recombinant predicted value vectors.

[0261] Method 1: Only select the recombinant predicted value vectors in the middle section;

[0262] Taking Figure 15 as an example, the value of N can be preset. Assuming N is 2, the two middle recombinant predicted value vectors are preferentially selected for training the classification network model, that is, the recombinant predicted value vector b and the recombinant predicted value vector c are selected.

[0263] Method 2: Select the recombinant predicted value vectors according to a threshold;

[0264] Taking Figure 15 as an example, an activation value gap threshold can be preset. Assuming the activation value gap threshold is 3, then the recombinant predicted value vectors less than or equal to this activation value gap threshold are selected, that is, the recombinant predicted value vector b, the recombinant predicted value vector c, and the recombinant predicted value vector d are selected.

[0265] Again, in the embodiments of the present application, a training method is provided that arranges label elements in ascending order and non-label elements in descending order. By this method, when the label elements are arranged in ascending order and the non-label elements are arranged in descending order, the first half of the recombined prediction value vectors in the recombined prediction value vector set can be directly selected, which is more in line with the characteristic that the activation values are close. Thereby, the gradient of backpropagation is increased, thus improving the training effect of the classification network model and enabling the classification network model to learn more efficiently. Furthermore, the recognition ability of the classification network model is improved.

[0266] The model training device in the present application will be described in detail below. Please refer to Figure 16 , Figure 16 which is a schematic diagram of an embodiment of the model training device in the embodiments of the present application. The model training device 20 includes:

[0267] An acquisition module 201, configured to acquire a set of data to be trained, where the set of data to be trained includes at least two data to be trained, and each data to be trained has the same true class label;

[0268] The acquisition module 201 is further configured to, based on the set of data to be trained, obtain a set of prediction value vectors through a classification network model, where the set of prediction value vectors includes at least two prediction value vectors, each prediction value vector corresponds to a data to be trained, each prediction value vector includes label elements and M non-label elements, the set of prediction value vectors includes a target prediction value vector, and the absolute value of the difference between the activation value of the label elements in the target prediction value vector and the maximum activation value among the M non-label elements is a first value, and M is an integer greater than or equal to 1;

[0269] A sorting module 202, configured to sort the label elements and the M non-label elements of each prediction value vector in the set of prediction value vectors to obtain a set of recombined prediction value vectors, where the set of recombined prediction value vectors includes a target recombined prediction value vector, and the absolute value of the difference between the activation value of the label elements in the target recombined prediction value vector and the maximum activation value among the M non-label elements is a second value, and the second value is less than the first value;

[0270] A training module 203, configured to train the classification network model according to the set of recombined prediction value vectors.

[0271] Optionally, on the basis of the above Figure 16 corresponding embodiment, in another embodiment of the model training device 20 provided in the embodiments of the present application,

[0272] The acquisition module 201 is specifically configured to acquire an initial set of training data, where the initial set of training data includes Q data to be trained, each of the Q data to be trained has a true class label, and Q is an integer greater than or equal to 2;

[0273] Obtain at least two data to be trained with the same true class label from the initial training data set as the data set to be trained.

[0274] Optionally, based on the corresponding embodiment above, Figure 16 In another embodiment of the model training device 20 provided by the embodiments of the present application,

[0275] The sorting module 202 is specifically configured to, for each label element of the prediction value vector set, arrange a column of activation values corresponding to the label elements in the prediction value vector set in descending order to obtain a first sequence;

[0276] For each of the M non-label elements of each prediction value vector in the prediction value vector set, arrange the M columns of activation values corresponding to the M non-label elements in the prediction value vector set in ascending order to obtain M second sequences;

[0277] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0278] Optionally, based on the corresponding embodiment above, Figure 16 In another embodiment of the model training device 20 provided by the embodiments of the present application,

[0279] The sorting module 202 is specifically configured to obtain a first sequence according to the label elements of each prediction value vector in the prediction value vector set;

[0280] For each of the M non-label elements of each prediction value vector in the prediction value vector set, arrange the M columns of activation values corresponding to the M non-label elements in the prediction value vector set in ascending order to obtain M second sequences;

[0281] Generate a recombined prediction value vector set according to the first sequence and the M second sequences.

[0282] Optionally, based on the corresponding embodiment above, Figure 16 In another embodiment of the model training device 20 provided by the embodiments of the present application,

[0283] The sorting module 202 is specifically configured to, for each label element of the prediction value vector set, arrange a column of activation values corresponding to the label elements in the prediction value vector set in descending order to obtain a first sequence;

[0284] For each of the M non-label elements of each prediction value vector in the prediction value vector set, obtain the maximum activation value of the M non-label elements in each prediction value vector;

[0285] For each predicted value vector in the set of predicted value vectors, replace the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain the updated predicted value vector corresponding to each predicted value vector;

[0286] Arrange the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0287] Generate a set of recombined predicted value vectors according to the first sequence and the M second sequences.

[0288] Optionally, based on the above Figure 16 In another embodiment of the model training device 20 provided by the embodiments of the present application,

[0289] The sorting module 202 is specifically configured to arrange the activation values corresponding to the label elements in the set of predicted value vectors in descending order for each predicted value vector in the set of predicted value vectors to obtain the first sequence;

[0290] For the M non-label elements of each predicted value vector in the set of predicted value vectors, obtain the K maximum activation values of the M non-label elements in each predicted value vector, where the K maximum activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M;

[0291] For each predicted value vector in the set of predicted value vectors, replace the K minimum activation values among the M non-label elements with the K maximum activation values to obtain the updated predicted value vector corresponding to each predicted value vector, where the K minimum activation values represent the first K activation values after arranging the M non-label elements in ascending order;

[0292] Arrange the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences;

[0293] Generate a set of recombined predicted value vectors according to the first sequence and the M second sequences.

[0294] Optionally, based on the above Figure 16 In another embodiment of the model training device 20 provided by the embodiments of the present application,

[0295] The training module 203 is specifically configured to obtain N recombined predicted value vectors from the set of recombined predicted value vectors, where the N recombined predicted value vectors are the last 1 to N recombined predicted value vectors in the set of recombined predicted value vectors, and N is an integer greater than or equal to 1;

[0296] Based on N recombined predicted value vectors, the network parameters of the classification network model are updated using a loss function.

[0297] Optionally, based on the corresponding embodiment above, in another embodiment of the model training device 20 provided by the embodiments of the present application, Figure 16 the sorting module 202 is specifically configured to sort the label elements of each predicted value vector in the predicted value vector set in ascending order to obtain a first sequence;

[0298] For the M non-label elements of each predicted value vector in the predicted value vector set, sort them in descending order to obtain M second sequences;

[0299] Generate a recombined predicted value vector set according to the first sequence and the M second sequences.

[0300] Based on the corresponding embodiment above, in another embodiment of the model training device 20 provided by the embodiments of the present application,

[0301] Optionally, Figure 16 the sorting module 202 is specifically configured to obtain a first sequence according to the label elements of each predicted value vector in the predicted value vector set;

[0302] For the M non-label elements of each predicted value vector in the predicted value vector set, sort them in descending order to obtain M second sequences;

[0303] Generate a recombined predicted value vector set according to the first sequence and the M second sequences.

[0304] Based on the corresponding embodiment above, in another embodiment of the model training device 20 provided by the embodiments of the present application,

[0305] Optionally, Figure 16 the sorting module 202 is specifically configured to sort a column of activation values corresponding to the label elements in the predicted value vector set in ascending order for the label elements of each predicted value vector in the predicted value vector set to obtain a first sequence;

[0306] For the M non-label elements of each predicted value vector in the predicted value vector set, obtain the maximum activation value of the M non-label elements in each predicted value vector;

[0307] For each predicted value vector in the predicted value vector set, replace the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain an updated predicted value vector corresponding to each predicted value vector;

[0308]

[0309] ​Arrange the M-column activation values corresponding to the updated prediction value vectors corresponding to each prediction value vector in ascending order to obtain M second sequences;

[0310] Generate a set of recombined prediction value vectors according to the first sequence and the M second sequences.

[0311] Optionally, based on the above Figure 16 corresponding embodiment, in another embodiment of the model training device 20 provided by the embodiments of the present application,

[0312] The sorting module 202 is specifically configured to, for each label element of the prediction value vectors in the prediction value vector set, arrange a column of activation values corresponding to the label elements in the prediction value vector set in ascending order to obtain a first sequence;

[0313] For the M non-label elements of each prediction value vector in the prediction value vector set, obtain the K largest activation values of the M non-label elements in each prediction value vector, where the K largest activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M;

[0314] For each prediction value vector in the prediction value vector set, replace the K smallest activation values among the M non-label elements with the K largest activation values to obtain the updated prediction value vector corresponding to each prediction value vector, where the K smallest activation values represent the first K activation values after arranging the M non-label elements in ascending order;

[0315] Arrange the M-column activation values corresponding to the updated prediction value vectors corresponding to each prediction value vector in ascending order to obtain M second sequences;

[0316] Generate a set of recombined prediction value vectors according to the first sequence and the M second sequences.

[0317] Optionally, based on the above Figure 16 corresponding embodiment, in another embodiment of the model training device 20 provided by the embodiments of the present application,

[0318] The training module 203 is specifically configured to obtain N recombined prediction value vectors from the set of recombined prediction value vectors, where the N recombined prediction value vectors are the first to N recombined prediction value vectors in the set of recombined prediction value vectors, and N is an integer greater than or equal to 1;

[0319] Update the network parameters of the classification network model by using a loss function according to the N recombined prediction value vectors.

[0320] The model training device provided by the present application can be deployed on a computer device. Taking this computer device as a server as an example for introduction, please refer toFigure 17 , Figure 17 is a schematic diagram of a server structure provided by an embodiment of the present application. The server 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 for storing application programs 342 or data 344 (for example, one or more mass storage devices). Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0321] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0322] The steps performed by the server in the above embodiments may be based on the Figure 17 shown server structure.

[0323] An embodiment of the present application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the methods described in the foregoing embodiments.

[0324] An embodiment of the present application also provides a computer program product including a program. When it runs on a computer, it causes the computer to execute the methods described in the foregoing embodiments.

[0325] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0326] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0327] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0328] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0329] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0330] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for model training, characterized in that, it includes: Obtain a set of data to be trained, wherein the set of data to be trained includes at least two data to be trained, and each data to be trained has the same true class label, and the data to be trained is one or more of images, texts, and voices; Based on the set of data to be trained, obtain a set of predicted value vectors through a classification network model, wherein the set of predicted value vectors includes at least two predicted value vectors, each predicted value vector corresponds to a data to be trained, and each predicted value vector includes a label element and M non-label elements, the set of predicted value vectors includes a target predicted value vector, and the absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among the M non-label elements is a first value, and M is an integer greater than or equal to 1; Sort the label element and the M non-label elements of each predicted value vector in the set of predicted value vectors to obtain a set of recombined predicted value vectors, wherein the set of recombined predicted value vectors includes a target recombined predicted value vector, and the absolute value of the difference between the activation value of the label element in the target recombined predicted value vector and the maximum activation value among the M non-label elements is a second value, and the second value is less than the first value; Train the classification network model according to the set of recombined predicted value vectors, wherein the classification network model is one or more of an image classification model, a text classification model, and an audio classification model.

2. The method according to claim 1, characterized in that, the obtaining of the set of data to be trained includes: Obtain an initial set of training data, wherein the initial set of training data includes Q data to be trained, and each data to be trained in the Q data to be trained has a true class label, and Q is an integer greater than or equal to 2; From the initial set of training data, obtain at least two data to be trained with the same true class label as the set of data to be trained.

3. The method according to claim 1, characterized in that, the sorting of the label element and the M non-label elements of each predicted value vector in the set of predicted value vectors to obtain a set of recombined predicted value vectors includes: For the label element of each predicted value vector in the set of predicted value vectors, arrange the column of activation values corresponding to the label elements in the set of predicted value vectors in descending order to obtain a first sequence; For the M non-label elements of each predicted value vector in the set of predicted value vectors, arrange the M columns of activation values corresponding to the M non-label elements in the set of predicted value vectors in ascending order to obtain M second sequences; Generate the set of recombined predicted value vectors according to the first sequence and the M second sequences.

4. The method according to claim 1, characterized in that, Sorting the label element and the M non-label elements of each predicted value vector in the predicted value vector set to obtain a recombined predicted value vector set, including: Obtaining a first sequence according to the label element of each predicted value vector in the predicted value vector set; For the M non-label elements of each predicted value vector in the predicted value vector set, arranging the M columns of activation values corresponding to the M non-label elements in the predicted value vector set in ascending order to obtain M second sequences; Generating the recombined predicted value vector set according to the first sequence and the M second sequences.

5. The method according to claim 1, wherein, Sorting the label element and the M non-label elements of each predicted value vector in the predicted value vector set to obtain a recombined predicted value vector set, including: For the label element of each predicted value vector in the predicted value vector set, arranging the column of activation values corresponding to the label element in the predicted value vector set in descending order to obtain a first sequence; For the M non-label elements of each predicted value vector in the predicted value vector set, obtaining the maximum activation value of the M non-label elements in each predicted value vector; For each predicted value vector in the predicted value vector set, replacing the activation values corresponding to the M non-label elements with the maximum activation value of the M non-label elements to obtain an updated predicted value vector corresponding to each predicted value vector; Arranging the M columns of activation values corresponding to the updated predicted value vector corresponding to each predicted value vector in ascending order to obtain M second sequences; Generating the recombined predicted value vector set according to the first sequence and the M second sequences.

6. The method according to claim 1, wherein, Sorting the label element and the M non-label elements of each predicted value vector in the predicted value vector set to obtain a recombined predicted value vector set, including: For the label element of each predicted value vector in the predicted value vector set, arranging the column of activation values corresponding to the label element in the predicted value vector set in descending order to obtain a first sequence; For the M non-label elements of each predicted value vector in the predicted value vector set, obtaining K maximum activation values of the M non-label elements in each predicted value vector, where the K maximum activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M; For each predicted value vector in the predicted value vector set, replacing the K minimum activation values among the M non-label elements with the K maximum activation values to obtain an updated predicted value vector corresponding to each predicted value vector, where the K minimum activation values represent the first K activation values after arranging the M non-label elements in ascending order; Arrange the M column activation values corresponding to the updated prediction value vectors corresponding to each of the prediction value vectors in ascending order to obtain M second sequences; Generate the set of recombined prediction value vectors according to the first sequence and the M second sequences.

7. The method according to any one of claims 3 to 6, characterized in that, The training of the classification network model according to the set of recombined prediction value vectors includes: Obtain N recombined prediction value vectors from the set of recombined prediction value vectors, where the N recombined prediction value vectors are the last 1 to N recombined prediction value vectors in the set of recombined prediction value vectors, and N is an integer greater than or equal to 1; Update the network parameters of the classification network model using a loss function according to the N recombined prediction value vectors.

8. The method according to claim 1, characterized in that, The sorting of the label elements and the M non-label elements of each prediction value vector in the set of prediction value vectors to obtain a set of recombined prediction value vectors includes: Arrange the label elements of each prediction value vector in the set of prediction value vectors in ascending order to obtain a first sequence; Arrange the M non-label elements of each prediction value vector in the set of prediction value vectors in descending order to obtain M second sequences; Generate the set of recombined prediction value vectors according to the first sequence and the M second sequences.

9. The method according to claim 1, characterized in that, The sorting of the label elements and the M non-label elements of each prediction value vector in the set of prediction value vectors to obtain a set of recombined prediction value vectors includes: Obtain a first sequence according to the label elements of each prediction value vector in the set of prediction value vectors; Arrange the M non-label elements of each prediction value vector in the set of prediction value vectors in descending order to obtain M second sequences; Generate the set of recombined prediction value vectors according to the first sequence and the M second sequences.

10. The method according to claim 1, characterized in that, The sorting of the label elements and the M non-label elements of each prediction value vector in the set of prediction value vectors to obtain a set of recombined prediction value vectors includes: Arrange the column activation values corresponding to the label elements in the set of prediction value vectors in ascending order for the label elements of each prediction value vector in the set of prediction value vectors to obtain a first sequence; For the M non-label elements of each prediction value vector in the set of prediction value vectors, obtain the maximum activation value of the M non-label elements in each prediction value vector; For each prediction value vector in the set of prediction value vectors, replace the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain the updated prediction value vector corresponding to each prediction value vector; Arrange the M-column activation values corresponding to the updated prediction value vectors corresponding to each of the prediction value vectors in ascending order to obtain M second sequences; Generate the set of recombined prediction value vectors according to the first sequence and the M second sequences.

11. The method according to claim 1, wherein, the sorting the label element and the M non-label elements of each prediction value vector in the set of prediction value vectors to obtain a set of recombined prediction value vectors includes: For the label element of each prediction value vector in the set of prediction value vectors, arrange the activation values of a column corresponding to the label element in the set of prediction value vectors in ascending order to obtain a first sequence; For the M non-label elements of each prediction value vector in the set of prediction value vectors, obtain the K largest activation values of the M non-label elements in each prediction value vector, where the K largest activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M; For each prediction value vector in the set of prediction value vectors, replace the K smallest activation values among the M non-label elements with the K largest activation values to obtain the updated prediction value vector corresponding to each prediction value vector, where the K smallest activation values represent the first K activation values after arranging the M non-label elements in ascending order; Arrange the M-column activation values corresponding to the updated prediction value vectors corresponding to each of the prediction value vectors in ascending order to obtain M second sequences; Generate the set of recombined prediction value vectors according to the first sequence and the M second sequences.

12. The method according to any one of claims 8 to 11, wherein, the training the classification network model according to the set of recombined prediction value vectors includes: Obtain N recombined prediction value vectors from the set of recombined prediction value vectors, where the N recombined prediction value vectors are the first to N recombined prediction value vectors in the set of recombined prediction value vectors, and N is an integer greater than or equal to 1; Update the network parameters of the classification network model by using a loss function according to the N recombined prediction value vectors.

13. A model training device, wherein, it includes: An acquisition module for acquiring a set of data to be trained, where the set of data to be trained includes at least two pieces of data to be trained, and each piece of data to be trained has the same true class label, and the data to be trained is one or more of image, text, and voice data; The obtaining module is further configured to obtain a set of predicted value vectors through a classification network model based on the set of data to be trained, where the set of predicted value vectors includes at least two predicted value vectors, each predicted value vector corresponds to a piece of data to be trained, each of the predicted value vectors includes a label element and M non-label elements, the set of predicted value vectors includes a target predicted value vector, and the absolute value of the difference between the activation value of the label element in the target predicted value vector and the maximum activation value among the M non-label elements is a first value, where M is an integer greater than or equal to 1; The sorting module is configured to sort the label element and the M non-label elements of each predicted value vector in the set of predicted value vectors to obtain a set of reorganized predicted value vectors, where the set of reorganized predicted value vectors includes a target reorganized predicted value vector, and the absolute value of the difference between the activation value of the label element in the target reorganized predicted value vector and the maximum activation value among the M non-label elements is a second value, and the second value is less than the first value; The training module is configured to train the classification network model according to the set of reorganized predicted value vectors, where the classification network model is one or more of an image classification model, a text classification model, and an audio classification model.

14. The apparatus according to claim 13, wherein, The obtaining module is specifically configured to: Obtain an initial set of training data, where the initial set of training data includes Q pieces of data to be trained, and each piece of data to be trained in the Q pieces of data to be trained has a true class label, and Q is an integer greater than or equal to 2; From the initial set of training data, obtain at least two pieces of data to be trained with the same true class label as the set of data to be trained.

15. The apparatus according to claim 13, wherein, The sorting module is specifically configured to: For the label element of each predicted value vector in the set of predicted value vectors, arrange the column of activation values corresponding to the label element in the set of predicted value vectors in descending order to obtain a first sequence; For the M non-label elements of each predicted value vector in the set of predicted value vectors, arrange the M columns of activation values corresponding to the M non-label elements in the set of predicted value vectors in ascending order to obtain M second sequences; Generate the set of reorganized predicted value vectors according to the first sequence and the M second sequences.

16. The apparatus according to claim 13, wherein, The sorting module is specifically configured to: Obtain a first sequence according to the label element of each predicted value vector in the set of predicted value vectors; For the M non-label elements of each predicted value vector in the set of predicted value vectors, arrange the M columns of activation values corresponding to the M non-label elements in the set of predicted value vectors in ascending order to obtain M second sequences; Generate the set of reorganized predicted value vectors according to the first sequence and the M second sequences.

17. The apparatus according to claim 13, It is characterized in that The sorting module is specifically used for: For the label element of each prediction value vector in the prediction value vector set, arranging a column of activation values corresponding to the label element in the prediction value vector set in descending order to obtain a first sequence; For the M non-label elements of each prediction value vector in the prediction value vector set, obtaining the maximum activation value of the M non-label elements in each prediction value vector; For each prediction value vector in the prediction value vector set, replacing the activation values corresponding to the M non-label elements with the maximum activation values of the M non-label elements to obtain the updated prediction value vector corresponding to each prediction value vector; Arranging the M columns of activation values corresponding to the updated prediction value vector corresponding to each prediction value vector in ascending order to obtain M second sequences; Generating the recombined prediction value vector set according to the first sequence and the M second sequences.

18. The device according to claim 13, It is characterized in that The sorting module is specifically used for: For the label element of each prediction value vector in the prediction value vector set, arranging a column of activation values corresponding to the label element in the prediction value vector set in descending order to obtain a first sequence; For the M non-label elements of each prediction value vector in the prediction value vector set, obtaining K maximum activation values of the M non-label elements in each prediction value vector, where the K maximum activation values represent the first K activation values after arranging the M non-label elements in descending order, and K is an integer greater than 1 and less than M; For each prediction value vector in the prediction value vector set, replacing the K minimum activation values among the M non-label elements with the K maximum activation values to obtain the updated prediction value vector corresponding to each prediction value vector, where the K minimum activation values represent the first K activation values after arranging the M non-label elements in ascending order; Arranging the M columns of activation values corresponding to the updated prediction value vector corresponding to each prediction value vector in ascending order to obtain M second sequences; Generating the recombined prediction value vector set according to the first sequence and the M second sequences.

19. The device according to any one of claims 15 to 18, It is characterized in that The training module is specifically used for: Obtaining N recombined prediction value vectors from the recombined prediction value vector set, where the N recombined prediction value vectors are the last 1 to N recombined prediction value vectors in the recombined prediction value vector set, and N is an integer greater than or equal to 1; Updating the network parameters of the classification network model by using a loss function according to the N recombined prediction value vectors.

20. A computer device, It is characterized in that Comprising: A memory, a processor, and a bus system; Wherein, the memory is used for storing programs; The processor is configured to execute the program in the memory, and the processor is configured to execute the method according to any one of claims 1 to 12 based on the instructions in the program code; The bus system is configured to connect the memory and the processor to enable communication between the memory and the processor.

21. A computer-readable storage medium comprising instructions that, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 12.

22. A computer program product, wherein, the computer program product comprises a program that, when run on a computer, causes the computer to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Face recognition method based on semi-supervised training

    CN110472533A

  • Model training method, text classification method and device and storage medium

    CN111368078A