A classification model training method, an image classification method, and related devices

By using a weight prediction model to train the target weights for each exit in a dynamic early termination neural network model, the adaptive inference characteristics of the model are optimized, improving the image recognition accuracy and efficiency on edge devices and solving the problem of low recognition accuracy in existing technologies.

CN116958616BActive Publication Date: 2026-06-19CHINA MOBILE COMM LTD RES INST +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210373908.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2026-06-19
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

Existing dynamic early termination neural network models neglect the adaptive inference characteristics of the model during training, resulting in poor performance, especially in video surveillance scenarios of edge devices where computing resources are limited, leading to low recognition accuracy.

Method used

A weight prediction model is used to train a multi-exit dynamic neural network model. By predicting the target weight of each exit, the image classification model is trained based on the loss value of the sample data to optimize the adaptive inference characteristics of the model.

Benefits of technology

It improves the recognition accuracy and efficiency of image classification models on edge devices, reduces resource costs, and is suitable for video surveillance systems on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958616B_ABST
    Figure CN116958616B_ABST
Patent Text Reader

Abstract

This application provides a classification model training method, an image recognition method, and related equipment. The classification model training method includes: using an image classification model to perform forward propagation calculations on multiple first sample data to obtain first loss values ​​for the multiple first sample data. The image classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​at the multiple exits. The multiple first sample data are image frame data obtained from the first sampling. A weight prediction model is used to predict weights on the first loss values ​​of the multiple first sample data to obtain target weights for each of the multiple first sample data at each exit. Based on the predicted weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the image classification model is trained to obtain an image classification model for classifying images. This application can improve model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a classification model training method, an image classification method, and related equipment. Background Technology

[0002] In video surveillance scenarios, due to bandwidth limitations, edge devices are typically used to receive video streams sent by cameras, perform inference using deep learning models, and send the inference results to the client or upload the inference results to a cloud server.

[0003] Edge devices, with their lower power consumption and weaker computing power, can utilize multi-exit networks with dynamic "early termination" mechanisms. These networks adaptively terminate computations early based on input samples, reducing unnecessary redundant calculations. However, existing dynamic "early termination" neural network models often employ the same training strategies as static neural network models, neglecting the model's adaptive inference characteristics and resulting in poor model performance. Summary of the Invention

[0004] This application provides a classification model training method, an image classification method, and related equipment to address the problem of poor model performance.

[0005] In a first aspect, embodiments of this application provide a classification model training method, including:

[0006] An image classification model is used to perform forward propagation calculations on multiple first sample data to obtain the first loss value of the multiple first sample data. The classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​of the multiple exits. The multiple first sample data are image frame data obtained by first sampling.

[0007] The weight prediction model is used to predict the weight of the first loss value of the plurality of first sample data, so as to obtain the target weight of the plurality of first sample data at each exit.

[0008] Based on the prediction weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the image classification model is trained to obtain an image classification model for classifying images.

[0009] Secondly, embodiments of this application also provide an image classification method, including:

[0010] Image frames from the video to be identified are input into an image classification model for category prediction to obtain prediction results. The image classification model includes multiple outputs.

[0011] If the confidence level of the predicted category of the target exit output among the plurality of exits is greater than a preset threshold, the predicted category of the target exit output is determined to be the category of the image frame.

[0012] The image classification model is trained using the classification model training method disclosed in the first aspect.

[0013] Thirdly, embodiments of this application also provide an electronic device, including: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method described in the first aspect of the embodiments of this application, or to implement the steps in the method described in the second aspect of the embodiments of this application.

[0014] Fourthly, embodiments of this application also provide a readable storage medium storing a program, which, when executed by a processor, implements the steps in the method described in the first aspect of embodiments of this application, or implements the steps in the method described in the second aspect of embodiments of this application.

[0015] In this embodiment, an image classification model is used to perform forward propagation calculations on multiple first sample data to obtain first loss values ​​for the multiple first sample data. A weight prediction model is then used to predict the weights of the first loss values ​​for the multiple first sample data, obtaining target weights for each exit point. Based on the target weights of the multiple first sample data at each exit point and the first loss values, the image classification model is trained to obtain an image classification model for image classification. By using the weight prediction network to predict the weights of the output loss values ​​of each sample data at different exit points, the classification model can be trained based on the target weights of the multiple first sample data at each exit point and the first loss values. This allows the obtained image classification model to utilize its adaptive inference characteristics during training, thereby improving its performance. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of this application, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a classification model training method provided in an embodiment of this application;

[0018] Figure 2 This is a schematic flowchart of an image recognition method provided in an embodiment of this application;

[0019] Figure 3 This is a schematic diagram illustrating the alternating optimization of a backbone network and a weight prediction network provided in an embodiment of this application;

[0020] Figure 4 This is a schematic diagram of the structure of a classification model training device provided in an embodiment of this application;

[0021] Figure 5 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0022] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0023] Figure 7 This is a schematic diagram of the structure of another electronic device provided in an embodiment of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.

[0026] Please see Figure 1 , Figure 1 This is a flowchart illustrating a classification model training method provided in an embodiment of this application, as shown below. Figure 1 As shown, it includes the following steps:

[0027] Step 101: Use an image classification model to perform forward propagation calculations on multiple first sample data to obtain the first loss value of the multiple first sample data. The classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​of the multiple exits. The multiple first sample data are image frame data obtained from the first sampling.

[0028] The image classification model described above includes multiple exits. For example, it can be a multi-scale dense network (MSDNet) model, a resolution-adaptive network (RANet) model, or other multi-exit dynamic neural network models. It should be understood that before applying the image classification model to image recognition, it needs to be trained multiple times to update its parameters and optimize its recognition performance.

[0029] It should be noted that the image classification model in step 101 above can be an initial image classification model (with initial model parameters) or a model obtained after one or more training iterations (with model parameters updated one or more times). For example, taking the example that the above image classification model needs to undergo T iterations of training before being applied to image recognition, in the t (0≤t≤T) training process, if t=0, then the image classification model used in step 101 in this training process is the initial classification model, that is, the parameters of the above classification model are the initial preset values; if 0<t≤T, then the classification model used in step 101 in this training process is the image classification model obtained after t-1 training iterations, and the parameters of the above classification model have been updated t-1 times.

[0030] During model training, the image classification model can first be forward-propagated on multiple sets of first sample data to obtain the output value of each first sample data at each exit, thereby determining the loss value of each first sample data at each exit. The loss values ​​of each first sample data at these multiple exits can then be used to update the classification model, completing one model training cycle.

[0031] The aforementioned image frame data can be multiple frames of image data from a video. The multiple first sample data can be randomly sampled from a pre-acquired training sample set and used as training data for this classification model training.

[0032] Step 102: Use a weight prediction model to predict the weights of the first loss values ​​of the multiple first sample data, and obtain the target weights of the multiple first sample data at each exit.

[0033] In existing technologies, the training process for dynamic neural network models typically involves summing the loss values ​​of all exits on the training data and then backpropagating to update the model. This means that regardless of the simplicity of the training data or which exit point it exits from, the loss value at each exit carries equal weight in the overall loss function. However, in multi-exit dynamic neural network models, the subnetworks represented by each exit have varying depths and expressive capabilities, and the difficulty of the samples they need to predict during the inference phase also differs. If the model is updated according to the above training strategy, the impact of the training samples output from each exit on the total loss value will be consistent, leading to the neglect of the model's adaptive inference characteristics and resulting in low accuracy.

[0034] In this embodiment, the weight prediction model is used to predict the weights of each sample data at different exits. Specifically, the first loss value of each first sample data is used as the input to the weight prediction model to obtain the weights of each first sample at different exits. By inputting the loss values ​​of a first sample data at different exits into the weight prediction model, the weights of that first sample data at different exits are obtained, thus distinguishing the weights of the loss values ​​of the same sample data at different exits, thereby conforming to the model structure of the classification network described above.

[0035] Step 103: Based on the target weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, train the classification model to obtain an image classification model for classifying images.

[0036] It can be understood that by obtaining the prediction weights of the multiple first sample data at each exit, the weighted sum of the losses of the multiple first sample data at each exit can be minimized during the training process, thereby achieving the training of the above classification model. In other words, the sample data used for training can be used to train the model based on the structure of each exit, updating the model parameters of the above image classification model, thereby improving the accuracy of the image classification model in recognizing images.

[0037] In this embodiment, a classification model is used to perform forward propagation calculations on multiple first sample data to obtain first loss values ​​for the multiple first sample data. A weight prediction model is then used to predict the weights of the first loss values ​​for the multiple first sample data, obtaining target weights for each exit point. Based on the target weights of the multiple first sample data at each exit point and the first loss values, the classification model is trained to obtain an image classification model for image classification. By using the weight prediction network to predict the weights of the output loss values ​​of each sample data at different exit points, the classification model can be trained based on the target weights of the multiple first sample data at each exit point and the first loss values. This allows the obtained image classification model to utilize its adaptive inference characteristics during training, thereby improving its performance.

[0038] Optionally, step 102, which involves using a weight prediction model to predict the weights of the first loss values ​​of the plurality of first sample data to obtain the target weights of the plurality of first sample data at each exit, includes:

[0039] The weight prediction model is used to predict the weights of the first loss values ​​of the multiple first sample data, so as to obtain the predicted weights of the multiple first sample data at each exit.

[0040] Based on the prediction weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the model parameters of the classification model are updated to obtain a reference classification model.

[0041] The reference classification model is used to classify and predict multiple second sample data respectively, so as to obtain a sample data set for each exit and a prediction result for each second sample data. The multiple second sample data are image frame data obtained by second sampling, and the multiple second sample data include sample data in the sample data set.

[0042] Based on the sample data set of each exit and the prediction result of each second sample data, the second loss value is calculated;

[0043] The weight prediction model is updated using the second loss value;

[0044] The updated weight prediction model is used to predict the weights of the first loss values ​​of the plurality of first sample data, so as to obtain the target weights of the plurality of first sample data at each exit.

[0045] The second sample data mentioned above can also be randomly sampled from a pre-acquired training sample set and used as training data for this classification model training. Furthermore, the second sample data and the first sample data can be sampled separately from the pre-acquired training sample set. For example, multiple sample data can be randomly sampled from the pre-acquired training sample set, and then these multiple sample data can be randomly divided into the first sample data and the second sample data.

[0046] The aforementioned reference classification model can be understood as a model obtained by pseudo-updating the aforementioned image classification model. That is, the obtained reference classification model is used to update the weight prediction network, and the updated weight prediction network is used to predict the weights of the sample data at each exit. The training and updating of the aforementioned image classification model is based on the loss value of the sample data at each exit, which is weighted according to the weights predicted by the updated weight prediction network to obtain the overall loss value. The overall loss value is then used to train and update the aforementioned image classification model.

[0047] It is understood that by obtaining the target weights of the multiple first sample data at each exit through the weight prediction network, the weight prediction network can be optimized, thereby improving the accuracy of the weight prediction network in predicting weights.

[0048] In this embodiment, the weight prediction network and the classification network are updated alternately. By simulating the adaptive inference process of the network on the multiple second sample data, a meta-loss function is defined, and the weight prediction network is optimized with the meta-loss function. Then, the updated weight prediction network is used to predict the weights of the first loss values ​​of the multiple first sample data, so as to obtain the target weights of the multiple first sample data at each exit, thereby improving the prediction accuracy of the weights.

[0049] Optionally, the step of using the reference classification model to classify and predict multiple second sample data to obtain a sample data set for each exit and a prediction result for each second sample data includes:

[0050] Obtain the output sample ratio for each outlet;

[0051] The reference classification model is used to classify and predict multiple second sample data respectively, and the results of each second sample data output at the multiple outlets are obtained respectively.

[0052] Based on the output sample ratio of each exit and the results of each second sample data being output at the multiple exits, the sample data set of each exit and the prediction result of each second sample data are determined.

[0053] The output sample ratio for each exit can be pre-determined based on empirical values ​​to match the structure of the image classification model with multiple exits. It can be understood that for different numbers of second sample data, the number of sample data sets for each exit can be determined based on the total amount of the second sample data and the output sample ratio for each exit.

[0054] It is understandable that the above classification network, as a multi-exit dynamic neural network model, can output recognition results at shallower exits for simple images, and at deeper exits for complex images, in order to achieve a balance between efficiency and accuracy.

[0055] Optionally, the results of each second sample data output at the plurality of outlets include the prediction confidence of each second sample data output at the plurality of outlets.

[0056] The determination of the sample data set for each exit and the prediction result for each second sample data based on the output sample ratio of each exit and the results of each second sample data output at the multiple exits includes:

[0057] Obtain the order of the multiple exits in the classification model;

[0058] Based on the prediction confidence of each second sample data output at the multiple exits and the order thereof, the sample data set of each exit and the prediction result of each second sample data are determined sequentially, and the number of samples in the sample data set of each exit matches the output sample ratio of each exit.

[0059] It should be noted that in the above multi-exit classification network, exits ranked earlier require less computation and are suitable for simple identification samples, while exits ranked later require more computation and are suitable for complex identification samples. For results that can be identified at an earlier exit with a confidence level greater than a certain threshold, the identification result can be output from that earlier exit to reduce the model's computational load. For cases where no result can be identified at an earlier exit, subsequent exits are then evaluated to determine whether they can yield identification results with a confidence level greater than a certain threshold, thus ensuring the accuracy of the identification results.

[0060] During training, the results of each second sample data point at multiple exits can be obtained, and the accuracy of each second sample data point's output at multiple exits can be evaluated using prediction confidence. For results where a prediction confidence greater than a certain threshold can be output at a shallower exit, these sample data points can be grouped into a sample data set for that shallower exit, and the image classification model can be updated via backpropagation. Furthermore, the sample data sets for each exit can be determined sequentially according to the order of each exit in the image classification model, i.e., from shallower exits to deeper exits. Thus, for results where a prediction confidence greater than a certain threshold can be obtained at a shallower exit, the output can be directly from the shallower exit and used to train the image classification model; for results where a prediction confidence greater than a certain threshold can only be obtained at a deeper exit, the output can be from the deeper exit and used to train the image classification model.

[0061] Please see Figure 2 , Figure 2 This application provides an image recognition method, such as... Figure 2 As shown, it includes the following steps:

[0062] Step 201: Input the image frames of the video to be identified into the image classification model for category prediction and obtain the prediction result. The image classification model includes multiple outputs.

[0063] Step 202: If the confidence level of the predicted category of the target exit output among the multiple exits indicated by the prediction result is greater than a preset threshold, determine the predicted category of the target exit output as the category of the image frame.

[0064] The image classification model is trained using the classification model training method described in the above embodiment.

[0065] It should be noted that this embodiment is as a comparison with... Figure 1 The implementation method of the image classification model corresponding to the illustrated embodiment can be found in the following examples. Figure 1 The relevant descriptions in the illustrated embodiments will not be repeated here to avoid repetition.

[0066] For ease of understanding, a specific example is as follows:

[0067] This application provides a sample weighting method for dynamic "early termination" neural networks, the specific steps of which are as follows:

[0068] (1) Establishing a system with Export classification network Let its parameters be ;

[0069] (2) Establish a weighted prediction network It consists of two fully connected layers, with the first layer having an input dimension of . The output dimension is ( The value is determined based on experience, for example, 500); the second layer input dimension is... The output dimension is .network The input is the network. of Export Classification ;

[0070] (3) Alternate optimization within the meta-learning framework and ,like Figure 3 As shown, the specific process includes the following:

[0071] (3-1) Let Let the maximum number of iterations be... ;

[0072] (3-2) When Perform the following steps:

[0073] i. Sampling to obtain batch data ;

[0074] ii. For the obtained batch data Randomly divide it into training data and metadata ;

[0075] iii. Use The classification network at each time step in the training data Perform forward propagation to obtain the classification loss:

[0076]

[0077] in, Indicates the first The loss value for each training data set. Indicates the first The training data at the th th The loss value of each export ( ), Representing training data The quantity.

[0078] iv. Input the training loss Weight prediction network at time step This yields the weight of each sample at each exit:

[0079]

[0080] in, Indicates the first The weights of each training data point at each exit. express Weight prediction network at time step, Indicates the first The loss value for each training data set. express The weights at each time step predict the network parameters of the network.

[0081] v. Weights generated based on predictions ; for classification networks Perform a pseudo-update ( (Learning rate):

[0082]

[0083] in, Representation of classification networks Network parameters after pseudo-update express Network parameters of the classification network at different times. This indicates the maximum number of iterations.

[0084] vi. Classification network based on pseudo-update (Network parameters are) (Classification network), in metadata The above simulates the dynamic reasoning process to obtain the sample sets that exit at each exit point from the metadata. Specifically, set the sample ratio for each outlet, and then determine the number of samples that each outlet should output based on this ratio. Further, based on the network output, we obtain... The prediction confidence scores of each sample at the first exit were calculated and ranked, with the highest confidence score being selected. A set of samples At the second exit, the confidence level is also ranked, and then... The middle belongs to the set Exclude samples from the first group, and then select the sample with the highest confidence level from the remaining samples. A set of samples ; and so on, we get It should be understood that for the last exit, its target set is all samples that were not selected by the preceding exits, i.e. ( This indicates removal, that is, the first... The sample set of each export is metadata. Remove the first To the (sample set of samples in the sample set of each export).

[0085] vii. Calculate the meta-loss function:

[0086]

[0087] in, Represents the meta-loss function. Indicates the first The number of samples that should be exported from each outlet. Indicates the first The set of samples that should be exported. Represents the sample set The first in One sample, Indicates sample The corresponding tag value.

[0088] viii. Based on meta-loss function The parameter prediction network is updated in one step. (Learning rate):

[0089]

[0090] in, express The weights at time points predict the network parameters of the network. express The weights at each time step predict the network parameters of the network.

[0091] ix. Prediction network based on updated weights Regenerate the weights of each sample at each exit:

[0092]

[0093] in, This indicates the use of the updated weights to predict the network. The generated first The weight of each sample at each exit Indicates after the update Weight prediction network at each time step.

[0094] x. Based on the new weights Classification networks Update ( (for learning rate)

[0095]

[0096] in, Indicates after the update Classification network parameters at each time step.

[0097] xi. Order Repeat the steps above (3-2), as follows: Figure 3 As shown;

[0098] (3-3) Obtain the parameters of the trained classification network. .

[0099] The weight prediction network is no longer used in the testing and inference phases.

[0100] In this embodiment, a weighted prediction network is employed. For dynamic multi-exit networks, the loss function of each exit is used as input to predict the training weights of samples at each exit. Compared to existing dynamic network training methods, this approach fully considers the adaptive computation mode of the model during the inference phase. A meta-learning algorithm is used to weight different training samples at each exit, thereby improving the network's performance in adaptive inference scenarios and achieving a better balance between accuracy and efficiency. Therefore, the classification model in this embodiment can be effectively applied to video surveillance systems on edge devices, improving the utilization rate of edge devices and reducing resource costs.

[0101] See Figure 4 , Figure 4 This is a schematic diagram of the structure of a classification model training device provided in an embodiment of this application. Figure 4 As shown, the classification model training device 400 includes:

[0102] The calculation module 401 is used to perform forward propagation calculation on multiple first sample data using an image classification model to obtain a first loss value of the multiple first sample data. The image classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​of the multiple exits. The multiple first sample data are image frame data obtained by first sampling.

[0103] The first prediction module 402 is used to use a weight prediction model to predict the weight of the first loss value of the plurality of first sample data, so as to obtain the target weight of the plurality of first sample data at each exit.

[0104] The training module 403 is used to train the image classification model based on the prediction weights of the plurality of first sample data at each exit and the first loss values ​​of the plurality of first sample data, so as to obtain an image classification model for classifying images.

[0105] Optionally, the first prediction module 402 includes:

[0106] The first prediction unit is used to use a weighted prediction model to predict the weights of the first loss values ​​of the plurality of first sample data, so as to obtain the predicted weights of the plurality of first sample data at each exit.

[0107] The first update unit is used to update the model parameters of the image classification model based on the prediction weights of the plurality of first sample data at each exit and the first loss values ​​of the plurality of first sample data, so as to obtain a reference classification model.

[0108] The second prediction unit is used to classify and predict multiple second sample data using the reference classification model to obtain a sample data set for each exit and a prediction result for each second sample data. The multiple second sample data are image frame data obtained from the second sampling, and the multiple second sample data include sample data in the sample data set.

[0109] The calculation unit is used to calculate the second loss value based on the sample data set of each exit and the prediction result of each second sample data.

[0110] The second update unit is used to update the weight prediction model using the second loss value;

[0111] The third prediction unit is used to predict the weights of the first loss values ​​of the plurality of first sample data using the updated weight prediction model, so as to obtain the target weights of the plurality of first sample data at each exit.

[0112] Optionally, the second prediction unit includes:

[0113] A sub-unit is used to obtain the output sample ratio of each outlet;

[0114] The prediction subunit is used to perform classification prediction on multiple second sample data using the reference classification model, and to obtain the results of each second sample data output at the multiple outlets.

[0115] A subunit is defined for determining the sample data set for each exit and the prediction result for each second sample data based on the output sample ratio of each exit and the results of each second sample data being output at the multiple exits.

[0116] Optionally, the results of each second sample data output at the plurality of outlets include the prediction confidence of each second sample data output at the plurality of outlets.

[0117] The determining subunit is used for:

[0118] Obtain the order of the multiple exits in the classification model;

[0119] Based on the prediction confidence of each second sample data output at the multiple exits and the order thereof, the sample data set of each exit and the prediction result of each second sample data are determined sequentially, and the number of samples in the sample data set of each exit matches the output sample ratio of each exit.

[0120] The classification model training device 400 can achieve the functions described in the embodiments of this application. Figure 1 The various processes in the method embodiments, and the ways to achieve the same beneficial effects, will not be repeated here to avoid repetition.

[0121] See Figure 5 , Figure 5 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application. Figure 5 As shown, the image recognition device 500 includes:

[0122] The second prediction module 501 is used to input the image frames of the video to be identified into the image classification model to predict the category and obtain the prediction result. The image classification model includes multiple outputs.

[0123] The determining module 502 is used to determine the predicted category of the target exit output as the category of the image frame when the confidence of the predicted category of the target exit output among the plurality of exits indicated by the prediction result is greater than a preset threshold.

[0124] The image classification model is trained using the classification model training method described above.

[0125] The image recognition device 500 can realize the embodiments of this application. Figure 2 The various processes in the method embodiments, and the ways to achieve the same beneficial effects, will not be repeated here to avoid repetition.

[0126] This application also provides an electronic device. Because the principle by which the electronic device solves the problem is similar to that in the embodiments of this application... Figure 1 The training method for the classification models shown is similar; therefore, the implementation of this electronic device can be found in the implementation of this method, and repeated details will not be elaborated upon. For example... Figure 6 As shown, the electronic device of this application embodiment includes a memory 620, a transceiver 610, and a processor 600;

[0127] The memory 620 is used to store computer programs; the transceiver 610 is used to send and receive data under the control of the processor 600; the processor 600 is used to read the computer program in the memory 620 and perform the following operations:

[0128] The image classification model is used to perform forward propagation calculations on multiple first sample data to obtain the first loss value of the multiple first sample data. The image classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​of the multiple exits. The multiple first sample data are image frame data obtained by first sampling.

[0129] The weight prediction model is used to predict the weight of the first loss value of the plurality of first sample data, so as to obtain the target weight of the plurality of first sample data at each exit.

[0130] Based on the prediction weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the image classification model is trained to obtain an image classification model for classifying images.

[0131] Among them, Figure 6 In this context, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 600) and memory (memory 620). The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface. Transceiver 610 may be multiple elements, including transmitters and transceivers, providing a unit for communicating with various other devices over a transmission medium. Processor 600 is responsible for managing the bus architecture and general processing, and memory 620 may store data used by processor 600 during operation.

[0132] The processor 600 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.

[0133] Optionally, the step of using a weight prediction model to predict the weights of the first loss values ​​of the plurality of first sample data to obtain the target weights of the plurality of first sample data at each exit includes:

[0134] The weight prediction model is used to predict the weights of the first loss values ​​of the multiple first sample data, so as to obtain the predicted weights of the multiple first sample data at each exit.

[0135] Based on the prediction weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the model parameters of the image classification model are updated to obtain a reference classification model.

[0136] The reference classification model is used to classify and predict multiple second sample data respectively, so as to obtain a sample data set for each exit and a prediction result for each second sample data. The multiple second sample data are image frame data obtained by second sampling, and the multiple second sample data include sample data in the sample data set.

[0137] Based on the sample data set of each exit and the prediction result of each second sample data, the second loss value is calculated;

[0138] The weight prediction model is updated using the second loss value;

[0139] The updated weight prediction model is used to predict the weights of the first loss values ​​of the plurality of first sample data, so as to obtain the target weights of the plurality of first sample data at each exit.

[0140] Optionally, the step of using the reference classification model to classify and predict multiple second sample data to obtain a sample data set for each exit and a prediction result for each second sample data includes:

[0141] Obtain the output sample ratio for each outlet;

[0142] The reference classification model is used to classify and predict multiple second sample data respectively, and the results of each second sample data output at the multiple outlets are obtained respectively.

[0143] Based on the output sample ratio of each exit and the results of each second sample data being output at the multiple exits, the sample data set of each exit and the prediction result of each second sample data are determined.

[0144] Optionally, the results of each second sample data output at the plurality of outlets include the prediction confidence of each second sample data output at the plurality of outlets.

[0145] The determination of the sample data set for each exit and the prediction result for each second sample data based on the output sample ratio of each exit and the results of each second sample data output at the multiple exits includes:

[0146] Obtain the order of the multiple exits in the classification model;

[0147] Based on the prediction confidence of each second sample data output at the multiple exits and the order thereof, the sample data set of each exit and the prediction result of each second sample data are determined sequentially, and the number of samples in the sample data set of each exit matches the output sample ratio of each exit.

[0148] This application also provides an electronic device. Because the principle by which the electronic device solves the problem is similar to that in the embodiments of this application... Figure 2 The image recognition methods shown are similar; therefore, the implementation of this electronic device can be found in the implementation of the method, and repeated details will not be elaborated further. For example... Figure 7 As shown, the electronic device of this application embodiment includes a memory 720, a transceiver 710, and a processor 700;

[0149] The memory 720 is used to store computer programs; the transceiver 710 is used to send and receive data under the control of the processor 700; the processor 700 is used to read the computer program in the memory 720 and perform the following operations:

[0150] Image frames from the video to be identified are input into an image classification model for category prediction to obtain prediction results. The image classification model includes multiple outputs.

[0151] If the confidence level of the predicted category of the target exit output among the plurality of exits is greater than a preset threshold, the predicted category of the target exit output is determined to be the category of the image frame.

[0152] The image classification model is trained using the classification model training method described above.

[0153] Among them, Figure 7 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 700) and memory (memory 720). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 710 can be multiple elements, including transmitters and transceivers, providing a unit for communicating with various other devices over a transmission medium. The processor 700 is responsible for managing the bus architecture and general processing, and the memory 720 can store data used by the processor 700 during operation.

[0154] The processor 700 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.

[0155] The electronic device provided in this application embodiment can perform the above-described functions. Figure 2 The method embodiments shown are similar in principle and technical effect, and will not be described again here.

[0156] This application also provides a readable storage medium storing a program that, when executed by a processor, implements the following... Figure 1 or Figure 2 The various processes in the Chinese method embodiment can achieve the same technical effect, and will not be described again here to avoid repetition.

[0157] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0159] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute some steps of the transmission and reception methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0160] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A classification model training method, characterized in that, include: The image classification model is used to perform forward propagation calculations on multiple first sample data to obtain the first loss value of the multiple first sample data. The image classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​of the multiple exits. The multiple first sample data are image frame data obtained by first sampling. The weight prediction model is used to predict the weight of the first loss value of the plurality of first sample data, so as to obtain the target weight of the plurality of first sample data at each exit. Based on the target weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the image classification model is trained to obtain an image classification model for classifying images. The step of using a weight prediction model to predict the weights of the first loss values ​​of the plurality of first sample data to obtain the target weights of the plurality of first sample data at each exit includes: The weight prediction model is used to predict the weights of the first loss values ​​of the multiple first sample data, so as to obtain the predicted weights of the multiple first sample data at each exit. Based on the prediction weights of the multiple first sample data at each exit and the first loss values ​​of the multiple first sample data, the model parameters of the image classification model are updated to obtain a reference classification model. The reference classification model is used to classify and predict multiple second sample data respectively, so as to obtain a sample data set for each exit and a prediction result for each second sample data. The multiple second sample data are image frame data obtained by second sampling, and the multiple second sample data include sample data in the sample data set. Based on the sample data set of each exit and the prediction result of each second sample data, the second loss value is calculated; The weight prediction model is updated using the second loss value; The updated weight prediction model is used to predict the weights of the first loss values ​​of the plurality of first sample data, so as to obtain the target weights of the plurality of first sample data at each exit.

2. The method of claim 1, wherein, The step of using the reference classification model to classify and predict multiple second sample data to obtain a sample data set for each exit and a prediction result for each second sample data includes: Obtain the output sample ratio for each outlet; The reference classification model is used to classify and predict multiple second sample data respectively, and the results of each second sample data output at the multiple outlets are obtained respectively. Based on the output sample ratio of each exit and the results of each second sample data being output at the multiple exits, the sample data set of each exit and the prediction result of each second sample data are determined.

3. The method of claim 2, wherein, The results of each second sample data output at the multiple exits include the prediction confidence of each second sample data output at the multiple exits. The determination of the sample data set for each exit and the prediction result for each second sample data based on the output sample ratio of each exit and the results of each second sample data output at the multiple exits includes: Obtain the order of the multiple exits in the image classification model; Based on the prediction confidence of each second sample data output at the multiple exits and the order thereof, the sample data set of each exit and the prediction result of each second sample data are determined sequentially, and the number of samples in the sample data set of each exit matches the output sample ratio of each exit.

4. An image classification method characterized by, include: Image frames from the video to be identified are input into an image classification model for category prediction to obtain prediction results. The image classification model includes multiple outputs. If the confidence level of the predicted category of the target exit output among the plurality of exits is greater than a preset threshold, the predicted category of the target exit output is determined to be the category of the image frame. The image classification model is trained using the classification model training method described in any one of claims 1 to 3.

5. A classification model training apparatus characterized by comprising: include: The calculation module is used to perform forward propagation calculation on multiple first sample data using an image classification model to obtain a first loss value for the multiple first sample data. The image classification model includes multiple exits, and the first loss value of each first sample data includes the loss values ​​of the multiple exits. The multiple first sample data are image frame data obtained by first sampling. The first prediction module is used to use a weight prediction model to predict the weight of the first loss value of the plurality of first sample data, so as to obtain the target weight of the plurality of first sample data at each exit. The training module is used to train the classification model based on the target weights of the plurality of first sample data at each exit and the first loss values ​​of the plurality of first sample data, so as to obtain an image classification model for classifying images. The first prediction module includes: The first prediction unit is used to use a weighted prediction model to predict the weights of the first loss values ​​of the plurality of first sample data, so as to obtain the predicted weights of the plurality of first sample data at each exit. The first update unit is used to update the model parameters of the classification model based on the prediction weights of the plurality of first sample data at each exit and the first loss values ​​of the plurality of first sample data, to obtain a reference classification model. The second prediction unit is used to classify and predict multiple second sample data using the reference classification model to obtain a sample data set for each exit and a prediction result for each second sample data. The multiple second sample data are image frame data obtained from the second sampling, and the multiple second sample data include sample data in the sample data set. The calculation unit is used to calculate the second loss value based on the sample data set of each exit and the prediction result of each second sample data. The second update unit is used to update the weight prediction model using the second loss value; The third prediction unit is used to predict the weights of the first loss values ​​of the plurality of first sample data using the updated weight prediction model, so as to obtain the target weights of the plurality of first sample data at each exit.

6. The apparatus of claim 5, wherein, The second prediction unit includes: A sub-unit is used to obtain the output sample ratio of each outlet; The prediction subunit is used to perform classification prediction on multiple second sample data using the reference classification model, and to obtain the results of each second sample data output at the multiple outlets. A subunit is defined for determining the sample data set for each exit and the prediction result for each second sample data based on the output sample ratio of each exit and the results of each second sample data being output at the multiple exits.

7. The apparatus of claim 6, wherein, The results of each second sample data output at the multiple exits include the prediction confidence of each second sample data output at the multiple exits. The determining subunit is used for: Obtain the order of the multiple exits in the classification model; Based on the prediction confidence of each second sample data output at the multiple exits and the order thereof, the sample data set of each exit and the prediction result of each second sample data are determined sequentially, and the number of samples in the sample data set of each exit matches the output sample ratio of each exit.

8. An image classification device, characterized in that, include: The second prediction module is used to input the image frames of the video to be identified into the image classification model to predict the category and obtain the prediction result. The image classification model includes multiple outputs. The determination module is used to determine the predicted category of the target exit output as the category of the image frame when the confidence level of the predicted category of the target exit output among the plurality of exits indicated by the prediction result is greater than a preset threshold. The image classification model is trained using the classification model training method described in any one of claims 1 to 3.

9. An electronic device, characterized in that, include: A transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor; characterized in that, The processor is configured to read a program from the memory to implement the steps of the method as described in any one of claims 1 to 3, or to implement the steps of the method as described in claim 4.

10. A computer readable storage medium for storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 3, or implements the steps of the method as described in claim 4.