Model updating method, data processing method, device, equipment, medium and product
By employing a differential update method between the student and teacher models and utilizing differential optimization of multi-scale sampling layers and classification layers, the problem of insufficient prediction performance of the model in scenarios such as contour point detection is solved, achieving better data processing results.
Patent Information
- Application Number
- CN202410642231.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2025-11-25
AI Technical Summary
In data processing scenarios such as contour point detection, the predictive performance of existing technologies needs improvement, especially in the extraction and integration of information at different scales.
A difference update method between the student model and the teacher model is adopted. By obtaining the difference between the first and second prediction results of the sample data, the sampling layer and classification layer in the student model are updated. Multi-scale information is obtained by using sampling layers with different sampling window sizes, and the model is optimized by combining label information and learning loss.
It improves the model's prediction performance at different scales and enhances the overall effect of data processing, especially the accuracy of key point detection in scenarios such as contour point detection.
Smart Images

Figure CN121010840A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular, relates to a model updating method, a data processing method, an apparatus, a device, a medium and a product. BACKGROUND
[0002] In some scenarios, such as a contour point detection scenario, it is possible to perform certain processing on certain data, such as an image, such as face contour point detection or body contour point detection, to meet the data processing requirements of these scenarios. The face contour point detection is used to detect key points that can describe the contour of a face from an image. The body contour point detection is used to detect key points that can describe the contour of a body from an image.
[0003] It should be noted that the technical solutions provided by the present application are not limited to the contour point detection scenario, and can also be applied to other data processing related scenarios, such as target recognition scenarios, image segmentation scenarios, edge detection scenarios, text detection scenarios, speech recognition scenarios, etc. In addition, the processing object of the present application is not limited to images, but can also be speech, text, etc. SUMMARY
[0004] The present application provides a model updating method, a data processing method, an apparatus, a device, a medium and a product, which are beneficial to improve the data processing effect.
[0005] In order to achieve the above-mentioned purpose, the technical solutions provided by the present application are as follows:
[0006] The present application provides a model updating method, which comprises:
[0007] Obtaining sample data;
[0008] Determining a first prediction result corresponding to the sample data by using a model; the model comprises a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain classification results of output data of each first sampling layer; and the first prediction result comprises classification results of output data of the plurality of first sampling layers.
[0009] Updating part or all of the model according to the first prediction result.
[0010] In a possible implementation, the model comprises a feature extraction network and a detection network; the detection network is used to process output data of the feature extraction network; and the detection network comprises the first classification layer and the plurality of first sampling layers.
[0011] In a possible implementation, the model is a student model.
[0012] The method further comprises:
[0013] determining a second prediction result corresponding to the sample data by using a teacher model; the teacher model comprises a second classification layer and a plurality of second sampling layers; for any second sampling layer, a sampling window size of the second sampling layer is same as a sampling window size of a corresponding first sampling layer of the plurality of first sampling layers; the second classification layer is used to obtain classification results of output data of each second sampling layer; the second prediction result comprises the classification results of the output data of the plurality of second sampling layers;
[0014] updating part or all of the model according to the first prediction result comprises:
[0015] updating part or all of the student model according to a difference between the first prediction result and the second prediction result.
[0016] In a possible implementation, the updating part or all of the student model according to the difference between the first prediction result and the second prediction result comprises:
[0017] for any first sampling layer, determining a loss corresponding to the first sampling layer according to a difference between a classification result of output data of the first sampling layer and a classification result of output data of a reference sampling layer corresponding to the first sampling layer; the plurality of second sampling layers comprise the reference sampling layer, and a sampling window size of the reference sampling layer is same as a sampling window size of the first sampling layer;
[0018] updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers.
[0019] In a possible implementation, the plurality of second sampling layers comprise a global sampling layer.
[0020] The method further comprises:
[0021] for any first sampling layer, determining a loss weight corresponding to the first sampling layer according to a comparison result between a classification result of output data of a reference sampling layer corresponding to the first sampling layer and a classification result of output data of the global sampling layer;
[0022] the updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers comprises:
[0023] updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers and the loss weights corresponding to the plurality of first sampling layers.
[0024] In a possible implementation, for any first sampling layer, if a classification result of output data of a reference sampling layer corresponding to the first sampling layer is different from a classification result of output data of the global sampling layer, a loss weight corresponding to the first sampling layer is determined according to a first preset weight; if the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is the same as the classification result of the output data of the global sampling layer, the loss weight corresponding to the first sampling layer is determined according to a second preset weight; the second preset weight is less than the first preset weight.
[0025] In a possible implementation, the method further includes:
[0026] obtaining label information of the sample data;
[0027] The updating, according to the difference between the first prediction result and the second prediction result, of part or all of the student model includes:
[0028] determining a prediction loss of the student model according to a difference between the first prediction result and the label information;
[0029] determining a response learning loss of the student model according to a difference between the first prediction result and the second prediction result;
[0030] updating part or all of the student model according to the prediction loss and the response learning loss.
[0031] In a possible implementation, before the updating, according to the prediction loss and the response learning loss, of part or all of the student model, the method further includes:
[0032] if the response learning loss is lower than a preset response loss threshold, obtaining a response learning decay parameter and a response learning weight;
[0033] updating the response learning weight according to the response learning decay parameter;
[0034] The updating, according to the prediction loss and the response learning loss, of part or all of the student model includes:
[0035] updating part or all of the student model according to a product of the response learning weight and the response learning loss and the prediction loss.
[0036] In a possible implementation, the response learning decay parameter is negatively correlated with a number of times of updating of the student model; and / or, the response learning decay parameter includes at least one of a weight coefficient, a weight reduction amount, and a weight decay probability.
[0037] In a possible implementation, the model is a student model.
[0038] The method further includes:
[0039] determining a feature learning loss of the student model according to a difference between output data of a first feature network in the student model and output data of a second feature network in a teacher model; the first feature network is configured to perform feature extraction processing on input data of the student model; and the second feature network is configured to perform feature extraction processing on input data of the teacher model.
[0040] The updating of the model according to the first prediction result includes:
[0041] The updating of the student model according to the first prediction result and the feature learning loss includes:
[0042] In a possible implementation, before the updating of the student model according to the first prediction result and the feature learning loss, the method further includes:
[0043] if the feature learning loss is lower than a preset feature loss threshold, obtaining a feature learning decay parameter and a feature learning weight;
[0044] updating the feature learning weight according to the feature learning decay parameter;
[0045] The updating of the student model according to the first prediction result and the feature learning loss includes:
[0046] updating part or all of the student model according to a product of the feature learning weight and the feature learning loss and the first prediction result.
[0047] In a possible implementation, the feature learning decay parameter is negatively correlated with a number of times of updating of the student model; and / or the feature learning decay parameter includes at least one of a weight coefficient, a weight reduction amount, and a weight decay probability.
[0048] In a possible implementation, the model is a student model; the model updating method is applied to a first training stage and / or a second training stage of the student model; the first training stage is implemented in an offline distillation manner; and the second training stage is implemented in a self-distillation manner, and a student model in the second training stage is initialized by using a student model obtained in the first training stage.
[0049] In a possible implementation, the offline distillation manner is implemented by using a prediction loss and a response learning loss of the student model in the first training stage, the prediction loss and the response learning loss being determined according to the first prediction result; and a feature learning loss of the student model, the feature learning loss being determined according to output data of a first feature network of the student model; and the self-distillation manner is implemented by using a response learning loss of the student model in the second training stage.
[0050] In a possible implementation, the sample data is an image, and the first prediction result is used to describe a predicted position of at least one contour point in the image.
[0051] The present application provides a data processing method, the method comprising:
[0052] obtaining to-be-processed data;
[0053] determining a prediction result corresponding to the to-be-processed data by using a model, the model being obtained by using a model updating method provided in the present application.
[0054] The present application provides a model updating device, comprising:
[0055] a first obtaining unit configured to obtain sample data;
[0056] a first prediction unit configured to determine a first prediction result corresponding to the sample data by using a model, the model comprising a first classification layer and a plurality of first sampling layers, sampling window sizes of different first sampling layers being different, the first classification layer being configured to obtain classification results of output data of each first sampling layer, and the first prediction result comprising the classification results of the output data of the plurality of first sampling layers;
[0057] a first updating unit configured to update part or all of the model according to the first prediction result.
[0058] The present application provides a data processing device, comprising:
[0059] a second obtaining unit configured to obtain to-be-processed data;
[0060] a second prediction unit configured to determine a prediction result corresponding to the to-be-processed data by using a model, the model being obtained by using a model updating method provided in the present application.
[0061] The present application provides an electronic device, the device comprising: a processor and a memory;
[0062] the memory is configured to store instructions or a computer program;
[0063] The processor is configured to execute the instructions or computer programs in the memory, so that the electronic device performs the model updating method or the data processing method provided in the present application.
[0064] The present application provides a computer readable medium, which stores instructions or computer programs, when the instructions or computer programs are run on a device, so that the device performs the model updating method or the data processing method provided in the present application.
[0065] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer readable medium, and the computer program includes program codes for executing the model updating method or the data processing method provided in the present application.
[0066] Compared with the related art, the present application has at least the following advantages:
[0067] In the technical solution provided in the present application, after obtaining the sample data, the model is used to determine the first prediction result corresponding to the sample data, such as the body contour point detection result, so that the first prediction result can represent the prediction information given by the model for the sample data, such as the prediction position information for describing the key points of the body contour; and then part or all of the model is updated according to the first prediction result, so that the updated model has better prediction performance. The model includes a first classification layer and a plurality of first sampling layers, and the first classification layer is used to obtain the classification results of the output data of each first sampling layer, so that the first prediction result includes the classification results of the output data of the plurality of first sampling layers. In addition, because the sampling window sizes of different first sampling layers are different, the output data of different first sampling layers can represent different scale information in the sample data, such as global information and local information at each scale, so that the classification results of the output data of these first sampling layers can better represent the prediction situation for the sample data, such as the overall prediction result, the local prediction result at each scale, and the like, and thus these classification results can better represent the performance of the model, so that the model updated based on these classification results has better performance, and thus better data processing effect can be achieved when the model is used for data processing. BRIEF DESCRIPTION OF DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments or the related art, the drawings needed in the embodiment or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.
[0069] Figure 1 A flowchart of a model updating method provided by an embodiment of the present application;
[0070] Figure 2 A schematic diagram of an offline distillation method provided by an embodiment of the present application;
[0071] Figure 3 A schematic diagram of a self-distillation method provided by an embodiment of the present application;
[0072] Figure 4 A schematic diagram of a multi-scale distillation method provided by an embodiment of the present application;
[0073] Figure 5 A flowchart of a data processing method provided by an embodiment of the present application;
[0074] Figure 6 A structural schematic diagram of a model updating device provided by an embodiment of the present application;
[0075] Figure 7 A structural schematic diagram of a data processing device provided by an embodiment of the present application;
[0076] Figure 8 A structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0077] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor fall within the scope of protection of the present application.
[0078] In order to better understand the technical solutions provided by the present application, the model updating method provided by the present application will be described below in combination with some drawings. As shown in Figure 1 The model updating method provided by the embodiment of the present application includes the following S101-S103. Among them, the model updating method provided by the embodiment of the present application includes the following S101-S103. Among them, Figure 1 A flowchart of a model updating method provided by an embodiment of the present application.
[0079] S101: Obtain sample data.
[0080] Among them, the sample data refers to the data that needs to be processed by the model in the model training process, such as Figure 2 The image 1 shown in Figure 3 The image 2 shown in Figure 4 The image 3 shown in.
[0081] In addition, the present application does not limit the implementation of the sample data, for example, the sample data can include one or more of images, texts, audios, and videos. For another example, the implementation of the sample data can be determined according to the actual application scenario. In order to facilitate understanding, some examples are described below.
[0082] Example 1, if the model updating method provided by the present application is applied to a certain image processing scene, such as face contour point detection processing, body contour point detection processing, target recognition scene, image segmentation scene, edge detection scene or text recognition processing, etc., the sample data can be implemented by using images.
[0083] Example 2, if the model updating method provided by the present application is applied to a certain text processing scene, such as information retrieval or text correction, etc., the sample data can be implemented by using texts.
[0084] Example 3, if the model updating method provided by the present application is applied to a certain audio processing scene, such as speech recognition processing or speech translation processing, etc., the sample data can be implemented by using audios.
[0085] Example 4, if the model updating method provided by the present application is applied to a certain video processing scene, such as video translation processing, foreground or background replacement processing in video, or posture adjustment processing in video, etc., the sample data can be implemented by using videos.
[0086] In addition, the present application does not limit the above-mentioned method of obtaining sample data, for example, it can be implemented by using any existing or future sample data acquisition method.
[0087] S102: determining a first prediction result corresponding to the sample data by using a model; the model includes a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain the classification results of the output data of each first sampling layer; and the first prediction result includes the classification results of the output data of the plurality of first sampling layers.
[0088] Among them, the first prediction result corresponding to the sample data refers to the prediction result given by the model in S102 for the sample data, such as contour point detection result, etc.
[0089] In addition, the present application does not limit the implementation of the first prediction result corresponding to the sample data. For example, the implementation of the first prediction result can be determined according to the actual application scenario. As an example, when the model updating method provided by the present application is applied to a body contour point detection processing scenario, and the sample data is an image, the first prediction result can be used to describe the predicted position of at least one contour point in the image, so that the first prediction result can represent the body contour point detection result given by the model in S102 for the sample data. Wherein, the at least one contour point refers to a key point used to describe the body contour; and the present application does not limit the implementation of the at least one contour point, for example, the at least one contour point can include some pre-set key points capable of describing the body contour. As can be seen, in one possible implementation, when the sample data is an image, the first prediction result can refer to the body contour point detection result given by the model in S102 for the sample data, so that the first prediction result can represent the predicted position of part or all of the key points used to describe the body contour in the image.
[0090] For example, the implementation of the first prediction result corresponding to the sample data can be determined according to the structural characteristics of the model in S102. Wherein, the model refers to the object to be updated, such as the student model shown in Figure 2 In addition, the present application does not limit the implementation of the model, for example, it can be implemented by using any existing or future machine learning model that needs to be updated. For example, the model can include at least a first classification layer and a plurality of first sampling layers, and the first classification layer is used to obtain the classification result of the output data of each first sampling layer. As can be seen, in one possible implementation, when the model includes a first classification layer and a plurality of first sampling layers, and the first classification layer is used to obtain the classification result of the output data of each first sampling layer, the first prediction result can include the classification result of the output data of the plurality of first sampling layers. For ease of understanding, the two network layers are introduced respectively as follows.
[0091] For the plurality of first sampling layers, the plurality of first sampling layers satisfy the following constraints: the input data of different first sampling layers is the same, but the sampling window size of different first sampling layers is different, so that the size of the output data of different first sampling layers is different, so that the amount of data in the output data of different first sampling layers is different, so that the output data of different first sampling layers respectively represents information in different scales extracted from the same input data, such as global information, local information in different scales, etc. For better understanding, the i-th first sampling layer is taken as an example for description, i is a positive integer, i≤I, I is a positive integer, I represents the number of sampling layers in the plurality of first sampling layers.
[0092] For the i-th first sampling layer, the i-th first sampling layer refers to a network layer existing in the model and used for sampling information according to an i-th sampling window size, so that the i-th first sampling layer is used to sample information at an i-th scale from input data of the i-th first sampling layer. The i-th sampling window size is used to represent the size of a sampling window required by the i-th first sampling layer when performing sampling processing. The application does not limit the manner of obtaining the i-th sampling window size. For example, the i-th sampling window size can be determined according to the scale requirement of output data of the i-th first sampling layer, so that the i-th first sampling layer can obtain output data meeting the requirement by means of the i-th sampling window size.
[0093] In addition, for the i-th scale in the above paragraph, the i-th scale is used to represent the size of output data of the i-th first sampling layer, so that the i-th scale can represent the number of data in the output data of the i-th first sampling layer, so that the i-th scale can represent the scale requirement set in advance for the output data of the i-th first sampling layer, and further, the i-th scale can represent the size of a receptive field of information sampled by the i-th first sampling layer to a certain extent. In addition, the application does not limit the implementation manner of the i-th scale. For example, the i-th scale can be implemented by 1x1, 2x2 or 3x3 as shown in the following table. Figure 4 It should be noted that, for the i-th scale in the above paragraph, the i-th scale is used to represent the size of output data of the i-th first sampling layer, so that the i-th scale can represent the number of data in the output data of the i-th first sampling layer, so that the i-th scale can represent the scale requirement set in advance for the output data of the i-th first sampling layer, and further, the i-th scale can represent the size of a receptive field of information sampled by the i-th first sampling layer to a certain extent. In addition, the application does not limit the implementation manner of the i-th scale. For example, the i-th scale can be implemented by 1x1, 2x2 or 3x3 as shown in the following table. Figure 4 It should be noted that, for the i-th scale in the above paragraph, the i-th scale is used to represent the size of output data of the i-th first sampling layer, so that the i-th scale can represent the number of data in the output data of the i-th first sampling layer, so that the i-th scale can represent the scale requirement set in advance for the output data of the i-th first sampling layer, and further, the i-th scale can represent the size of a receptive field of information sampled by the i-th first sampling layer to a certain extent. In addition, the application does not limit the implementation manner of the i-th scale. For example, the i-th scale can be implemented by 1x1, 2x2 or 3x3 as shown in the following table.
[0094] Furthermore, this application does not limit the implementation of the i-th first sampling layer. For example, the i-th first sampling layer can be implemented using any existing or future pooling layer, such as the average pooling layer at the i-th scale. As can be seen, in some scenarios, such as the scenario shown in Figure 4, if i = 1, the i-th first sampling layer can be implemented using a 1×1 average pooling layer, so that the i-th first sampling layer performs 1×1 average pooling on the input data of the i-th first sampling layer to obtain and output 1 data point; if i = 2, the i-th first sampling layer can be implemented using a 2×2 average pooling layer, so that the i-th first sampling layer performs 2×2 average pooling on the input data of the i-th first sampling layer to obtain and output 4 data points; if i = 4, the i-th first sampling layer can be implemented using a 4×4 average pooling layer, so that the i-th first sampling layer performs 4×4 average pooling on the input data of the i-th first sampling layer to obtain and output 16 data points.
[0095] For the first classification layer mentioned above, this first classification layer is used to classify the input data of the first classification layer. It can be seen that when the input data of the first classification layer includes the output data of the i-th first sampling layer, where i is a positive integer, i≤I, and I is a positive integer, the first classification layer can be used to classify the output data of the i-th first sampling layer to obtain and output the classification result of the output data of the i-th first sampling layer, so that the classification result can represent the classification of each data point in the output data of the i-th first sampling layer. It can be seen that, in one possible implementation, when the output data of the i-th first sampling layer includes J data points, such as through... Figure 4 The data obtained by the 1×1 average pooling process shown is obtained by... Figure 4 The four data points obtained from the 2×2 average pooling process shown, or through... Figure 4 When 16 data points are obtained from the 4×4 average pooling process shown, the classification result of the output data of the i-th first sampling layer can include the classification results of the J data points. The classification result of the j-th data point is used to represent the prediction information for the j-th data point in the output data of the i-th first sampling layer. This application does not limit the implementation of the classification result of the j-th data point. For example, when the j-th data point includes both horizontal and vertical coordinate information, the classification result of the j-th data point can include the classification results of the horizontal coordinate information and the classification results of the vertical coordinate information, so that the classification result of the j-th data point can better represent the prediction information given for the j-th data point, thereby improving the data processing effect. Here, j is a positive integer, j≤J, and J is a positive integer.
[0096] In addition, the present application does not limit the implementation of the first classification layer, for example, it can be implemented by using any existing or future classification layer, such as a fully connected layer. For example, in some scenarios, such as certain image processing scenarios or Figure 4 In order to better improve the effect, the first classification layer can include a horizontal coordinate classifier and a vertical coordinate classifier in some scenarios, such as certain image processing scenarios or
[0097] Based on the above-mentioned first classification layer and the plurality of first sampling layers, in some scenarios, the model in S102 can include a first classification layer and a plurality of first sampling layers. The first classification layer is used for classification processing of the output data of each first sampling layer. Different first sampling layers in the plurality of first sampling layers are used to sample information of different receptive fields, so that the information obtained by the plurality of first sampling layers is more comprehensive, thereby enabling the classification results of the output data of the first sampling layers to better represent the performance of the model.
[0098] In addition, the present application does not limit the model structure of the model in S102, for example, in some scenarios, the model can include a feature extraction network and a detection network. The feature extraction network is used for feature extraction processing of the input data of the feature extraction network, such as the above-mentioned sample data. The present application does not limit the implementation of the feature extraction network, for example, it can be implemented by using a backbone network. The detection network is used for processing the output data of the feature extraction network. The detection network includes a first classification layer and a plurality of first sampling layers. It can be seen that for the i-th first sampling layer in the detection network, the input data of the i-th first sampling layer is determined according to the output data of the feature extraction network. The present application does not limit the determination process of the input data of the i-th first sampling layer, for example, the input data of the i-th first sampling layer can be determined by using a module upstream of the i-th first sampling layer in the detection network, i is a positive integer, i≤I, I is a positive integer. In addition, the present application does not limit the implementation of the detection network, for example, it can be implemented by using a detection head including a first classification layer and a plurality of first sampling layers.
[0099] Based on the above content, in one possible implementation, when the model in S102 above includes a feature extraction network and a detection network, and the detection network includes a first classification layer and a plurality of first sampling layers, the determination of the first prediction result corresponding to the sample data can include: first, the feature extraction network performs feature extraction processing on the sample data to obtain a feature extraction result of the sample data, such as an image feature, so that the feature extraction result can better represent the information carried by the sample data, such as image information; and then the detection network performs some processing on the feature extraction result, such as sampling processing described by each first sampling layer and classification processing described by the first classification layer, to obtain the first prediction result corresponding to the sample data, so that the first prediction result includes classification results of output data of the plurality of first sampling layers, so that the first prediction result can better represent the prediction information given by the model for the sample data, such as detection results of body contour points and other information, and further make the first prediction result represent the performance of the model to some extent.
[0100] S103: Update part or all of the model according to the first prediction result.
[0101] It should be noted that the present application does not limit the implementation of S103 above, for example, in some scenarios, such as the true value guidance scenario, S103 can specifically include: updating part or all of the model according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data, such as updating all networks in the model or only updating the detection network in the model, so that the updated model has better performance, so that subsequent steps can continue to be executed based on the updated model to implement the next round of training for the model, and the cycle is iterated until the training process for the model is ended when the pre-set training stop condition is reached.
[0102] In addition, for the label information shown in the above paragraph, the label information refers to the guidance information pre-annotated for the sample data, so that the label information can represent the true value corresponding to the first prediction result. For example, when the sample data is an image and the first prediction result is used to describe the predicted position of at least one contour point in the image, the label information can be used to describe the actual position of the at least one contour point in the image; and the present application does not limit the acquisition method of the label information, for example, it can be implemented by manual annotation.
[0103] Further, for the above training stop condition, the training stop condition refers to a condition required to be satisfied when stopping the model training, and the application does not limit the implementation of the training stop condition, for example, the training stop condition can include that the model loss of the model is lower than a preset loss threshold. For another example, the training stop condition can include that the change rate of the model loss of the model is lower than a preset change rate threshold. For another example, the training stop condition can include that the number of model updates of the model reaches a preset number threshold. Wherein, the model loss of the model is used to represent the performance of the model, and the application does not limit the calculation manner of the model loss, for example, the model loss can be determined according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data.
[0104] Further, the application does not limit the implementation of the step of "updating part or all of the model", and the implementation of the step can be determined according to the actual application scene. For example, in some scenes, such as the scene where the performance of each network in the model needs to be optimized or the scene where the performance of one or more networks in the model needs to be optimized, the step can be specifically: updating all networks in the model, so that the performance of each network in the updated model is improved. Figure 2 For another example, in some scenes, such as the scene where the performance of each network in the model needs to be optimized or the scene where the performance of one or more networks in the model needs to be optimized, the step can be specifically: freezing part of the networks in the model, and updating other networks in the model except the part of the networks. Wherein, the part of the networks refers to the network in the model that does not need to be optimized in performance, such as the feature extraction network 3 shown in Figure 3 For another example, in some scenes, such as the scene where the performance of each network in the model needs to be optimized or the scene where the performance of one or more networks in the model needs to be optimized, the step can be specifically: freezing part of the networks in the model, and updating other networks in the model except the part of the networks. Wherein, the part of the networks refers to the network in the model that does not need to be optimized in performance, such as the feature extraction network 3 shown in Figure 3 For another example, in some scenes, such as the scene where the performance of each network in the model needs to be optimized or the scene where the performance of one or more networks in the model needs to be optimized, the step can be specifically: freezing part of the networks in the model, and updating other networks in the model except the part of the networks. Wherein, the part of the networks refers to the network in the model that does not need to be optimized in performance, such as the feature extraction network 3 shown in Figure 3 For another example, in some scenes, such as the scene where the performance of each network in the model needs to be optimized or the scene where the performance of one or more networks in the model needs to be optimized, the step can be specifically: freezing part of the networks in the model, and updating other networks in the model except the part of the networks. Wherein, the part of the networks refers to the network in the model that does not need to be optimized in performance, such as the feature extraction network 3 shown in
[0105] Based on the related content of S101 to S103 above, it can be known that for the model updating method provided by the embodiments of the present application, after the sample data is obtained, the first prediction result corresponding to the sample data is determined by using the model, such as the body contour point detection result, so that the first prediction result can represent the prediction information given by the model for the sample data, such as the prediction position information of the key points for describing the body contour; and then part or all of the model is updated according to the first prediction result, so that the updated model has better prediction performance. The model includes a first classification layer and a plurality of first sampling layers, and the first classification layer is used to obtain the classification result of the output data of each first sampling layer, so that the first prediction result includes the classification result of the output data of the plurality of first sampling layers. In addition, because the sampling window sizes of different first sampling layers are different, the output data of different first sampling layers can represent information of different scales in the sample data, such as global information and local information at each scale, so that the classification results of the output data of these first sampling layers can better represent the prediction situation for the sample data, such as the overall prediction result, the local prediction result at each scale, and the like, and thus these classification results can better represent the performance of the model, so that the model updated based on these classification results has better performance, thereby achieving better data processing effect when using the model for data processing.
[0106] In addition, the present application does not limit the execution subject of the model updating method provided by the embodiments of the present application. For example, the model updating method provided by the embodiments of the present application can be applied to a terminal device or a server. For another example, the model updating method provided by the embodiments of the present application can also be implemented by means of the data interaction process between the terminal device and the server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server or a cloud server.
[0107] In addition, in some scenarios, such as body contour point detection scenarios, the model training can be implemented by means of distillation technology. Based on this, the present application also provides a possible implementation of the above model updating method, in which the model updating method can include the following steps 11 to 14.
[0108] Step 11: Obtain sample data.
[0109] It should be noted that the related content of step 11 can be referred to the related content of S101 above.
[0110] Step 12: determining a first prediction result corresponding to the sample data by using a student model; the student model comprises a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is configured to obtain classification results of output data of each first sampling layer; and the first prediction result comprises the classification results of the output data of the plurality of first sampling layers.
[0111] It should be noted that the related content of step 12 is similar to the related content of S102, and the related content of step 12 can be obtained by replacing the “model” in the related content of S102 with “student model”.
[0112] Step 13: determining a second prediction result corresponding to the sample data by using a teacher model; the teacher model comprises a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of the corresponding first sampling layer of the second sampling layer in the plurality of first sampling layers; the second classification layer is configured to obtain classification results of output data of each second sampling layer; and the second prediction result comprises the classification results of the output data of the plurality of second sampling layers.
[0113] The second prediction result corresponding to the sample data refers to the prediction result given by the teacher model for the sample data, such as the contour point detection result. In addition, the present application does not limit the implementation of the second prediction result, for example, the implementation of the second prediction result is similar to the implementation of the first prediction result. Furthermore, the present application does not limit the relationship between the second prediction result and the first prediction result, for example, the relationship can include that the second prediction result and the first prediction result are prediction information given by different models for the same sample data. It can be seen that in a possible implementation, when the sample data is an image, the second prediction result can be used to describe the predicted position of at least one contour point in the image.
[0114] The teacher model is used to provide some guidance information, such as response guidance information, in the training process of the student model. It should be noted that the “response” appearing in the present application can refer to model output data.
[0115] In addition, the present application does not limit the association relationship between the teacher model and the student model, for example, the two models at least satisfy the following constraint: the number of parameters in the student model is less than the number of parameters in the teacher model.
[0116] In addition, the present application does not limit the implementation of the teacher model. For example, in order to better guide the first classification layer and the plurality of first sampling layers in the student model, the teacher model can at least include a second classification layer and a plurality of second sampling layers, so that in the subsequent training process, the second classification layer is used to guide the first classification layer, and the plurality of second sampling layers are used to guide the plurality of first sampling layers, so as to promote the first classification layer to learn the classification performance of the second classification layer, and promote each first sampling layer to learn the sampling performance of the corresponding second sampling layer.
[0117] In addition, in a possible implementation, there is a certain correspondence between the plurality of second sampling layers and the plurality of first sampling layers, and the present application does not limit the correspondence. For example, the correspondence satisfies the following constraint: for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of the corresponding first sampling layer in the plurality of first sampling layers. That is, for any pair of first sampling layer and second sampling layer with correspondence, the sampling window sizes of the two sampling layers are the same, which is conducive to promoting the first sampling layer to learn the sampling performance of the corresponding second sampling layer. For ease of understanding, the following is described with examples.
[0118] For example, when the student model includes I first sampling layers, and the teacher model includes I second sampling layers, if the i-th first sampling layer refers to the network layer existing in the student model for sampling information according to the i-th sampling window size, and the i-th second sampling layer refers to the network layer existing in the teacher model for sampling information according to the i-th sampling window size, then the sampling window size of the i-th first sampling layer and the sampling window size of the i-th second sampling layer are both the i-th sampling window size, so that the sampling window size of the i-th first sampling layer is the same as the sampling window size of the i-th second sampling layer, thereby the i-th first sampling layer and the i-th second sampling layer have a correspondence, so that in the subsequent training process, the i-th second sampling layer is used to guide the i-th first sampling layer, so as to promote the i-th first sampling layer to learn the sampling performance of the i-th second sampling layer. Wherein, i is a positive integer, i≤I, I is a positive integer. It should be noted that the implementation of the i-th second sampling layer is similar to that of the i-th first sampling layer, and for the sake of brevity, it will not be repeated here.
[0119] In addition, for the second classification layer in the teacher model, the second classification layer is configured to perform classification processing on input data of the second classification layer. It can be seen that when the input data of the second classification layer includes output data of the i-th second sampling layer, i is a positive integer, i≤I, and I is a positive integer, the second classification layer can be used to perform classification processing on the output data of the i-th second sampling layer to obtain and output the classification result of the output data of the i-th second sampling layer, so that the classification result can represent the classification of each data in the output data of the i-th second sampling layer. It should be noted that the implementation of the classification result of the output data of the i-th second sampling layer is similar to the implementation of the classification result of the output data of the i-th first sampling layer, and will not be described here for brevity.
[0120] Based on the above-mentioned related content of the second classification layer and the plurality of second sampling layers, it can be known that in some scenarios, the teacher model can include a second classification layer and a plurality of second sampling layers. Among them, the second classification layer is configured to perform classification processing on the output data of each second sampling layer. Different second sampling layers in the plurality of second sampling layers are configured to sample information of different receptive fields, so that the information obtained by means of the plurality of second sampling layers is more comprehensive, so that the classification result of the output data of these second sampling layers can better represent the performance of the model, so as to guide the classification result of the output data of the corresponding first sampling layer by the classification result of the output data of each second sampling layer in the subsequent training process, so as to realize the training process of the student model by means of multi-scale information such as global information and local information of various scales, and robust the various scale information of the student model, so as to effectively improve the accuracy of the student model.
[0121] In addition, the present application does not limit the implementation of the teacher model, for example, when the student model includes a first feature network and a first detection network, the teacher model can include a second feature network and a second detection network, so as to guide the first feature network by the second feature network and guide the first detection network by the second detection network in the subsequent training process. The following will introduce the four networks respectively.
[0122] For the above-mentioned first feature network, the first feature network refers to a feature extraction network in the student model, such as Figure 2 the feature extraction network 2 shown in Figure 3 the feature extraction network 3 shown in or Figure 4 the backbone network 2 shown in, so that the first feature network is configured to perform feature extraction processing on input data of the first feature network, such as the above-mentioned sample data.
[0123] For the above-mentioned first detection network, the first detection network refers to a detection network in the student model, such as Figure 2 the detection network 2 shown in orFigure 3 The first detection network 4 is shown to be configured to process the output data of the first feature network; and the first detection network comprises a first classification layer and a plurality of first sampling layers. The first classification layer and the plurality of first sampling layers are the same as those described above.
[0124] For the second feature network described above, the second feature network refers to a feature extraction network in the teacher model, such as Figure 2 The feature extraction network 1 shown in FIG. 1, Figure 3 The feature extraction network 3 shown in FIG. 3, or Figure 4 The backbone network 1 shown in FIG. 1, so that the second feature network is configured to perform feature extraction processing on the input data of the second feature network, such as the sample data described above, so as to guide the first feature network described above in the subsequent training process by using the second feature network, so as to promote the first feature network to learn the feature extraction performance of the second feature network.
[0125] For the second detection network described above, the second detection network refers to a detection network in the teacher model, such as Figure 2 The detection network 1 shown in FIG. 1, or Figure 3 The detection network 3 shown in FIG. 3, so that the second detection network is configured to process the output data of the second feature network, so as to guide the first detection network described above in the subsequent training process by using the second detection network, so as to promote the first detection network to learn the detection performance of the second detection network. In addition, the second detection network can comprise a second classification layer and a plurality of second sampling layers. It can be seen that for the i-th second sampling layer in the second detection network, the input data of the i-th second sampling layer is determined according to the output data of the second feature network; and the present application does not limit the determination process of the input data of the i-th second sampling layer, for example, the input data of the i-th second sampling layer can be determined by means of a module upstream of the i-th second sampling layer in the second detection network, i is a positive integer, i≤I, I is a positive integer. In addition, the present application does not limit the implementation of the detection network, for example, it can be implemented by using a detection head (Head) comprising a second classification layer and a plurality of second sampling layers.
[0126] Based on the related content of the above teacher model, in one possible implementation, when the teacher model includes a second feature network and a second detection network, and the second detection network includes a second classification layer and a plurality of second sampling layers, the determination process of the second prediction result corresponding to the above sample data can include: first, the second feature network is used to perform feature extraction processing on the sample data to obtain a feature extraction result of the sample data, such as an image feature, so that the feature extraction result can better represent the information carried by the sample data, such as image information; then, the second detection network is used to perform some processing on the feature extraction result, such as sampling processing described by each second sampling layer and classification processing described by the second classification layer, etc., to obtain the second prediction result corresponding to the sample data, so that the second prediction result includes the classification result of the output data of the plurality of second sampling layers, so that the second prediction result can represent the prediction information given by the teacher model for the sample data, such as the detection result of the body contour point, and further make the second prediction result can represent the performance of the teacher model to a certain extent.
[0127] In addition, the present application does not limit the relationship between the execution time of the above step 13 and the execution time of the above step 12, for example, both can be the same. For another example, the former is earlier than the latter. For another example, the latter is earlier than the former.
[0128] Step 14: updating part or all of the student model according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data.
[0129] The difference between the first prediction result and the second prediction result is used to represent the difference between the student model and the teacher model in response, so that the difference can represent the performance gap between the two models to a certain extent, so that the student model can be updated based on the difference, so that the performance of the updated student model is closer to the performance of the teacher model, so as to facilitate the realization of the data processing performance of the student model with less parameters simulating the teacher model with more parameters.
[0130] In addition, the present application does not limit the implementation of the above step 14, for example, it can be implemented by means of KL divergence (Kullback-Leibler Divergence) loss, such as the loss function shown in the following formula (1).
[0131]
[0132] In the formula, L response represents the loss of the student model and the teacher model in response; S l represents the response given by the student model to the lth sample data, such as the above first prediction result; Tl represents a response given by the teacher model for the l-th sample data, as the second prediction result above, l is a positive integer, l≤L, L, K and N are preset values, and the present application does not limit the implementation of L, K and N.
[0133] It can be seen that in a possible implementation, step 14 above can be specifically: first, according to the difference between the first prediction result and the second prediction result, determine the model loss of the student model, such as the loss shown in formula (1) above; then update part or all of the student model according to the loss, such as updating all networks in the student model or only updating the detection network in the student model, so that the performance of the updated student model is closer to the performance of the teacher model, so that subsequent can continue to execute step 11 above and subsequent steps based on the updated student model, to realize the next round of training for the student model, and so on. Iteration, until the training process for the student model is ended when the preset training stopping condition is reached.
[0134] In fact, in order to better improve the performance of the model, the present application also provides a possible implementation of step 14 above, in which the step 14 can specifically include steps 141-142 below.
[0135] Step 141: for any first sampling layer, according to the difference between the classification result of the output data of the first sampling layer and the classification result of the output data of the reference sampling layer corresponding to the first sampling layer, determine the loss corresponding to the first sampling layer; the plurality of second sampling layers include the reference sampling layer, and the sampling window size of the reference sampling layer is the same as the sampling window size of the first sampling layer.
[0136] Wherein, the reference sampling layer corresponding to the i-th first sampling layer refers to the second sampling layer existing in the teacher model and having a corresponding relationship with the i-th first sampling layer, such as the i-th second sampling layer above, so that the reference sampling layer and the i-th first sampling layer satisfy the following constraint: the sampling window size of the reference sampling layer is the same as the sampling window size of the i-th first sampling layer. It can be seen that the reference sampling layer is used to guide the i-th first sampling layer, so as to make the i-th first sampling layer learn the sampling performance of the reference sampling layer.
[0137] In addition, for the reference sampling layer corresponding to the i th first sampling layer, the classification result of the output data of the reference sampling layer refers to a result obtained by performing classification processing on the output data of the reference sampling layer by the second classification layer in the teacher model, so that the classification result can represent the prediction information determined according to the output data of the reference sampling layer, such as the body contour point detection result and the like. It should be noted that the present application does not limit the implementation of the "classification result of the output data of the reference sampling layer", for example, when the output data of the reference sampling layer includes at least one data, the "classification result of the output data of the reference sampling layer" can include the classification result of the at least one data, so that the classification result of each data in the "classification result of the output data of the reference sampling layer" is respectively used to guide the classification result of the corresponding data in the output data of the i th first sampling layer, which is beneficial to robust the scale information of the student model, thereby improving the prediction accuracy of the student model.
[0138] In addition, for the loss corresponding to the i th first sampling layer, the loss is used to represent the gap between the sampling performance of the i th first sampling layer and the sampling performance of the reference sampling layer corresponding to the i th first sampling layer, and the loss is determined according to the difference between the classification result of the output data of the i th first sampling layer and the classification result of the output data of the reference sampling layer. It should be noted that the present application does not limit the implementation of the loss, for example, it can be implemented by using the KL divergence loss.
[0139] Based on the related content of the above step 141, for the i th first sampling layer in the student model, the reference sampling layer corresponding to the i th first sampling layer is first determined from the teacher model, so that the sampling window size of the reference sampling layer is the same as the sampling window size of the i th first sampling layer; then, the difference between the classification result of the output data of the i th first sampling layer and the classification result of the output data of the reference sampling layer is calculated; then, according to the difference, the loss corresponding to the first sampling layer is determined, so that the loss can represent the performance gap between the i th first sampling layer and the reference sampling layer. Wherein, i is a positive integer, i≤I, I is a positive integer.
[0140] Step 142: updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers.
[0141] It should be noted that the present application does not limit the implementation of the step 142, for example, the step 142 can be specifically: adding the losses corresponding to the plurality of first sampling layers to obtain the model loss of the student model; then, updating part or all of the student model according to the model loss.
[0142] Based on the related content of steps 141-142, in some scenarios, when the student model includes multiple first sampling layers and the teacher model includes multiple second sampling layers, the performance gap between each first sampling layer and its corresponding second sampling layer can be calculated first; then, based on these performance gaps, part or all of the student model is updated, such as updating all networks in the student model or only updating the first detection network in the student model, so that the updated student model has better performance, such as sampling performance of each scale information, so that subsequent can continue to perform step 11 and subsequent steps based on the updated student model to implement the next round of training for the student model, and so on until the training process for the student model is ended when the pre-set training stop condition is reached.
[0143] In fact, in order to better improve the training effect, the present application also provides a possible implementation of step 14, in which when the teacher model includes multiple second sampling layers and the multiple second sampling layers include a global sampling layer, step 14 can specifically include steps 143-145.
[0144] Step 143: For any first sampling layer, the loss corresponding to the first sampling layer is determined according to the difference between the classification result of the output data of the first sampling layer and the classification result of the output data of the reference sampling layer corresponding to the first sampling layer; the multiple second sampling layers include the reference sampling layer, and the sampling window size of the reference sampling layer is the same as the sampling window size of the first sampling layer.
[0145] It should be noted that the related content of step 143 can be referred to the related content of step 141.
[0146] Step 144: For any first sampling layer, the loss weight corresponding to the first sampling layer is determined according to the comparison result between the classification result of the output data of the reference sampling layer corresponding to the first sampling layer and the classification result of the output data of the global sampling layer.
[0147] The global sampling layer refers to the second sampling layer in the teacher model that is used to obtain global information, such as the sampling layer used to implement the 1x1 average pooling processing as shown in the figure. It can be seen that in a possible implementation, the size of the output data of the global sampling layer is 1x1. Figure 4
[0148] In addition, for the reference sampling layer corresponding to the i-th first sampling layer, such as the i-th second sampling layer, the classification result of the output data of the reference sampling layer can be compared with the classification result of the output data of the global sampling layer to obtain a comparison result, so that the comparison result can represent whether the two classification results are consistent, thereby enabling the comparison result to represent whether there is a large gap between the local information described by the output data of the reference sampling layer and the global information described by the output data of the global sampling layer, and further enabling the comparison result to represent whether the local information can supplement the global information. Wherein, i is a positive integer, i≤I, I is a positive integer.
[0149] In addition, the present application does not limit the implementation of step 144 above, for example, it can be specifically: for any first sampling layer, if the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is different from the classification result of the output data of the global sampling layer, then the loss weight corresponding to the first sampling layer is determined according to the first preset weight, such as the weight 2; if the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is the same as the classification result of the output data of the global sampling layer, then the loss weight corresponding to the first sampling layer is determined according to the second preset weight, such as the weight 1. Wherein, the second preset weight is less than the first preset weight.
[0150] It can be seen that in a possible implementation, for the i-th first sampling layer, after obtaining the classification result of the output data of the reference sampling layer corresponding to the i-th first sampling layer, it can be determined whether the classification result of the output data of the reference sampling layer is the same as the classification result of the output data of the global sampling layer, if the same, it can be determined that part or all of the local information described by the output data of the reference sampling layer is consistent with the global information described by the output data of the global sampling layer, thereby it can be determined that the local information cannot supplement the global information, so the second preset weight, such as the weight 1, can be directly determined as the loss weight corresponding to the i-th first sampling layer, to reduce the impact of the loss corresponding to the i-th first sampling layer on the updating process of the student model; but if different, it can be determined that the local information does not exist in the global information, thereby it can be determined that the local information can supplement the global information, so the first preset weight, such as the weight 2, can be directly determined as the loss weight corresponding to the first sampling layer, to increase the impact of the loss corresponding to the i-th first sampling layer on the updating process of the student model, which is conducive to supplementing some local information provided by the teacher model to the student model, to perfect the local information provided by the student model, thereby facilitating to improve the performance of the student model.
[0151] In addition, the present application does not limit the relationship between the execution time of the above step 144 and the execution time of the above step 143, for example, both can be the same. For another example, the former is earlier than the latter. For another example, the latter is earlier than the former.
[0152] Step 145: update part or all of the student model according to the losses corresponding to the plurality of first sampling layers and the loss weights corresponding to the plurality of first sampling layers.
[0153] It should be noted that the present application does not limit the implementation of the above step 145, for example, it can be specifically: first, the loss corresponding to the plurality of first sampling layers is weighted and summed according to the loss weight corresponding to the plurality of first sampling layers, to obtain the model loss of the student model; then, part or all of the student model is updated according to the model loss.
[0154] Based on the related content of the above steps 143 to 145, in some scenarios, when the student model includes a plurality of first sampling layers and the teacher model includes a plurality of second sampling layers, the performance gap and its weight between each first sampling layer and its corresponding second sampling layer can be calculated first; then, part or all of the student model is updated according to the performance gap and its weight, such as updating all networks in the student model or only updating the first detection network in the student model, so that the updated student model has better performance, such as the sampling performance of each scale information, so that subsequent can continue to execute the above step 11 and its subsequent steps based on the updated student model, to realize the next round of training for the student model, and so on. Iteration until the training process for the student model ends when the pre-set training stop condition is reached.
[0155] Based on the related content of the above steps 11 to 14, in some scenarios, when the student model includes a plurality of first sampling layers and the teacher model includes a plurality of second sampling layers, the sample data is first processed using the student model to obtain the classification result of the output data of the plurality of first sampling layers, and the sample data is processed using the teacher model to obtain the classification result of the output data of the plurality of second sampling layers; then, the classification result of the output data of the corresponding first sampling layer is analyzed with reference to the classification result of the output data of each second sampling layer, to obtain the loss and its weight corresponding to the corresponding first sampling layer; finally, part or all of the student model is updated according to the loss and its weight, such as updating all networks in the student model or only updating the first detection network in the student model, so that the updated student model has better performance, such as the sampling performance of each scale information, so that subsequent can continue to execute the next round of training for the student model based on the updated student model, and so on. Iteration until the training process for the student model ends when the pre-set training stop condition is reached.
[0156] In fact, in some scenarios, such as offline distillation scenarios, in order to better improve the performance of the model, not only the degree of simulation of the response of the student model to the teacher model needs to be considered, but also the prediction performance of the student model itself, such as the body contour point detection performance, needs to be considered. Based on this, the present application also provides a possible implementation of the above model updating method, in which the model updating method can include the following steps 21-26.
[0157] Step 21: obtaining sample data and label information of the sample data.
[0158] It should be noted that the related content of the sample data in step 21 can be referred to the related content of S101 above, and the related content of the label information in step 21 can be referred to the related content of S103 above.
[0159] Step 22: determining a first prediction result corresponding to the sample data by using a student model; the student model includes a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain classification results of output data of each first sampling layer; and the first prediction result includes classification results of output data of the plurality of first sampling layers.
[0160] It should be noted that the related content of step 22 can be referred to the related content of step 12 above.
[0161] Step 23: determining a second prediction result corresponding to the sample data by using a teacher model; the teacher model includes a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of the corresponding first sampling layer of the second sampling layer in the plurality of first sampling layers above; the second classification layer is used to obtain classification results of output data of each second sampling layer; and the second prediction result includes classification results of output data of the plurality of second sampling layers.
[0162] It should be noted that the related content of step 23 can be referred to the related content of step 13 above.
[0163] It should be further noted that the present application does not limit the relationship between the execution time of step 23 above and the execution time of step 22 above, for example, the former can be the same as the latter. For another example, the former is earlier than the latter. For another example, the latter is earlier than the former.
[0164] Step 24: determining a prediction loss of the student model according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data.
[0165] The prediction loss of the student model is used to represent the prediction performance of the student model, and the prediction loss is determined according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data. It should be noted that the application does not limit the calculation method of the prediction loss, for example, it can be implemented by using any method that can calculate the loss between the prediction result and the ground truth.
[0166] Step 25: determining the response learning loss of the student model according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data.
[0167] The response learning loss of the student model is used to represent the difference between the student model and the teacher model in response, so that the response learning loss can represent the learning situation of the response of the student model to the teacher model, such as the simulation degree.
[0168] In addition, the response learning loss of the student model is determined according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data; and the response learning loss can be implemented by using the loss calculation method involved in step 14 above. In order to facilitate understanding, some examples are described below.
[0169] Example 1, in some scenarios, the above step 25 can include the following steps 251-252.
[0170] Step 251: for any first sampling layer, determining the loss corresponding to the first sampling layer according to the difference between the classification result of the output data of the first sampling layer and the classification result of the output data of the reference sampling layer corresponding to the first sampling layer; the plurality of second sampling layers include the reference sampling layer, and the sampling window size of the reference sampling layer is the same as the sampling window size of the first sampling layer.
[0171] It should be noted that the related content of step 251 can be referred to the related content of step 141 above.
[0172] Step 252: determining the response learning loss of the student model according to the losses corresponding to the plurality of first sampling layers.
[0173] It should be noted that the application does not limit the implementation of the step 252, for example, the step 252 can be specifically: summing the losses corresponding to the plurality of first sampling layers to obtain the response learning loss of the student model.
[0174] Based on the related content of steps 251-252, in some scenarios, when the student model includes a plurality of first sampling layers and the teacher model includes a plurality of second sampling layers, the performance gap between each first sampling layer and its corresponding second sampling layer can be calculated first; then, the response learning loss of the student model is determined according to the performance gap, so that the response learning loss can represent the simulation degree of the response of the student model to the teacher model.
[0175] In example 2, in some scenarios, when the teacher model includes a plurality of second sampling layers, and the plurality of second sampling layers includes a global sampling layer, step 25 can include steps 253-255.
[0176] Step 253: For any first sampling layer, determine the loss corresponding to the first sampling layer according to the difference between the classification result of the output data of the first sampling layer and the classification result of the output data of the reference sampling layer corresponding to the first sampling layer; the plurality of second sampling layers include the reference sampling layer, and the sampling window size of the reference sampling layer is the same as the sampling window size of the first sampling layer.
[0177] It should be noted that the related content of step 253 can be referred to the related content of step 141.
[0178] Step 254: For any first sampling layer, determine the loss weight corresponding to the first sampling layer according to the comparison result between the classification result of the output data of the reference sampling layer corresponding to the first sampling layer and the classification result of the output data of the global sampling layer.
[0179] It should be noted that the related content of step 254 can be referred to the related content of step 144.
[0180] It should be further noted that the present application does not limit the relationship between the execution time of step 254 and the execution time of step 253, for example, they can be the same. For another example, the former is earlier than the latter. For another example, the latter is earlier than the former.
[0181] Step 255: Determine the response learning loss of the student model according to the loss corresponding to the plurality of first sampling layers and the loss weight corresponding to the plurality of first sampling layers.
[0182] It should be noted that the present application does not limit the implementation of the above step 255, for example, it can be specifically: according to the loss weight corresponding to the plurality of first sampling layers, the loss corresponding to the plurality of first sampling layers is weighted and summed to obtain the response learning loss of the student model. As can be seen, in one possible implementation, the response learning loss = the loss weight corresponding to the first first sampling layer x the loss corresponding to the first first sampling layer + the loss weight corresponding to the second first sampling layer x the loss corresponding to the second first sampling layer + … (and so on) + the loss weight corresponding to the Ith first sampling layer x the loss corresponding to the Ith first sampling layer.
[0183] Based on the above steps 253 to 255, in some scenarios, when the student model includes a plurality of first sampling layers, and the teacher model includes a plurality of second sampling layers, the performance gap between each first sampling layer and its corresponding second sampling layer and its weight can be calculated first; then according to these performance gaps and their weights, the response learning loss of the student model is determined, so that the response learning loss can more accurately represent the simulation degree of the response of the student model to the teacher model.
[0184] In addition, the present application does not limit the relationship between the execution time of the above step 25 and the execution time of the above step 24, for example, both can be the same. For example, the former is earlier than the latter. For example, the latter is earlier than the former.
[0185] Step 26: updating part or all of the student model according to the prediction loss of the student model and the response learning loss of the student model.
[0186] It should be noted that the present application does not limit the implementation of the above step 26, for example, the step 26 can be specifically: first, the prediction loss of the student model and the response learning loss of the student model are added to obtain the model loss of the student model; then, according to the model loss, part or all of the student model is updated.
[0187] For example, in order to better improve the training effect, the above step 26 can be specifically: first, the prediction loss of the student model and the response learning loss of the student model are weighted and summed to obtain the model loss of the student model; then, according to the model loss, part or all of the student model is updated. It should be noted that the present application does not limit the weight corresponding to the prediction loss, for example, the weight corresponding to the prediction loss can be 1. In addition, the present application also does not limit the weight corresponding to the response learning loss, for example, the weight corresponding to the response learning loss can be β. The β can be set according to the actual application scenario.
[0188] Based on the related content of steps 21-26 above, in some scenarios, such as offline distillation scenarios, not only the degree of simulation of the response of the student model to the teacher model needs to be calculated, but also the gap between the output data of the student model and the true label needs to be calculated, so that based on these two kinds of information, part or all of the student model can be updated, such as updating all networks in the student model or only updating the first detection network in the student model, so that the updated student model has better performance, such as the sampling performance of each scale information, so that the next round of training of the student model can continue to be performed based on the updated student model, and the training process of the student model is ended when the pre-set training stopping condition is reached.
[0189] In fact, in some scenarios, such as offline distillation scenarios, in order to better improve the performance of the model, the learning of the response of the student model to the teacher model can be gradually weakened as the performance of the student model is continuously improved, so that the student model can focus more on improving its own prediction performance. Based on this, the present application also provides a possible implementation of the above model updating method, in which the model updating method can include steps 31-39 below.
[0190] Step 31: Obtain sample data and label information of the sample data.
[0191] It should be noted that the related content of step 31 can be referred to the related content of step 21 above.
[0192] Step 32: Determine the first prediction result corresponding to the sample data by using the student model; the student model includes a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain the classification results of the output data of each first sampling layer; and the first prediction result includes the classification results of the output data of a plurality of first sampling layers.
[0193] It should be noted that the related content of step 32 can be referred to the related content of step 12 above.
[0194] Step 33: Determine the second prediction result corresponding to the sample data by using the teacher model; the teacher model includes a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of the corresponding first sampling layer of the second sampling layer in the plurality of first sampling layers above; the second classification layer is used to obtain the classification results of the output data of each second sampling layer; and the second prediction result includes the classification results of the output data of a plurality of second sampling layers.
[0195] It should be noted that the related content of step 33 can be referred to the related content of step 13 above.
[0196] It should be further noted that the present application does not limit the relationship between the execution time of the above step 33 and the execution time of the above step 32, such as both can be the same. For example, the former is earlier than the latter. Also, for example, the latter is earlier than the former.
[0197] Step 34: According to the difference between the first prediction result corresponding to the sample data and the label information of the sample data, determine the prediction loss of the student model.
[0198] It should be noted that the related content of step 34 can refer to the related content of the above step 24.
[0199] Step 35: According to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data, determine the response learning loss of the student model.
[0200] It should be noted that the related content of step 35 can refer to the related content of the above step 25.
[0201] Step 36: Determine whether the response learning loss of the student model is lower than the preset response loss threshold, if yes, then execute the following steps 37-39 in turn; if not, then execute the following step 39.
[0202] Among them, the preset response loss threshold refers to a threshold value preset for measuring whether to weaken the learning of the response of the student model to the teacher model; and the preset response loss threshold can be determined according to the model update requirement in the actual application scene.
[0203] Based on the related content of the above step 36, after obtaining the response learning loss of the student model, it is determined whether the response learning loss is lower than the preset response loss threshold, if it is lower, it can be determined that the student model can better simulate the response of the teacher model, therefore, in order to better improve the model performance, the response learning loss can be gradually weakened by reducing the corresponding weighted weight, such as the weight decay method shown in the following steps 37-38, to make the student model more focused on improving its own prediction performance; however, if it is not lower, it can be determined that the gap between the student model and the teacher model in response is still relatively large, so it is necessary to continue learning the response of the teacher model according to the previous intensity, such as the intensity represented by the initial value of the preset response learning weight, so there is no need to adjust the weighted weight corresponding to the response learning loss.
[0204] Step 37: Obtain the response learning decay parameter and the response learning weight.
[0205] The response learning weight refers to a weighting weight corresponding to the response learning loss of the student model in the current round, such as the above β, so that the response learning weight can represent the influence degree of the response learning loss on the updating process of the student model in the current round. In addition, the initial value of the response learning weight can be determined according to the model updating requirement in the actual application scenario.
[0206] The response learning decay parameter refers to a parameter required for weakening the learning of the student model in response to the teacher model in the current round.
[0207] In addition, the present application does not limit the implementation of the response learning decay parameter, for example, the response learning decay parameter can include at least one of a weight coefficient, a weight reduction amount, and a weight decay probability. The weight coefficient is used to represent the decay degree of the weighting weight corresponding to the response learning loss of the student model. The weight reduction amount is used to represent the decay amount of the weighting weight corresponding to the response learning loss. The weight decay probability is used to represent the possibility of the decay processing of the weighting weight corresponding to the response learning loss.
[0208] In addition, in order to better improve the model performance, the present application also provides a possible implementation of the above response learning decay parameter, in which the response learning decay parameter satisfies the following constraint: the response learning decay parameter is negatively correlated with the updating times of the student model. It can be seen that when the response learning decay parameter includes the weight coefficient, the weight coefficient is negatively correlated with the updating times of the student model, that is, as the updating times of the student model gradually increase, the weight coefficient gradually decreases, so that the decay degree of the influence degree of the response learning loss of the student model on the updating process of the student model gradually decreases; when the response learning decay parameter includes the weight reduction amount, the weight reduction amount is negatively correlated with the updating times of the student model, that is, as the updating times of the student model gradually increase, the weight reduction amount gradually decreases, so that the decay amount of the influence degree of the response learning loss of the student model on the updating process of the student model gradually decreases. When the response learning decay parameter includes the weight decay probability, the weight decay probability is negatively correlated with the updating times of the student model, that is, as the updating times of the student model gradually increase, the weight decay probability gradually decreases, so that the occurrence probability of the decay processing of the influence degree of the response learning loss of the student model on the updating process of the student model gradually decreases.
[0209] Step 38: updating the response learning weight according to the response learning decay parameter.
[0210] It should be noted that the present application does not limit the implementation of step 38, for example, when the response learning decay parameter above includes a weight coefficient, step 38 can be specifically: multiplying the weight coefficient and the response learning weight before updating to obtain the response learning weight after updating, so that the response learning weight after updating is less than the response learning weight before updating.
[0211] For another example, when the response learning decay parameter above includes a weight reduction amount, step 38 above can be specifically: subtracting the weight reduction amount from the response learning weight before updating to obtain the response learning weight after updating, so that the response learning weight after updating is less than the response learning weight before updating.
[0212] For another example, when the response learning decay parameter above includes a weight decay probability, step 38 above can be specifically: performing decay processing on the response learning weight before updating according to the weight decay probability to obtain the response learning weight after updating.
[0213] Step 39: updating part or all of the student model according to the product of the response learning weight and the response learning loss of the student model, and the prediction loss of the student model.
[0214] It should be noted that the present application does not limit the implementation of step 39, for example, it can be specifically: first calculating the product of the response learning weight and the response learning loss of the student model to obtain the weighted loss; then adding the weighted loss and the prediction loss of the student model to obtain the model loss of the student model; then, updating part or all of the student model according to the model loss.
[0215] Based on the related content of steps 31 to 39 above, for some scenarios, such as offline distillation scenarios, in the early stage of training of the learning model, the learning model can be prompted to implement response learning for the teacher model with a larger weight, so that the learning model can learn the response performance of the teacher model as quickly as possible. However, in the later stage of training of the learning model, because the response performance of the learning model is close to the response performance of the teacher model, in order to better improve the model performance and the model training efficiency, the learning model can be prompted to implement response learning for the teacher model with a smaller weight, so that the learning model can focus on improving its own prediction performance, thereby making the finally trained student model have better prediction performance, which is conducive to improving the model performance.
[0216] In fact, in some scenarios, such as offline distillation scenarios, in order to better improve the performance of the model, the learning model can be further prompted to learn some intermediate data of the teacher model, such as feature extraction results and the like. Based on this, the present application also provides a possible implementation manner of the above model updating method, in which the model updating method can include the following steps 41-44.
[0217] Step 41: obtaining sample data.
[0218] It should be noted that the related content of step 51 can be referred to the related content of S101 above.
[0219] Step 42: determining a first prediction result corresponding to the sample data by using a student model; the student model includes a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain the classification results of the output data of each first sampling layer; and the first prediction result includes the classification results of the output data of the plurality of first sampling layers.
[0220] It should be noted that the related content of step 42 can be referred to the related content of step 12 above.
[0221] Step 43: determining a feature learning loss of the student model according to the difference between the output data of a first feature network in the student model and the output data of a second feature network in the teacher model; the first feature network is used for feature extraction processing on the input data of the student model; and the second feature network is used for feature extraction processing on the input data of the teacher model.
[0222] The feature learning loss of the student model is used to represent the gap between the student model and the teacher model in feature extraction, so that the learning loss can represent the learning situation of the student model on the feature extraction of the teacher model, such as the simulation degree.
[0223] In addition, the feature learning loss of the student model is determined according to the difference between the output data of the first feature network in the student model and the output data of the second feature network in the teacher model; and the present application does not limit the determination manner of the feature learning loss, for example, it can be implemented by means of a mean squared error (MSE) loss, such as a loss function shown in the following formula (2).
[0224]
[0225] In the formula, L feature represents the feature learning loss of the student model; F t represents the output data of the second feature network in the teacher model, and the size of F t is CxHxW.s represents output data of the first feature network in the student model, and the F s has a size of CxHxW; represents the F t data in the F represents the F s data in the F is a 1x1 convolution, so that the f(·) is used to transform the dimension of the F to make the dimension of the F
[0226] Step 44: updating part or all of the student model according to the first prediction result corresponding to the sample data and the feature learning loss of the student model.
[0227] It should be noted that the present application does not limit the implementation of step 44 above, and the following examples are described for the convenience of understanding.
[0228] Example 1: in some scenarios, step 44 above can specifically include: first determining the prediction loss of the student model according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data; and then updating part or all of the student model according to the prediction loss and the feature learning loss of the student model. It should be noted that the related content of the label information and the prediction loss is described above. In addition, the present application does not limit the implementation of the updating process, for example, it can specifically be: first, the prediction loss and the feature learning loss are weighted and summed to obtain the model loss of the student model; and then, part or all of the student model is updated according to the model loss. Wherein, the weighted weight corresponding to the prediction loss can be 1, and the weighted weight corresponding to the feature learning loss can be a. In addition, the a can be set according to the actual application scenario.
[0229] Example 2: in some scenarios, step 44 above can specifically include: first, determining the response learning loss of the student model according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data; and then, updating part or all of the student model according to the response learning loss and the feature learning loss of the student model. It should be noted that the related content of the second prediction result and the response learning loss is described above. In addition, the present application does not limit the implementation of the updating process, for example, it can specifically be: first, the response learning loss and the feature learning loss are weighted and summed to obtain the model loss of the student model; and then, part or all of the student model is updated according to the model loss. Wherein, the weighted weight corresponding to the response learning loss can be β, and the weighted weight corresponding to the feature learning loss can be a.
[0230] In some scenarios, step 44 can specifically include: first determining the prediction loss of the student model according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data, and determining the response learning loss of the student model according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data; and then updating part or all of the student model according to the prediction loss, the response learning loss and the feature learning loss of the student model. It should be noted that the application does not limit the implementation of the updating process. For example, it can specifically be: first, the prediction loss, the response learning loss and the feature learning loss are weighted and summed to obtain the model loss of the student model; and then, part or all of the student model is updated according to the model loss. The weighted weight corresponding to the prediction loss can be 1, the weighted weight corresponding to the response learning loss can be β, and the weighted weight corresponding to the feature learning loss can be α.
[0231] Based on the related content of steps 41 to 44, for some scenarios, such as offline distillation scenarios, not only the first prediction result corresponding to the sample data is needed to determine the model loss of the student model, but also the simulation degree of the response of the student model to the teacher model is further used to determine the model loss of the student model, so that the model loss can better represent the performance of the student model in the current round, so that the student model updated based on the model loss has better performance, such as prediction performance, feature extraction performance, and sampling performance of each scale information, so that the next round of training of the student model can be continued based on the updated student model, and the cycle is iterated until the training process of the student model is ended when the pre-set training stop condition is reached.
[0232] In fact, in some scenarios, such as offline distillation scenarios, in order to better improve the model performance, the learning of the intermediate data of the teacher model by the student model can be gradually weakened as the performance of the student model is continuously improved. Based on this, the application also provides a possible implementation of the above model updating method, in which the model updating method can include steps 51-57 below.
[0233] Step 51: obtaining sample data.
[0234] It should be noted that the related content of step 51 can be referred to the related content of S101 above.
[0235] Step 52: determining a first prediction result corresponding to the sample data by using the student model; the student model comprises a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain classification results of output data of the first sampling layers; and the first prediction result comprises the classification results of the output data of the first sampling layers.
[0236] It should be noted that the related content of step 52 can be referred to the related content of step 12 in the foregoing.
[0237] Step 53: determining a feature learning loss of the student model according to a difference between output data of a first feature network in the student model and output data of a second feature network in the teacher model; the first feature network is used to perform feature extraction processing on input data of the student model; and the second feature network is used to perform feature extraction processing on input data of the teacher model.
[0238] It should be noted that the related content of step 53 can be referred to the related content of step 43 in the foregoing.
[0239] Step 54: determining whether the feature learning loss of the student model is lower than a preset feature loss threshold, if yes, sequentially executing steps 55-57 in the following; and if no, executing step 57 in the following.
[0240] The preset feature loss threshold refers to a threshold that is preset and used to measure whether to weaken the learning of the student model on the feature extraction result of the teacher model; and the preset feature loss threshold can be determined according to the model updating requirement in the actual application scenario.
[0241] Based on the related content of step 54 in the foregoing, after obtaining the feature learning loss of the student model in the foregoing, it is determined whether the feature learning loss is lower than the preset feature loss threshold, if yes, it can be determined that the student model can better simulate the feature extraction result of the teacher model, and therefore, in order to better improve the model performance, the learning of the student model on the feature extraction result of the teacher model can be gradually weakened by means of reducing the weighting weight corresponding to the feature learning loss, such as the weight decay manner shown in steps 55-56 in the following; but if no, it can be determined that the gap between the student model and the teacher model in the feature extraction result is still relatively large, and therefore, the learning of the feature extraction result of the teacher model needs to be continued according to the previous intensity, such as the intensity represented by the initial value of the preset feature learning weight, and therefore, the weighting weight corresponding to the feature learning loss does not need to be adjusted.
[0242] Step 55: obtaining a feature learning decay parameter and a feature learning weight.
[0243] The feature learning weight refers to a weighting weight corresponding to the feature learning loss of the student model in the current round, such as the above a, so that the feature learning weight can represent the influence degree of the feature learning loss on the updating process of the student model in the current round. In addition, the initial value of the feature learning weight can be determined according to the model updating requirement in the actual application scene.
[0244] The feature learning attenuation parameter refers to a parameter required for weakening the learning of the student model to the feature extraction result of the teacher model in the current round.
[0245] In addition, the present application does not limit the implementation of the feature learning attenuation parameter, for example, the feature learning attenuation parameter can include at least one of a weight coefficient, a weight reduction amount, and a weight attenuation probability. The weight coefficient is used to represent the attenuation degree of the weighting weight corresponding to the feature learning loss of the student model. The weight reduction amount is used to represent the attenuation amount of the weighting weight corresponding to the feature learning loss. The weight attenuation probability is used to represent the possibility of the attenuation processing of the weighting weight corresponding to the feature learning loss.
[0246] In addition, in order to better improve the model performance, the present application also provides a possible implementation of the above feature learning attenuation parameter, in which the feature learning attenuation parameter satisfies the following constraint: the feature learning attenuation parameter is negatively correlated with the updating times of the student model. It can be seen that when the feature learning attenuation parameter includes the weight coefficient, the weight coefficient is negatively correlated with the updating times of the student model, that is, as the updating times of the student model gradually increase, the weight coefficient gradually decreases, so that the attenuation degree of the influence degree of the feature learning loss of the student model on the updating process of the student model gradually decreases; when the feature learning attenuation parameter includes the weight reduction amount, the weight reduction amount is negatively correlated with the updating times of the student model, that is, as the updating times of the student model gradually increase, the weight reduction amount gradually decreases, so that the attenuation amount of the influence degree of the feature learning loss of the student model on the updating process of the student model gradually decreases. When the feature learning attenuation parameter includes the weight attenuation probability, the weight attenuation probability is negatively correlated with the updating times of the student model, that is, as the updating times of the student model gradually increase, the weight attenuation probability gradually decreases, so that the occurrence probability of the attenuation processing of the influence degree of the feature learning loss of the student model on the updating process of the student model gradually decreases.
[0247] Step 56: updating the feature learning weight according to the feature learning attenuation parameter.
[0248] It should be noted that the present application does not limit the implementation of step 56, for example, when the feature learning decay parameter comprises a weight coefficient, step 56 can specifically be: multiplying the weight coefficient and the feature learning weight before updating to obtain the feature learning weight after updating, so that the feature learning weight after updating is less than the feature learning weight before updating.
[0249] For another example, when the feature learning decay parameter comprises a weight reduction amount, step 56 above can specifically be: subtracting the weight reduction amount from the feature learning weight before updating to obtain the feature learning weight after updating, so that the feature learning weight after updating is less than the feature learning weight before updating.
[0250] For another example, when the feature learning decay parameter comprises a weight decay probability, step 56 above can specifically be: performing decay processing on the feature learning weight before updating according to the weight decay probability to obtain the feature learning weight after updating.
[0251] Step 57: updating part or all of the student model according to the product of the feature learning weight and the feature learning loss of the student model, and the first prediction result corresponding to the sample data.
[0252] It should be noted that the present application does not limit the implementation of step 57 above, and the following examples are provided for ease of understanding.
[0253] Example 1: In some scenarios, step 57 above can specifically include: first determining the prediction loss of the student model according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data; and then updating part or all of the student model according to the product of the feature learning weight and the feature learning loss of the student model, and the prediction loss. It should be noted that the relevant content of the label information and the prediction loss is described above. In addition, the present application does not limit the implementation of the updating process, for example, it can specifically be: first adding the product and the prediction loss to obtain the model loss of the student model; and then updating part or all of the student model according to the model loss.
[0254] In some scenarios, step 57 in Example 2 can include determining a response learning loss of the student model according to a difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data, and updating part or all of the student model according to a product of the feature learning weight and the feature learning loss of the student model and the response learning loss. It should be noted that the second prediction result and the response learning loss are described above. In addition, the application does not limit the implementation of the updating process. For example, the updating process can include weighting and summing the response learning loss and the feature learning loss according to the feature learning weight and the response learning weight described above to obtain a model loss of the student model, and updating part or all of the student model according to the model loss. In addition, the response learning weight is described above.
[0255] In some scenarios, step 57 in Example 3 can include determining a prediction loss of the student model according to a difference between the first prediction result corresponding to the sample data and the label information of the sample data, and determining a response learning loss of the student model according to a difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data, and updating part or all of the student model according to a product of the feature learning weight and the feature learning loss of the student model, the prediction loss, and the response learning loss. It should be noted that the application does not limit the implementation of the updating process. For example, the updating process can include weighting and summing the prediction loss, the feature learning loss, and the response learning loss according to 1, the feature learning weight, and the response learning weight described above to obtain a model loss of the student model, and updating part or all of the student model according to the model loss. In addition, the response learning weight is described above.
[0256] L = L ori + αL feature + βL response (3)
[0257] In the formula, L represents the model loss of the student model; L ori represents the prediction loss of the student model; L feature represents the feature learning loss of the student model; α represents the feature learning weight; L response represents the response learning loss of the student model; and β represents the response learning weight. α and β are both hyperparameters, and the two weights gradually decrease during the distillation process, which helps the student model focus more on the true label in the later training period.
[0258] Based on the above steps 51 to 57, for some scenarios, such as offline distillation scenarios, in the early stage of training of the learning model, the learning model can be prompted to learn the feature extraction result of the teacher model with a larger weight, so that the learning model can learn the feature extraction performance of the teacher model as quickly as possible. However, in the later stage of training of the learning model, because the feature extraction performance of the learning model is close to that of the teacher model, in order to better improve the model performance and the model training efficiency, the learning model can be prompted to learn the feature extraction result of the teacher model with a smaller weight, so that the learning model can focus on improving its own prediction performance, thereby making the finally trained student model have better prediction performance, which is beneficial to improve the model performance.
[0259] In fact, in some scenarios, such as body contour point detection scenarios, in order to better improve the model performance, the present application also provides a training method of the above student model, in which method, the model updating method provided by the present application is applied to the first training stage and / or the second training stage of the student model; wherein the first training stage is realized by using offline distillation method; the second training stage is realized by using self-distillation method, and the student model in the second training stage is initialized by using the student model obtained in the first training stage.
[0260] In addition, in a possible implementation, in order to better improve the model performance, the losses involved in different training stages of the learning model are different. It can be seen that, in a possible implementation, the two training stages of the learning model have the following characteristics: in the first training stage, the offline distillation method is realized by using the prediction loss and the response learning loss of the student model in the first training stage, and the feature learning loss, the prediction loss and the response learning loss are all determined according to the first prediction result; the feature learning loss is determined according to the output data of the first feature network in the student model; and in the first training stage, the self-distillation method is realized by using the response learning loss of the student model in the second training stage.
[0261] In order to facilitate the understanding of the above two paragraphs, the relevant content of the two training stages of the learning model is introduced as follows.
[0262] For the above first training stage, an already trained teacher model with good performance is first obtained, such as the teacher model shown in Figure 2 Then, the offline distillation method is used to distill the student model using the teacher model, such as Figure 2The student model is shown to learn the feature extraction results and responses of the teacher model. It should be noted that the application does not limit the way of obtaining the initial value of the parameter in the student model in the first training stage. For example, a random acquisition method can be used for implementation, or a manual configuration method can be used for implementation.
[0263] In addition, the application does not limit the implementation of the above first training stage. For example, in a possible implementation, the first training stage can be implemented by using the training process shown in steps 61-73 below.
[0264] Step 61: Obtain sample data and label information of the sample data.
[0265] It should be noted that the related content of step 61 can be referred to the related content of step 21 above.
[0266] Step 62: Determine the first prediction result corresponding to the sample data by using the student model; the student model includes a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain the classification results of the output data of each first sampling layer; and the first prediction result includes the classification results of the output data of the plurality of first sampling layers.
[0267] It should be noted that the related content of step 62 can be referred to the related content of step 12 above.
[0268] Step 63: Determine the second prediction result corresponding to the sample data by using the teacher model; the teacher model includes a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of the corresponding first sampling layer of the second sampling layer in the plurality of first sampling layers above; the second classification layer is used to obtain the classification results of the output data of each second sampling layer; and the second prediction result includes the classification results of the output data of the plurality of second sampling layers.
[0269] It should be noted that the related content of step 63 can be referred to the related content of step 13 above.
[0270] Step 64: Determine the prediction loss of the student model according to the difference between the first prediction result corresponding to the sample data and the label information of the sample data.
[0271] It should be noted that the related content of step 64 can be referred to the related content of step 24 above.
[0272] Step 65: Determine the response learning loss of the student model according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data.
[0273] It should be noted that the related content of step 65 can be referred to the related content of step 25 in the above.
[0274] Step 66: determining a feature learning loss of the student model according to a difference between output data of a first feature network in the student model and output data of a second feature network in the teacher model; the first feature network is used for feature extraction processing on input data of the student model; and the second feature network is used for feature extraction processing on input data of the teacher model.
[0275] It should be noted that the related content of step 66 can be referred to the related content of step 43 in the above. In addition, the present application does not limit the correlation between the execution time of step 66, the execution time of step 65 in the above and the execution time of step 64 in the above, such as the same. For example, the three can be implemented according to a certain pre-set arrangement order.
[0276] Step 67: judging whether the response learning loss of the student model is lower than a preset response loss threshold, if yes, sequentially executing steps 68-70 in the below; if no, executing step 70 in the below.
[0277] It should be noted that the related content of step 67 can be referred to the related content of step 36 in the above. In addition, the present application does not limit the execution time of step 67, as long as the execution time of step 67 is later than the execution time of step 65 in the above.
[0278] Step 68: obtaining a response learning decay parameter and a response learning weight.
[0279] It should be noted that the related content of step 68 can be referred to the related content of step 37 in the above.
[0280] Step 69: updating the response learning weight according to the response learning decay parameter.
[0281] It should be noted that the related content of step 69 can be referred to the related content of step 38 in the above.
[0282] Step 70: judging whether the feature learning loss of the student model is lower than a preset feature loss threshold, if yes, sequentially executing steps 71-73 in the below; if no, executing step 73 in the below.
[0283] It should be noted that the related content of step 70 can be referred to the related content of step 54 in the above.
[0284] Step 71: obtaining a feature learning decay parameter and a feature learning weight.
[0285] It should be noted that the related content of step 71 can be referred to the related content of step 55 in the above.
[0286] Step 72: updating the feature learning weight according to the feature learning decay parameter.
[0287] It should be noted that the related content of step 72 can be referred to the related content of step 55.
[0288] Step 73: updating the student model according to the product of the response learning weight and the response learning loss of the student model, the product of the feature learning weight and the feature learning loss of the student model, and the prediction loss of the student model, and continuing to perform step 61 and the subsequent steps until a pre-set training stop condition is reached.
[0289] It should be noted that the application does not limit the implementation of step 73, for example, step 73 can be specifically: adding the product of the response learning weight and the response learning loss of the student model, the product of the feature learning weight and the feature learning loss of the student model, and the prediction loss of the student model to obtain the model loss of the student model; then updating all parameters in the student model according to the model loss, and returning to continue to perform step 61 and the subsequent steps until a pre-set training stop condition is reached.
[0290] Based on the related content of steps 61-73, in some scenarios, the first training phase of the student model can be implemented by means of Figure 2 The offline distillation mode shown in the figure can be used to implement the first training phase of the student model, so that the student model obtained in the first training phase not only learns the feature extraction result and the response of the teacher model, but also has good prediction performance, which is conducive to improving the model performance of the student model.
[0291] It should be noted that the application does not limit the relationship between the execution time of the above-mentioned step of “judging whether the response learning loss of the student model is lower than the pre-set response loss threshold” and the execution time of the above-mentioned step of “judging whether the feature learning loss of the student model is lower than the pre-set feature loss threshold” in the first training phase, for example, the former is earlier than the latter, such as the relationship between the two shown in steps 67-73. For example, the latter is earlier than the former. For example, the two are the same.
[0292] For the above-mentioned second training phase, the student model obtained in the first training phase can be used as the teacher model required for the second training phase, such as including Figure 3The teacher model of the feature extraction network 3 and the detection network 3 shown and the student model obtained in the first training stage initialize the student model required to be used in the second training stage, so that the initial state of the student model required to be used in the second training stage satisfies the following three constraints: constraint one is that the feature extraction network in the student model required to be used in the second training stage is consistent with the feature extraction network in the student model obtained in the first training stage, constraint two is that the network framework of the detection network in the student model required to be used in the second training stage is consistent with the network framework of the detection network in the student model obtained in the first training stage, and constraint three is that the parameters in the detection network in the student model required to be used in the second training stage are randomly generated. Then, the feature extraction network in the student model required to be used in the second training stage is frozen, and the updating process of the detection network in the student model required to be used in the second training stage is realized in a self-distillation manner under the guidance of the teacher model required to be used in the second training stage.
[0293] In addition, the present application does not limit the implementation of the second training stage described above, for example, in a possible implementation, the second training stage can be implemented by using the training process shown in steps 81-88 below.
[0294] Step 81: Obtain sample data.
[0295] It should be noted that the related content of step 81 can be referred to the related content of S101 described above.
[0296] Step 82: Determine the first prediction result corresponding to the sample data by using the student model; the student model includes a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to obtain the classification results of the output data of each first sampling layer; and the first prediction result includes the classification results of the output data of a plurality of first sampling layers.
[0297] It should be noted that the related content of step 82 can be referred to the related content of step 12 described above.
[0298] Step 83: Determine the second prediction result corresponding to the sample data by using the teacher model; the teacher model includes a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of the corresponding first sampling layer in the plurality of first sampling layers; the second classification layer is used to obtain the classification results of the output data of each second sampling layer; and the second prediction result includes the classification results of the output data of a plurality of second sampling layers.
[0299] It should be noted that the related content of step 83 can be referred to the related content of step 13 described above.
[0300] Step 84: determining the response learning loss of the student model according to the difference between the first prediction result corresponding to the sample data and the second prediction result corresponding to the sample data.
[0301] It should be noted that the related content of step 84 can be referred to the related content of step 25 in the above.
[0302] Step 85: determining whether the response learning loss of the student model is lower than the preset response loss threshold, if yes, sequentially executing steps 86-88 in the following; if no, executing step 88 in the following.
[0303] It should be noted that the related content of step 85 can be referred to the related content of step 36 in the above.
[0304] Step 86: obtaining the response learning decay parameter and the response learning weight.
[0305] It should be noted that the related content of step 86 can be referred to the related content of step 37 in the above.
[0306] Step 87: updating the response learning weight according to the response learning decay parameter.
[0307] It should be noted that the related content of step 87 can be referred to the related content of step 38 in the above.
[0308] Step 88: updating the detection network in the student model according to the product of the response learning weight and the response learning loss of the student model, and continuing to execute step 81 in the above and the subsequent steps until the preset training stop condition is reached.
[0309] It should be noted that the application does not limit the implementation of step 88 in the above, for example, it can be specifically: first, determine the product of the response learning weight and the response learning loss of the student model as the model loss of the student model; then, update the detection network in the student model according to the model loss, and return to continue to execute step 81 in the above and the subsequent steps until the preset training stop condition is reached.
[0310] Based on the related content of steps 81-88 in the above, in some scenarios, the second training phase for the student model can be implemented by means of the self-distillation mode shown in Figure 3 , so that the student model obtained in the second training phase has better prediction performance, which is conducive to improving the model training effect.
[0311] Based on the related content of the model updating method in the above, the application further provides a data processing method, as shown in Figure 5 , the data processing method includes steps S501-S502 in the following.
[0312] S501: Obtain to-be-processed data.
[0313] The to-be-processed data refers to data that needs to be processed by means of a pre-trained model, such as a student model obtained through two training stages.
[0314] In addition, the present application does not limit the implementation of the to-be-processed data, and the following is described in combination with some examples for the convenience of understanding.
[0315] Example 1: If the technical solution provided by the present application is applied to a certain image processing scene, such as face contour point detection processing, body contour point detection processing, target recognition scene, image segmentation scene, edge detection scene, or text recognition processing, etc., the to-be-processed data can be implemented by using an image.
[0316] Example 2: If the technical solution provided by the present application is applied to a certain text processing scene, such as information retrieval or text error correction, etc., the to-be-processed data can be implemented by using a text.
[0317] Example 3: If the technical solution provided by the present application is applied to a certain audio processing scene, such as speech recognition processing or speech translation processing, etc., the to-be-processed data can be implemented by using an audio.
[0318] Example 4: If the technical solution provided by the present application is applied to a certain video processing scene, such as video translation processing, foreground or background replacement processing in a video, or posture adjustment processing in a video, etc., the to-be-processed data can be implemented by using a video.
[0319] In addition, the present application does not limit the above-mentioned obtaining method of the to-be-processed data, for example, it can be implemented by using any kind of to-be-processed data obtaining method existing or appearing in the future, such as data provided by a user by means of some input device.
[0320] S502: Determine the prediction result corresponding to the to-be-processed data by using a model; the model is obtained by using any embodiment of the model updating method provided by the present application.
[0321] The prediction result corresponding to the to-be-processed data refers to the prediction result given by the model trained by the present application, such as a student model obtained through two training stages, for the to-be-processed data, such as a contour point detection result, etc.
[0322] In addition, the present application does not limit the implementation of the prediction result corresponding to the to-be-processed data. For example, the implementation of the prediction result can be determined according to the application scenario of the technical solution provided by the present application. As an example, when the technical solution provided by the present application is applied to the body contour point detection processing scenario, and the to-be-processed data is an image, the prediction result can be used to describe the predicted position of at least one contour point in the image, so that the prediction result is used to represent the body contour detection result given by the model trained by the present application for the to-be-processed data, so that the prediction result can represent the position of all or part of the key points used to describe the body contour in the image. It should be noted that the implementation of the prediction result is similar to the implementation of the first prediction result corresponding to the sample data in the above example, and will not be repeated here for brevity.
[0323] Based on the above S501 to S502, for the data processing method provided by the present application, after obtaining the trained model, such as the student model, the model can be used to predict the to-be-processed data to obtain the prediction result corresponding to the to-be-processed data, such as the body contour point detection result. Wherein, because the model has good prediction performance, so that the prediction result determined by the model is more accurate, so as to improve the data processing performance, such as the body contour point detection performance.
[0324] In addition, the present application does not limit the execution subject of the data processing method provided by the present application. For example, the data processing method provided by the present application can be applied to a terminal device or a server. For another example, the data processing method provided by the present application can also be realized by means of data interaction process between the terminal device and the server.
[0325] Based on the model updating method provided by the present application, the present application further provides a model updating device. The following will be explained and described in combination with Figure 6 The technical details of the model updating device provided by the present application are described in the above model updating method. Figure 6 A structural schematic diagram of a model updating device provided by the present application. It should be noted that the technical details of the model updating device provided by the present application are described in the above model updating method.
[0326] As Figure 6 shown, the model updating device 600 provided by the present application comprises:
[0327] The first acquisition unit 601 is configured to acquire sample data.
[0328] The first prediction unit 602 is configured to determine a first prediction result corresponding to the sample data by using a model; the model comprises a first classification layer and a plurality of first sampling layers; the sampling window size of different first sampling layers is different; the first classification layer is configured to obtain a classification result of output data of each first sampling layer; and the first prediction result comprises classification results of output data of the plurality of first sampling layers.
[0329] The first updating unit 603 is configured to update part or all of the model according to the first prediction result.
[0330] In a possible implementation, the model comprises a feature extraction network and a detection network; the detection network is configured to process output data of the feature extraction network; and the detection network comprises the first classification layer and the plurality of first sampling layers.
[0331] In a possible implementation, the model is a student model.
[0332] The model updating apparatus 600 further comprises:
[0333] The third prediction unit is configured to determine a second prediction result corresponding to the sample data by using a teacher model; the teacher model comprises a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of a corresponding first sampling layer of the second sampling layer in the plurality of first sampling layers; the second classification layer is configured to obtain a classification result of output data of each second sampling layer; and the second prediction result comprises classification results of output data of the plurality of second sampling layers.
[0334] The first updating unit 603 is specifically configured to update part or all of the student model according to a difference between the first prediction result and the second prediction result.
[0335] In a possible implementation, the first updating unit 603 is specifically configured to, for any first sampling layer, determine a loss corresponding to the first sampling layer according to a difference between a classification result of output data of the first sampling layer and a classification result of output data of a reference sampling layer corresponding to the first sampling layer; the plurality of second sampling layers comprise the reference sampling layer, and the sampling window size of the reference sampling layer is the same as the sampling window size of the first sampling layer; and update part or all of the student model according to the losses corresponding to the plurality of first sampling layers.
[0336] In a possible implementation, the plurality of second sampling layers comprise a global sampling layer.
[0337] The model updating apparatus 600 further comprises:
[0338] a weight determination unit configured to determine, for any first sampling layer, a loss weight corresponding to the first sampling layer according to a comparison result between a classification result of output data of a reference sampling layer corresponding to the first sampling layer and a classification result of output data of the global sampling layer;
[0339] The first updating unit 603 is specifically configured to update part or all of the student model according to the loss corresponding to the plurality of first sampling layers and the loss weight corresponding to the plurality of first sampling layers.
[0340] In a possible implementation, for any first sampling layer, if the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is different from the classification result of the output data of the global sampling layer, the loss weight corresponding to the first sampling layer is determined according to a first preset weight; if the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is the same as the classification result of the output data of the global sampling layer, the loss weight corresponding to the first sampling layer is determined according to a second preset weight; the second preset weight is less than the first preset weight.
[0341] In a possible implementation, the model updating apparatus 600 further includes:
[0342] a third obtaining unit configured to obtain label information of the sample data;
[0343] The first updating unit 603 is specifically configured to determine a prediction loss of the student model according to a difference between the first prediction result and the label information; determine a response learning loss of the student model according to a difference between the first prediction result and the second prediction result; and update part or all of the student model according to the prediction loss and the response learning loss.
[0344] In a possible implementation, the model updating apparatus 600 further includes:
[0345] a fourth obtaining unit configured to, if the response learning loss is lower than a preset response loss threshold, obtain a response learning decay parameter and a response learning weight;
[0346] a second updating unit configured to update the response learning weight according to the response learning decay parameter;
[0347] The first updating unit 603 is specifically configured to update part or all of the student model according to a product of the response learning weight and the response learning loss and the prediction loss.
[0348] In a possible implementation, the response learning decay parameter is negatively correlated with the number of updates of the student model; and / or, the response learning decay parameter comprises at least one of a weight coefficient, a weight reduction amount, and a weight decay probability.
[0349] In a possible implementation, the model is a student model.
[0350] The first updating unit 603 is further configured to determine a feature learning loss of the student model according to a difference between output data of a first feature network in the student model and output data of a second feature network in the teacher model, the first feature network being configured to perform feature extraction processing on input data of the student model, and the second feature network being configured to perform feature extraction processing on input data of the teacher model.
[0351] The first updating unit 603 is specifically configured to update part or all of the student model according to the first prediction result and the feature learning loss.
[0352] In a possible implementation, the model updating apparatus 600 further comprises:
[0353] The fifth obtaining unit is configured to obtain a feature learning decay parameter and a feature learning weight if the feature learning loss is lower than a preset feature loss threshold.
[0354] The third updating unit is configured to update the feature learning weight according to the feature learning decay parameter.
[0355] The first updating unit 603 is specifically configured to update part or all of the student model according to a product of the feature learning weight and the feature learning loss and the first prediction result.
[0356] In a possible implementation, the feature learning decay parameter is negatively correlated with the number of updates of the student model; and / or, the feature learning decay parameter comprises at least one of a weight coefficient, a weight reduction amount, and a weight decay probability.
[0357] In a possible implementation, the model is a student model; the model updating method is applied to a first training stage and / or a second training stage of the student model; the first training stage is implemented in an offline distillation manner; the second training stage is implemented in a self-distillation manner, and the student model in the second training stage is initialized by using the student model obtained in the first training stage.
[0358] In a possible implementation, the offline distillation manner is implemented by using a prediction loss and a response learning loss of the student model in the first training stage, the prediction loss and the response learning loss being determined according to the first prediction result; and the self-distillation manner is implemented by using a response learning loss of the student model in the second training stage.
[0359] In a possible implementation, the sample data is an image, and the first prediction result is used to describe a predicted position of at least one contour point in the image.
[0360] Based on the related content of the model updating apparatus 600, the working principle of the model updating apparatus 600 provided in the application includes: after the sample data is acquired, the model is used to determine a first prediction result corresponding to the sample data, such as a body contour point detection result, so that the first prediction result can represent the prediction information given by the model for the sample data, such as the prediction position information of the key points of the body contour; and then part or all of the model is updated according to the first prediction result, so that the updated model has better prediction performance. The model includes a first classification layer and a plurality of first sampling layers, and the first classification layer is used to acquire the classification result of the output data of each first sampling layer, so that the first prediction result includes the classification result of the output data of the plurality of first sampling layers. In addition, because the sampling window sizes of different first sampling layers are different, the output data of different first sampling layers can represent information of different scales in the sample data, such as global information and local information at each scale, so that the classification result of the output data of these first sampling layers can better represent the prediction situation for the sample data, such as the overall prediction result, the local prediction result at each scale, and the like, and thus the classification results can better represent the performance of the model, so that the model updated based on the classification results has better performance, thereby achieving better data processing effect when the model is used for data processing.
[0361] Based on the data processing method provided in the application, the application further provides a data processing apparatus, which will be explained and described below in combination with Figure 7 the drawings. In the drawings, Figure 7 a structure diagram of a data processing apparatus provided in the application is shown. It should be noted that the technical details of the data processing apparatus provided in the application are described in the related content of the data processing method above.
[0362] As Figure 7 shown in the drawings, the data processing apparatus 700 provided in the application includes:
[0363] The second obtaining unit 701 is configured to obtain to-be-processed data.
[0364] The second prediction unit 702 is configured to determine a prediction result corresponding to the to-be-processed data by using a model.
[0365] Based on the above-mentioned related content of the data processing apparatus 700, the working principle of the data processing apparatus 700 provided in the present application includes: after obtaining a trained model, such as a student model, the model can be used to predict the to-be-processed data to obtain a prediction result corresponding to the to-be-processed data, such as a body contour point detection result. Wherein, because the model has good prediction performance, the prediction result determined by the model is more accurate, which is conducive to improving the data processing performance, such as the body contour point detection performance.
[0366] In addition, the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any embodiment of the model updating method provided in the present application, or executes any embodiment of the data processing method provided in the present application.
[0367] Referring to Figure 8 , a structural schematic diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 8 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0368] As shown in Figure 8 , the electronic device 800 can include a processing device (such as a central processor, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 802 or loaded into a random access memory (RAM) 803 from a storage device 808. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0369] Generally, the following devices can be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 808 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 809. The communication devices 809 can allow the electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0370] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 809, or installed from the storage devices 808, or installed from the ROM 802. When the computer program is executed by the processing devices 801, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0371] The electronic device provided by the embodiments of the present disclosure and the method provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiments can be referred to the above embodiments, and the present embodiments have the same beneficial effects as the above embodiments.
[0372] The embodiments of the present disclosure also provide a computer readable medium, wherein instructions or a computer program are stored in the computer readable medium, and when the instructions or the computer program are executed on a device, the device is caused to perform any of the embodiments of the model updating method provided by the embodiments of the present disclosure, or perform any of the embodiments of the data processing method provided by the embodiments of the present disclosure.
[0373] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, in which the computer-readable program code is contained. Such a propagated data signal can take any of a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.
[0374] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0375] The computer-readable medium described above can be included in the electronic device; or exist separately from the electronic device, and not be assembled into the electronic device.
[0376] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method described above.
[0377] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0378] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0379] The units involved in the embodiments of the present disclosure can be implemented by software, or can be implemented by hardware. In some cases, the name of the unit / module does not constitute a limitation on the unit itself.
[0380] The functions described in the above description herein can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0381] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0382] It should be noted that the various embodiments described in the specification are progressive and each embodiment focuses on the differences from other embodiments. The same and similar parts between embodiments can be mutually referred to. For the system or device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0383] It should be understood that in this application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0384] It is also to be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless otherwise indicated. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or "contains" are used in either the detailed description and the claims, such terms are intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or other elements.
[0385] The embodiments disclosed herein can each be implemented as a method, apparatus, or article of manufacture using programming instructions. The embodiments disclosed herein can be implemented using software, firmware, hardware, or a combination thereof. The various elements of the disclosed embodiments, as well as the embodiments themselves, can be constructed from any combination of hardware, software, and / or firmware. The software implementation can be implemented by one or more software modules using object-oriented design methodology, among other techniques. The software modules can be stored on any computer-readable medium, including RAM, ROM, EEPROM, flash memory, or a hard disk, to name a few. The software modules can include one or more routines.
[0386] The above description of disclosed embodiments provides enough information to enable those skilled in the art to make and use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model updating method characterized by, The method comprises: acquiring sample data; determining a first prediction result corresponding to the sample data by using a model; the model comprises a first classification layer and a plurality of first sampling layers; the sampling window sizes of different first sampling layers are different; the first classification layer is used to acquire classification results of output data of each first sampling layer; the first prediction result comprises classification results of output data of the plurality of first sampling layers; updating part or all of the model according to the first prediction result.
2. The method of claim 1, wherein, The model comprises a feature extraction network and a detection network; the detection network is used to process output data of the feature extraction network; the detection network comprises the first classification layer and the plurality of first sampling layers.
3. The method of claim 1, wherein, The model is a student model; The method further comprises: determining a second prediction result corresponding to the sample data by using a teacher model; the teacher model comprises a second classification layer and a plurality of second sampling layers; for any second sampling layer, the sampling window size of the second sampling layer is the same as the sampling window size of a corresponding first sampling layer of the second sampling layer in the plurality of first sampling layers; the second classification layer is used to acquire classification results of output data of each second sampling layer; the second prediction result comprises classification results of output data of the plurality of second sampling layers; The updating part or all of the model according to the first prediction result comprises: updating part or all of the student model according to the difference between the first prediction result and the second prediction result.
4. The method of claim 3, wherein, The updating part or all of the student model according to the difference between the first prediction result and the second prediction result comprises: for any first sampling layer, determining a loss corresponding to the first sampling layer according to the difference between the classification result of the output data of the first sampling layer and the classification result of the output data of a reference sampling layer corresponding to the first sampling layer; the plurality of second sampling layers comprise the reference sampling layer, and the sampling window size of the reference sampling layer is the same as the sampling window size of the first sampling layer; updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers.
5. The method of claim 4, wherein, The plurality of second sampling layers comprise a global sampling layer; The method further comprises: for any first sampling layer, determining a loss weight corresponding to the first sampling layer according to a comparison result between the classification result of the output data of the reference sampling layer corresponding to the first sampling layer and the classification result of the output data of the global sampling layer; The updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers comprises: updating part or all of the student model according to the losses corresponding to the plurality of first sampling layers and the loss weights corresponding to the plurality of first sampling layers.
6. The method of claim 5, wherein, For any first sampling layer, if the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is different from the classification result of the output data of the global sampling layer, the loss weight corresponding to the first sampling layer is determined according to a first preset weight. If the classification result of the output data of the reference sampling layer corresponding to the first sampling layer is the same as the classification result of the output data of the global sampling layer, the loss weight corresponding to the first sampling layer is determined according to the second preset weight. The second preset weight is smaller than the first preset weight.
7. The method of claim 3, wherein, The method further comprises: obtaining label information of the sample data; the updating part or all of the student model according to the difference between the first prediction result and the second prediction result comprises: determining a prediction loss of the student model according to the difference between the first prediction result and the label information; determining a response learning loss of the student model according to the difference between the first prediction result and the second prediction result; updating part or all of the student model according to the prediction loss and the response learning loss.
8. The method of claim 7, wherein, The method further comprises: if the response learning loss is lower than a preset response loss threshold, obtaining a response learning decay parameter and a response learning weight; updating the response learning weight according to the response learning decay parameter; the updating part or all of the student model according to the prediction loss and the response learning loss comprises: updating part or all of the student model according to the product of the response learning weight and the response learning loss and the prediction loss.
9. The method of claim 8, wherein, The response learning decay parameter is negatively correlated with the number of updates of the student model. and / or The response learning decay parameter comprises at least one of a weight coefficient, a weight reduction amount and a weight decay probability.
10. The method of claim 1, wherein, The model is a student model. The method further comprises: determining a feature learning loss of the student model according to the difference between the output data of a first feature network in the student model and the output data of a second feature network in a teacher model; the first feature network is used for feature extraction processing on input data of the student model; and the second feature network is used for feature extraction processing on input data of the teacher model; the updating part or all of the student model according to the first prediction result comprises: updating part or all of the student model according to the first prediction result and the feature learning loss.
11. The method of claim 10, wherein, The method further comprises: if the feature learning loss is lower than a preset feature loss threshold, obtaining a feature learning decay parameter and a feature learning weight; updating the feature learning weight according to the feature learning decay parameter; the updating part or all of the student model according to the first prediction result and the feature learning loss comprises: updating part or all of the student model according to the product of the feature learning weight and the feature learning loss and the first prediction result.
12. The method of claim 11, wherein, The feature learning decay parameter is negatively correlated with the number of updates of the student model. and / or The feature learning decay parameter comprises at least one of a weight coefficient, a weight reduction amount and a weight decay probability. The characteristic learning decay parameter comprises at least one of a weight coefficient, a weight reduction amount, and a weight decay probability.
13. The method of claim 1, wherein, The model is a student model. The model updating method is applied to a first training stage and / or a second training stage of the student model. The first training stage is implemented in an offline distillation manner. The second training stage is implemented in a self-distillation manner, and the student model in the second training stage is initialized by using the student model obtained in the first training stage.
14. The method of claim 13, wherein, The offline distillation manner is implemented by using a prediction loss, a response learning loss, and a characteristic learning loss of the student model in the first training stage, the prediction loss and the response learning loss are determined according to the first prediction result, and the characteristic learning loss is determined according to output data of a first characteristic network in the student model. The self-distillation manner is implemented by using a response learning loss of the student model in the second training stage.
15. The method according to any one of claims 1 to 14, characterized in that, The sample data is an image. The first prediction result is used to describe a predicted position of at least one contour point in the image.
16. A data processing method, characterized by, The method comprises: Obtaining to-be-processed data; Determining a prediction result corresponding to the to-be-processed data by using a model, wherein the model is obtained by using the model updating method in any one of claims 1-15.
17. A model updating apparatus characterized by comprising: Comprise: A first obtaining unit is configured to obtain sample data. A first prediction unit is configured to determine a first prediction result corresponding to the sample data by using a model, wherein the model comprises a first classification layer and a plurality of first sampling layers, and sampling window sizes of different first sampling layers are different. The first classification layer is configured to obtain classification results of output data of each first sampling layer, and the first prediction result comprises the classification results of the output data of the plurality of first sampling layers. A first updating unit is configured to update part or all of the model according to the first prediction result.
18. A data processing apparatus, characterized by Comprise: A second obtaining unit is configured to obtain to-be-processed data. A second prediction unit is configured to determine a prediction result corresponding to the to-be-processed data by using a model. The model is obtained by using the model updating method in any one of claims 1-15.
19. An electronic device, comprising: The device comprises a processor and a memory. The memory is configured to store instructions or computer programs. The processor is configured to execute the instructions or computer programs in the memory, so that the electronic device executes the method in any one of claims 1-16.
20. A computer readable medium characterized by The computer readable medium stores instructions or computer programs, when the instructions or computer programs run on the device, the device executes the method in any one of claims 1-16.
21. A computer program product, characterised in that, It comprises a computer program carried on a non-transitory computer readable medium, and the computer program comprises program code for executing the method in any one of claims 1-16.