Network model training method and device

By mixing and self-distillation training on the original image, the problem of low neural network training efficiency and accuracy is solved, and more efficient training and better image recognition effect is achieved.

CN120599435APending Publication Date: 2025-09-05BEIJING JINGDONG YUANSHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410251355.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-09-05

Smart Images

  • Figure CN120599435A_ABST
    Figure CN120599435A_ABST
Patent Text Reader

Abstract

The invention discloses a network model training method and device, and relates to the technical field of computers. A specific embodiment of the training method of the network model comprises the following steps: in response to a plurality of received original images, mixing the plurality of original images to obtain a mixed image; performing feature extraction on the plurality of original images and the mixed image according to a preset neural network to obtain a plurality of feature maps; according to the neural network, classifying the plurality of feature maps to obtain probability distribution; and according to a preset loss function, determining the loss of the plurality of feature maps and the probability distribution, according to the loss, carrying out back propagation on the neural network, and updating the weight of the neural network. According to the embodiment, self-distillation training is performed on the neural network based on the original image and the mixed image, so that the training efficiency and training effect of the neural network can be improved, and the precision of the trained neural network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a training method and device for a network model. Background Art

[0002] Knowledge distillation is a research area in deep learning. When performing knowledge distillation on a neural network, a teacher model is typically used to supervise and guide the training of a student model. For example, the teacher model is pre-trained, and then the knowledge it possesses is used to assist in training the student model. Self-distillation technology builds on knowledge distillation by omitting the teacher model and training based on individual input samples. The student model then mines the knowledge between input samples based on its own knowledge.

[0003] In the process of implementing the present invention, the inventors discovered that the prior art has at least the following problems:

[0004] The amount of knowledge mined from a single input sample is limited, resulting in low accuracy of the trained network model and low training efficiency. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method and apparatus for training a network model, which can improve the training effect and training efficiency of a neural network and improve the accuracy of the trained neural network.

[0006] To achieve the above object, according to a first aspect of an embodiment of the present invention, a method for training a network model is provided, comprising:

[0007] In response to receiving a plurality of original images, mixing the plurality of original images to obtain a mixed image;

[0008] According to a preset neural network, feature extraction is performed on the multiple original images and the mixed image respectively to obtain multiple feature maps;

[0009] Classifying the plurality of feature maps according to the neural network to obtain a probability distribution;

[0010] According to a preset loss function, the loss between the multiple feature maps and the probability distribution is determined, and the neural network is back-propagated according to the loss to update the weight of the neural network.

[0011] Optionally, the original image includes: image information and label information; mixing the multiple original images to obtain a mixed image includes:

[0012] Mixing the image information of the plurality of original images according to a preset mixing weight;

[0013] The label information of the plurality of original images is mixed according to the mixing weight, and the mixed image information and the mixed label information are used as a mixed image.

[0014] Optionally, the multiple feature maps include: original feature maps and mixed feature maps; and determining the loss of the multiple feature maps according to a preset loss function includes:

[0015] Determining the consistency loss between the original feature map and the hybrid feature map according to a loss function;

[0016] Image information of the original image is obtained, a linear transformation is performed on the consistency loss according to the image information, and the consistency loss after the linear transformation is used as the loss of the multiple feature maps.

[0017] Optionally, the probability distribution includes: an original probability distribution and a mixed probability distribution; and determining the loss of the probability distribution according to a preset loss function includes:

[0018] According to the loss function, the loss between the original probability distribution and the mixed probability distribution is determined, and the loss between the original probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

[0019] Optionally, the original probability distribution includes: a probability distribution of each original image; and determining the loss between the original probability distribution and the mixed probability distribution according to a loss function includes:

[0020] Performing linear interpolation on the probability distributions of the multiple original images according to preset probability weights to obtain an interpolated probability distribution;

[0021] According to the loss function, a loss between the interpolated probability distribution and the mixed probability distribution is determined, and the loss between the interpolated probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

[0022] Optionally, before backpropagating the neural network according to the loss, the method further includes:

[0023] Determine the cross entropy loss for each original probability distribution;

[0024] According to the preset loss weights, the cross entropy loss, the loss of the feature map and the loss of the probability distribution are weighted and summed to obtain a comprehensive loss, which is used to perform backpropagation on the neural network.

[0025] According to a second aspect of an embodiment of the present invention, there is provided a method for image recognition, comprising:

[0026] In response to receiving a target image, extracting features of the target image according to a preset image recognition model to obtain a feature map of the target image, wherein the image recognition model is obtained using any training method described in the first aspect of the embodiments of the present invention;

[0027] Classifying the feature map according to the image recognition model to obtain a probability distribution of the feature map;

[0028] A classification result of the target image is determined according to the probability distribution.

[0029] According to a third aspect of an embodiment of the present invention, there is provided a network model training device, comprising:

[0030] a mixing module, configured to, in response to receiving a plurality of original images, mix the plurality of original images to obtain a mixed image;

[0031] A feature module, configured to extract features from the plurality of original images and the mixed image according to a preset neural network to obtain a plurality of feature maps;

[0032] A probability module, configured to classify the plurality of feature maps according to the neural network to obtain a probability distribution;

[0033] An updating module is used to determine the loss between the multiple feature maps and the probability distribution according to a preset loss function, perform backpropagation on the neural network according to the loss, and update the weights of the neural network.

[0034] Optionally, the original image includes: image information and label information; mixing the multiple original images to obtain a mixed image includes:

[0035] Mixing the image information of the plurality of original images according to a preset mixing weight;

[0036] The label information of the plurality of original images is mixed according to the mixing weight, and the mixed image information and the mixed label information are used as a mixed image.

[0037] Optionally, the multiple feature maps include: original feature maps and mixed feature maps; and determining the loss of the multiple feature maps according to a preset loss function includes:

[0038] Determining the consistency loss between the original feature map and the hybrid feature map according to a loss function;

[0039] Image information of the original image is obtained, a linear transformation is performed on the consistency loss according to the image information, and the consistency loss after the linear transformation is used as the loss of the multiple feature maps.

[0040] Optionally, the probability distribution includes: an original probability distribution and a mixed probability distribution; and determining the loss of the probability distribution according to a preset loss function includes:

[0041] According to the loss function, the loss between the original probability distribution and the mixed probability distribution is determined, and the loss between the original probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

[0042] Optionally, the original probability distribution includes: a probability distribution of each original image; and determining the loss between the original probability distribution and the mixed probability distribution according to a loss function includes:

[0043] Performing linear interpolation on the probability distributions of the multiple original images according to preset probability weights to obtain an interpolated probability distribution;

[0044] According to the loss function, a loss between the interpolated probability distribution and the mixed probability distribution is determined, and the loss between the interpolated probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

[0045] Optionally, the device further comprises:

[0046] The cross-loss module is used to determine the cross-entropy loss of each original probability distribution;

[0047] The comprehensive loss module is used to perform weighted summation of the cross entropy loss, the loss of the feature map and the loss of the probability distribution according to a preset loss weight to obtain a comprehensive loss, and the comprehensive loss is used to perform backpropagation on the neural network.

[0048] According to a fourth method of an embodiment of the present invention, a device for image recognition is provided, including:

[0049] a feature module configured to, in response to receiving a target image, extract features of the target image according to a preset image recognition model to obtain a feature map of the target image, wherein the image recognition model is obtained using any training method described in the first aspect of the embodiments of the present invention;

[0050] A probability module, configured to classify the feature map according to the image recognition model to obtain a probability distribution of the feature map;

[0051] A recognition module is used to determine a classification result of the target image according to the probability distribution.

[0052] According to a fifth aspect of an embodiment of the present invention, there is provided an electronic device, including:

[0053] one or more processors;

[0054] a storage device for storing one or more programs,

[0055] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the above embodiments.

[0056] According to a sixth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in any one of the above embodiments is implemented.

[0057] One embodiment of the above invention has the following advantages or beneficial effects: self-distillation training of the neural network based on the original image and the mixed image can improve the training efficiency and training effect of the neural network, and improve the accuracy of the trained neural network; according to the mixing weight, the image information and label information of the original image are mixed respectively, and then the mixed image information and label information are combined to obtain a mixed image, which can improve the flexibility of image mixing, provide more input samples for the neural network, and facilitate the neural network to mine more knowledge during the training process, thereby improving the accuracy of the neural network; determining the consistency loss of the original feature map and the mixed feature map is equivalent to self-distilling the feature map, which can enable the neural network to learn more semantic features and guide the neural network to produce consistent outputs for the original image and the mixed image; using the KL divergence loss function to determine the loss between the original probability distribution and the mixed probability distribution is equivalent to self-distilling the probability distribution, which can enable the neural network to have robust mixed prediction capabilities; using the trained neural network as an image recognition model to perform image recognition on the target image can improve image recognition efficiency and recognition accuracy.

[0058] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.

[0060] Figure 1 is a schematic diagram of the main process of the training method of the network model according to an embodiment of the present invention;

[0061] Figure 2 is a schematic diagram of a training process of a network model according to a reference embodiment of the present invention;

[0062] Figure 3 1 is a schematic diagram of the main process of a network model training method according to a reference embodiment of the present invention;

[0063] Figure 4 is a schematic diagram of the main process of a network model training method according to another reference embodiment of the present invention;

[0064] Figure 5 is a schematic diagram of the main process of the image recognition method according to an embodiment of the present invention;

[0065] Figure 6 is a schematic diagram of main modules of a network model training device according to an embodiment of the present invention;

[0066] Figure 7 is a schematic diagram of main modules of an apparatus for image recognition according to an embodiment of the present invention;

[0067] Figure 8 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;

[0068] Figure 9 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0069] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0070] It should be noted that in the technical solution of the present invention, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the user's consent or authorization. When applicable, the user's personal information is de-identified and / or anonymized and / or encrypted.

[0071] Knowledge distillation is a research area in deep learning. When performing knowledge distillation on a neural network, a teacher model is typically used to supervise and guide the training of a student model. For example, the teacher model is pre-trained, and then the knowledge it possesses is used to assist in training the student model. Self-distillation technology builds on knowledge distillation by omitting the teacher model and training based on individual input samples. The student model then mines the knowledge between input samples based on its own knowledge.

[0072] The amount of knowledge mined from a single input sample is limited, resulting in low accuracy and training efficiency for the trained network model. Before using the teacher model to assist in training the student model, the teacher model must be pre-trained, which prolongs the training iteration cycle, reduces training efficiency, consumes more resources, and increases R&D costs.

[0073] In view of this, according to a first aspect of an embodiment of the present invention, a method for training a network model is provided.

[0074] Figure 1 Schematic diagram of the main process of the training method of the network model according to an embodiment of the present invention. Figure 1 As shown, the training method of the network model according to the embodiment of the present invention mainly includes the following steps S101 to S104.

[0075] Step S101 : In response to receiving a plurality of original images, the plurality of original images are mixed to obtain a mixed image.

[0076] The original image is an input sample for training a neural network. The multiple original images have the same height, width, and number of image channels, where the number of image channels is used to describe each pixel included in the image. For example, when grayscale values ​​are used to describe pixels, the corresponding number of image channels is 1. When RGB (Red, Green, Blue) values ​​are used to describe pixels, the corresponding number of image channels is 3. After receiving the multiple original images, the execution subject of the embodiment of the present invention mixes the multiple original images to obtain a mixed image. Specifically, the execution subject of the embodiment of the present invention performs image processing on the multiple original images and combines the multiple processed original images to obtain a mixed image. The number of mixed images can be one or more.

[0077] For example, the execution subject of an embodiment of the present invention inserts a mask into the original image A1, that is, covers up part of the information of the original image A1, determines the position of the mask insertion in the original image A1, and adds the information at the corresponding position in the original image A2 to the original image A1, so that the original image A1 has part of the information of the original image A2, which is equivalent to mixing the original image A1 and the original image A2; then repeats the above steps of inserting the mask, determining the position of the mask, adding other image information, etc., and continues to mix the original image A1 with other images, or mix between other images, and uses the mixed original image as the mixed image. It should be noted that before mixing the original images, the execution subject of the embodiment of the present invention stores the original images, so that the original images can still be obtained while the mixed image is obtained.

[0078] As another example, the execution subject of an embodiment of the present invention selects original images for image blending from a plurality of original images, that is, the original images participating in the image blending may be all original images or a portion of all original images, and the selected original images are blended to obtain a blended image. Preset filtering conditions are provided, for example, the filtering conditions include: selecting two original images for image blending from every five original images in the order in which the original images are received; the selected original images belonging to the same scene (e.g., a vehicle driving scene, a commodity recognition scene, a cargo warehousing scene, etc.); and the number of original images for blending is greater than or equal to five; and the original images that meet the filtering conditions are blended to obtain a blended image.

[0079] Mixing original images into mixed images can improve the diversity of input samples of the neural network, facilitate improving the training efficiency and training effect of the neural network, and improve the accuracy of the trained neural network; setting filtering conditions can improve the flexibility of image mixing and improve the efficiency of image mixing.

[0080] According to a reference embodiment of the present invention, an original image includes image information and label information. The image information includes information such as the original image's height, width, number of image channels, and a pixel value matrix. The label information includes the original image's classification label. For example, if the original image depicts a vehicle driving scene, the label information for the original image may include two vehicle labels, three pedestrian labels, one traffic light label, three guardrail labels, four road labels, and so on. When blending multiple original images to obtain a blended image, pre-set blending weights are first obtained. The blending weights include blending weights for different images, with each image having a corresponding blending weight. The image information of the multiple original images is blended according to the blending weights. The label information of the multiple original images is then blended according to the blending weights. The weights for blending the image information and the weights for blending the label information are the same. The sum of all blending weights is equal to 1. The blended image information and the blended label information are used as the blended image. Exemplarily, the image information x1 of the original image B1 and the image information x2 of the original image B2 are mixed to obtain x3, and the specific calculation formula is: x3 = βx1 + (1-β)x2; the label information y1 of the original image B1 and the label information y2 of the original image B2 are mixed to obtain y3, and the specific calculation formula is: y3 = βy1 + (1-β)y2; the mixed image information x3 and the mixed label information y3 are used as a mixed image.

[0081] It should be noted that the number of mixed images can be one or more, and the mixed image includes: mixed image information and mixed label information.

[0082] For example, the preset image information mixing formula is: Among them, x h Represents the mixed image information, n represents the number of original images involved in image mixing, x i represents the image information of the original image participating in the image mixing, λ i Represents the mixing weight of the original image participating in the image mixing, and the image information of the original image is mixed according to the above image information mixing formula to obtain the mixed image information. The preset label information mixing formula is: Among them, y h Represents the mixed label information, n represents the number of original images involved in image mixing, y i Indicates the label information of the original image participating in the image mixing, λ i Represents the mixing weight of the i-th original image participating in the image mixing. According to the above label information mixing formula, the label information of the original image is mixed to obtain the mixed label information. The mixed image information and the mixed label information are taken as the mixed image. It should be noted that the mixing weight of the original image satisfies the following formula: That is, the sum of all mixing weights is equal to 1.

[0083] According to the mixing weights, the image information and label information of the original image are mixed respectively, and then the mixed image information and label information are combined to obtain a mixed image. This can improve the flexibility of image mixing, provide more input samples for the neural network, facilitate the neural network to mine more knowledge during the training process, and improve the accuracy of the neural network.

[0084] Step S102 : performing feature extraction on the multiple original images and the mixed image respectively according to a preset neural network to obtain multiple feature maps.

[0085] After obtaining the mixed image, multiple original images and mixed images are respectively input into the neural network, and the neural network is used to perform feature extraction on the multiple original images and mixed images respectively to obtain multiple original feature maps (i.e., feature maps of the original images) and mixed feature maps (i.e., feature maps of the mixed images). Specifically, the neural network includes a feature extractor, and the feature extractors suitable for the embodiment of the present invention include: a residual neural network (ResNet), and a network improved thereon. Two neural networks with the same structure are pre-set, and the two neural networks include the same feature extractor. The execution subject of the embodiment of the present invention uses the feature extractor of one of the neural networks to extract features from each original image to obtain the original feature map corresponding to each original image, and uses the feature extractor of the other neural network to extract features from each mixed image to obtain the mixed feature map corresponding to the mixed image.

[0086] Figure 2Schematic diagram of the training process of a network model according to a reference embodiment of the present invention. Figure 2 As shown, the original images received by the execution subject of the embodiment of the present invention include: original image 1 and original image 2, and the original image 1 and original image 2 are mixed to obtain a mixed image; the pre-set neural network 1 includes: feature extractor 1 and classifier 1, and the pre-set neural network 2 includes: feature extractor 2 and classifier 2, the two feature extractors have the same structure and share weights, and the two classifiers have the same structure and share weights; the original image 1 and the original image 2 are input into the feature extractor 1 to obtain the original feature map, and the mixed image is input into the feature extractor 2 to obtain the mixed feature map.

[0087] Using two neural networks with the same structure to extract features from the original image and the mixed image respectively can obtain multiple feature maps, so that the feature extractor in the neural network can learn more image features, which is convenient for improving the training efficiency of the feature extractor.

[0088] Step S103: Classify the multiple feature maps according to the neural network to obtain a probability distribution.

[0089] After obtaining the original feature map and the mixed feature map, a neural network is used to classify the multiple original feature maps and mixed feature maps respectively to obtain multiple probability distributions. Specifically, the neural network includes a classifier, and the classifiers suitable for the embodiment of the present invention include: KNN (K-Nearest Neighbor) algorithm classifier, CNN (Convolutional Neural Networks) classifier, and the like. Classifiers with the same structure are set in two neural networks with the same structure. The execution subject of the embodiment of the present invention uses the feature extractor of the neural network that extracts features from the original image to classify each original feature map to obtain the original probability distribution (that is, the probability distribution corresponding to each original feature map), and uses the feature extractor of the neural network that extracts features from the mixed image to classify each mixed feature map to obtain a mixed probability distribution (that is, the feature map corresponding to each mixed image).

[0090] Figure 2 Schematic diagram of the training process of a network model according to a reference embodiment of the present invention. Figure 2 As shown, the original feature map is input into classifier 1 to obtain the probability distribution of the original feature map, that is, the original probability distribution, and the mixed feature map is input into classifier 2 to obtain the probability distribution of the mixed feature map, that is, the mixed probability distribution.

[0091] Using two neural networks with the same structure to classify the original feature map and the mixed feature map respectively can obtain multiple probability distributions, so that the classifier in the neural network can learn more probability distributions, which is convenient for improving the training efficiency of the classifier.

[0092] Step S104: determining the loss between the plurality of feature maps and the probability distribution according to a preset loss function, performing backpropagation on the neural network according to the loss, and updating the weights of the neural network.

[0093] Multiple loss functions are pre-set, including: a loss function for calculating feature map loss and a probability loss function for calculating probability distribution loss. For example, the pre-set loss functions include: mean square error (MSE) loss function, mean absolute error (MAE) loss function, binary cross entropy (BCE) loss function, categorical cross entropy (Categorical Cross Entropy) loss function, etc. The execution subject of the embodiment of the present invention uses the mean square error loss function or the mean absolute error loss function to calculate the loss of multiple feature maps, and uses the binary cross entropy loss function or the categorical cross entropy loss function to calculate the loss of multiple probability distributions.

[0094] The execution subject of the embodiment of the present invention uses the losses of multiple feature maps to perform backpropagation on the feature extractor in the neural network and update the weights of the feature extractor in the neural network. The execution subject of the embodiment of the present invention uses the losses of multiple probability distributions to perform backpropagation on the classifier in the neural network and update the weights of the classifier in the neural network. The above steps of calculating the loss, backpropagation, and weight updating are repeated until the loss of the feature map and the loss of the probability distribution meet a preset loss condition, for example, until the loss of the feature map and the loss of the probability distribution are less than a preset loss threshold. If the loss meets the loss condition, the training of the neural network is stopped, and the trained neural network is used to perform image recognition, object detection, semantic segmentation, and other image applications on the input image. It should be noted that when performing backpropagation and weight updating on the neural network, the execution subject of the embodiment of the present invention performs backpropagation and weight update on two pre-set neural networks with the same structure respectively, or the execution subject of the embodiment of the present invention performs backpropagation and weight update on one neural network and then uses the updated weights to update the weights of the other neural network.

[0095] Backpropagation and weight update of the neural network based on the feature map loss, probability distribution loss of the original image and the feature map loss and probability distribution loss of the mixed image can improve the training efficiency and training effect of the neural network, improve the accuracy of the trained neural network, accelerate the iteration speed of the neural network, shorten the training time of the neural network, and reduce R&D costs.

[0096] According to a reference embodiment of the present invention, the multiple feature maps include: original feature maps and mixed feature maps. When determining the loss of multiple feature maps according to a pre-set loss function, first determine the consistency loss between the original feature map and the mixed feature map according to the loss function. For example, the loss function includes: Least Squares Error (LSE) loss function, that is, L2 distance. According to the least squares error loss function, the consistency loss between the original feature map and the mixed feature map is determined, that is, the L2 distance is used to constrain the consistency between the original image and the mixed image, which is equivalent to self-distillation of the original feature map. The image information of the original image is obtained, and the image information includes: the height, width, number of image channels, etc. of the original image. After determining the consistency loss between the original feature map and the mixed feature map, the consistency loss is linearly transformed according to the image information, and the consistency loss after the linear transformation is used as the loss of multiple feature maps.

[0097] Exemplarily, the original image received by the execution subject of the embodiment of the present invention includes the original image u i and the original image u j , the above two original images are mixed to obtain the mixed image u ij , the calculation formula for the loss of multiple pre-set feature maps is: Among them, L feature represents the consistency loss between the original feature map and the mixed feature map, F(u i ,u j ) indicates that the feature extractor is used to extract the original image u i and the original image u j The original feature map obtained after feature extraction, F(u ij ) represents the use of feature extractor to analyze the mixed image u ij After feature extraction, the mixed feature map is obtained, where H represents the height of the original image, W represents the width of the original image, and C represents the number of image channels of the original image. i ,u j )-F(u ij )|| 2 Represents the consistency loss between the original feature map and the hybrid feature map, HWC represents the parameter for linear transformation of the consistency loss; the loss of multiple feature maps is determined according to the above calculation formula.

[0098] For example, Figure 2 As shown, feature map self-distillation is performed on the original feature map and the mixed feature map. Specifically, the consistency loss of the original feature map and the mixed feature map is calculated, the consistency loss is linearly transformed, and feature extractor 1 and feature extractor 2 are back-propagated according to the consistency loss after linear transformation, and the weights of feature extractor 1 and feature extractor 2 are updated. The weights of feature extractor 1 and feature extractor 2 after update are the same.

[0099] Determining the consistency loss of the original feature map and the mixed feature map is equivalent to self-distillation of the feature map, which enables the neural network to learn more semantic features and guides the neural network to produce consistent output for the original image and the mixed image.

[0100] According to another reference embodiment of the present invention, the probability distribution includes: an original probability distribution and a mixed probability distribution. When determining the loss of the probability distribution according to a pre-set loss function, the loss between the original probability distribution and the mixed probability distribution is determined according to the loss function. For example, the loss function includes: a KL divergence (Kullback-LeiblerDivergence) loss function. According to the KL divergence loss function, the loss between the original probability distribution and the mixed probability distribution is determined, which is equivalent to using the KL divergence loss function to mutually distill the probability distribution of the original feature map and the probability distribution of the mixed feature map, and taking the loss between the original probability distribution and the mixed probability distribution as the loss of the probability distribution. Among them, the original probability distribution can be the probability distribution of a single original feature map, and the loss of the probability distribution of each original feature map and the mixed probability distribution is calculated; the original probability distribution can also be the sum of the probability distributions of multiple original feature maps. The loss of the sum of the probability distributions of multiple original feature maps and the mixed probability distribution is calculated.

[0101] Exemplarily, the original image received by the execution subject of the embodiment of the present invention includes the original image w i and the original image w j , the above two original images are mixed to obtain the mixed image w ij , the calculation formula for the loss of multiple pre-set feature maps is: L logit =L KL (f(w i ,w j ),f(w ij )), where L logit represents the loss of probability distribution, f(w i ,w j ) represents the original probability distribution of the original feature map, f(w ij ) represents the mixed probability distribution of the mixed feature map, L KL (f(w i,w j ),f(w ij )) represents the calculation of the KL divergence loss between the original probability distribution and the mixed probability distribution; the loss of the probability distribution is determined according to the above calculation formula.

[0102] For example, Figure 2 As shown, the original probability distribution and the mixed probability distribution are subjected to probability distribution self-distillation. Specifically, the KL divergence loss of the original probability distribution and the mixed probability distribution is calculated, and classifier 1 and classifier 2 are backpropagated according to the KL divergence loss. The weights of classifier 1 and classifier 2 are updated, and the updated weights of classifier 1 and classifier 2 are the same.

[0103] Using the KL divergence loss function to determine the loss between the original probability distribution and the mixed probability distribution is equivalent to self-distilling the probability distribution, which can enable the neural network to have robust mixed prediction capabilities.

[0104] According to another reference embodiment of the present invention, the original probability distribution includes: a probability distribution for each original image. When determining the loss between the original probability distribution and the mixed probability distribution according to the loss function, the probability distributions of multiple original images are first linearly interpolated according to pre-set probability weights to obtain an interpolated probability distribution. Specifically, the probability distributions of multiple original feature maps are weighted and summed using the probability weights, and the weighted summation result is used as the interpolated probability distribution. Then, according to the loss function, the loss between the interpolated probability distribution and the mixed probability distribution is determined, and the loss between the interpolated probability distribution and the mixed probability distribution is used as the loss of the probability distribution. For example, the loss function includes: a KL divergence loss function, and the loss between the interpolated probability distribution and the mixed probability distribution is determined according to the KL divergence loss function. It should be noted that the sum of all probability weights is equal to 1. The interpolated probability distribution can be used as a pseudo-teacher probability distribution, and the mixed probability distribution can be used as a data augmentation probability distribution, using the mixed probability distribution to approximate the interpolated probability distribution. The interpolated probability distribution incorporates the prediction information of all original feature maps.

[0105] For example, the original image received by the execution subject of the embodiment of the present invention includes the original image v i and the original image v j , the preset linear interpolation formula is: f(v i ,v j )=α*f(v i )+(1-α)*f(v j ), where f(v i ,v j ) represents the interpolation probability distribution, f(v i ) represents the original image v i The probability distribution of f(vj ) represents the original image v j The probability distribution of α represents the original image v i The probability weight of the original image v is equal to 1. j The probability weight is (1-α), and the above calculation formula is used to determine the interpolation probability distribution of the original feature map.

[0106] Linear interpolation and weighted summation of the probability distributions of multiple original feature maps can enable the classifier to learn robust hybrid prediction results, thereby improving the training efficiency and classification accuracy of the classifier.

[0107] According to another reference embodiment of the present invention, before backpropagating the neural network according to the loss, the method further includes: using a preset cross entropy loss function to determine the cross entropy loss of each original probability distribution, and according to the preset loss weight, performing weighted summation on the obtained cross entropy loss, the loss of the feature map, and the loss of the probability distribution to obtain a comprehensive loss, and the comprehensive loss is used to backpropagate the neural network.

[0108] Exemplarily, the original image received by the execution subject of the embodiment of the present invention includes the original image p i (including label information i ) and the original image p j (including label information j ), the preset linear interpolation formula is: L mixSKD =L ce (f(p i ),q i )+L ce (f(p j ),q j )+β1*L feature +β2*L logit , where L mixSKD represents the comprehensive loss, f(p i ) represents the original image p i The probability distribution of L ce (f(p i ),q i ) represents the original image p i The cross entropy loss, f(p j ) represents the original image p j The probability distribution of L ce (f(p j ),q j ) represents the original image p j The cross entropy loss, L featureRepresents the loss of the feature map, β1 represents the loss weight corresponding to the loss of the feature map, L logit Represents the loss of probability distribution, β2 represents the loss weight corresponding to the loss of probability distribution; the above calculation formula is used to determine the comprehensive loss.

[0109] Backpropagation and weight updating of the neural network based on comprehensive loss can improve the training efficiency and training effect of the neural network, improve the accuracy of the trained neural network, accelerate the iteration speed of the neural network, shorten the training time of the neural network, and reduce R&D costs.

[0110] Figure 3 FIG. 1 is a schematic diagram of the main process of a training method for a network model according to a reference embodiment of the present invention. Figure 3 As shown, the training method of the network model may include:

[0111] Step S301, in response to receiving a plurality of original images, mixing image information of the plurality of original images according to a preset mixing weight;

[0112] Step S302, mixing label information of multiple original images according to the mixing weight;

[0113] Step S303, combining the mixed image information and label information to obtain a mixed image;

[0114] Step S304, using the feature extractor of the first neural network to extract features from the multiple original images to obtain an original feature map;

[0115] Step S305, using the feature extractor of the second neural network to extract features from the mixed image to obtain a mixed feature map, wherein the feature extractor of the first neural network and the feature extractor of the second neural network have the same structure and share weights;

[0116] Step S306, using the classifier of the first neural network to classify the original feature map to obtain an original probability distribution;

[0117] Step S307: using the classifier of the second neural network to classify the mixed feature map to obtain a mixed probability distribution, the classifier of the first neural network and the classifier of the second neural network have the same structure and share weights;

[0118] Step S308, according to a preset loss function, determine the losses of the original feature map, the mixed feature map, the original probability distribution and the mixed probability distribution, perform backpropagation on the neural network according to the above losses, and update the weights of the neural network.

[0119] The specific implementation content of the training method of the network model of the above-mentioned reference embodiment of the present invention has been described in detail in the training method of the network model described above, so the repeated content will not be described again here.

[0120] Figure 4 FIG. 1 is a schematic diagram of the main process of a training method for a network model according to another reference embodiment of the present invention. Figure 4 As shown, the training method of the network model may include:

[0121] Step S401, in response to receiving a plurality of original images, mixing the plurality of original images to obtain a mixed image;

[0122] Step S402: performing feature extraction on the multiple original images and the mixed image according to a preset neural network to obtain multiple feature maps;

[0123] Step S403: classify the multiple feature maps according to the neural network to obtain a probability distribution;

[0124] Step S404, determining the loss between multiple feature maps according to a preset least square error loss function, and determining the loss between multiple probability distributions according to a KL divergence loss function;

[0125] Step S405 , backpropagating the neural network according to the loss between the multiple feature maps and the loss between the multiple probability distributions, and updating the weights of the neural network.

[0126] The specific implementation content of the network model training method of another reference embodiment of the present invention has been described in detail in the above-mentioned network model training method, so the repeated content will not be described again here.

[0127] According to a second aspect of an embodiment of the present invention, a method for image recognition is provided.

[0128] Figure 5 FIG. 1 is a schematic diagram of the main process of the image recognition method according to an embodiment of the present invention. Figure 5 As shown, the image recognition method according to the embodiment of the present invention mainly includes the following steps S501 to S503.

[0129] Step S501: In response to receiving a target image, extract features from the target image using a pre-set image recognition model to obtain a feature map of the target image. The image recognition model is obtained using any training method described in the first aspect of the embodiments of the present invention. Specifically, a feature extractor of the image recognition model is used to extract features from the target image to obtain a feature map of the target image.

[0130] Step S502: classify the feature map according to the image recognition model to obtain a probability distribution of the feature map. Specifically, the feature map of the target image is classified using a classifier of the image recognition model to obtain a probability distribution of the feature map.

[0131] Step S503: Determine the classification result of the target image based on the probability distribution. The probability distribution is a probability prediction of the category to which the target image belongs. For example, the probability distribution of the target image includes: the probability that the target image includes a vehicle is 0.8, the probability that the target image includes a pedestrian is 0.6, the probability that the target image includes a traffic light is 0.7, and the probability that the target image includes a guardrail is 0.3. The preset probability threshold is 0.5. If the above probabilities are greater than the probability threshold, it is determined that the target image includes the corresponding entity, and the target image is determined to include a vehicle, a pedestrian, and a traffic light.

[0132] Exemplarily, the image recognition model obtained by using any training method in the first aspect of the embodiment of the present invention is compared with the ResNet-50 model, and the above model is experimentally verified in image recognition and downstream tasks (including target detection and semantic segmentation). The experimental results are shown in Table 1. The model used by the execution subject of the embodiment of the present invention is named "ResNet-50+image mixing+self-distillation". The experimental results show that the execution subject of the embodiment of the present invention achieves data enhancement by mixing the original image, combines the image mixing technology with the self-distillation framework, adds the feature map self-distillation module and the probability distribution self-distillation module, so that the neural network can effectively mine the knowledge between samples, and has consistency in the intermediate feature layer of the feature extractor and the probability distribution space of the classifier. Compared with traditional neural networks and traditional self-distillation methods, the embodiment of the present invention can more effectively guide the self-learning of the neural network, and at the same time, it also achieves better performance than traditional neural networks and traditional self-distillation methods in downstream target detection and semantic segmentation tasks.

[0133] Table 1

[0134]

[0135] According to a third aspect of an embodiment of the present invention, a network model training device is provided.

[0136] Figure 6 Schematic diagram of the main modules of the training device of the network model according to an embodiment of the present invention, Figure 6 As shown, the network model training device 600 mainly includes:

[0137] A mixing module 601 is configured to, in response to receiving a plurality of original images, mix the plurality of original images to obtain a mixed image;

[0138] A feature module 602 is configured to extract features from the plurality of original images and the mixed image according to a preset neural network to obtain a plurality of feature maps;

[0139] A probability module 603 is configured to classify the plurality of feature maps according to the neural network to obtain a probability distribution;

[0140] The updating module 604 is used to determine the loss between the multiple feature maps and the probability distribution according to a preset loss function, perform backpropagation on the neural network according to the loss, and update the weights of the neural network.

[0141] According to a reference embodiment of the present invention, the original image includes: image information and label information; mixing the multiple original images to obtain a mixed image includes:

[0142] Mixing the image information of the plurality of original images according to a preset mixing weight;

[0143] The label information of the plurality of original images is mixed according to the mixing weight, and the mixed image information and the mixed label information are used as a mixed image.

[0144] According to another reference embodiment of the present invention, the multiple feature maps include: original feature maps and mixed feature maps; determining the loss of the multiple feature maps according to a preset loss function includes:

[0145] Determining the consistency loss between the original feature map and the hybrid feature map according to a loss function;

[0146] Image information of the original image is obtained, a linear transformation is performed on the consistency loss according to the image information, and the consistency loss after the linear transformation is used as the loss of the multiple feature maps.

[0147] According to another reference embodiment of the present invention, the probability distribution includes: an original probability distribution and a mixed probability distribution; determining the loss of the probability distribution according to a preset loss function includes:

[0148] According to the loss function, the loss between the original probability distribution and the mixed probability distribution is determined, and the loss between the original probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

[0149] According to another reference embodiment of the present invention, the original probability distribution includes: a probability distribution of each original image; and determining the loss between the original probability distribution and the mixed probability distribution according to a loss function includes:

[0150] Performing linear interpolation on the probability distributions of the multiple original images according to preset probability weights to obtain an interpolated probability distribution;

[0151] According to the loss function, a loss between the interpolated probability distribution and the mixed probability distribution is determined, and the loss between the interpolated probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

[0152] According to another reference embodiment of the present invention, the network model training device 600 further includes:

[0153] The cross-loss module is used to determine the cross-entropy loss of each original probability distribution;

[0154] The comprehensive loss module is used to perform weighted summation of the cross entropy loss, the loss of the feature map and the loss of the probability distribution according to a preset loss weight to obtain a comprehensive loss, and the comprehensive loss is used to perform backpropagation on the neural network.

[0155] It should be noted that the specific implementation content of the network model training device in the embodiment of the present invention has been described in detail in the network model training method described above, so the repeated content will not be described again here.

[0156] According to a fourth aspect of an embodiment of the present invention, a device for image recognition is provided.

[0157] Figure 7 is a schematic diagram of the main modules of the image recognition device according to an embodiment of the present invention, such as Figure 7 As shown, the image recognition device 700 mainly includes:

[0158] a feature module 701 configured to, in response to receiving a target image, extract features of the target image according to a preset image recognition model to obtain a feature map of the target image, wherein the image recognition model is obtained using any training method described in the first aspect of the embodiments of the present invention;

[0159] A probability module 702 is configured to classify the feature map according to the image recognition model to obtain a probability distribution of the feature map;

[0160] The recognition module 703 is configured to determine a classification result of the target image according to the probability distribution.

[0161] It should be noted that the specific implementation content of the image recognition device in the embodiment of the present invention has been described in detail in the image recognition method described above, so the repeated content will not be described again here.

[0162] According to the technical solution of the embodiment of the present invention, self-distillation training of the neural network is performed based on the original image and the mixed image, which can improve the training efficiency and training effect of the neural network and improve the accuracy of the trained neural network; according to the mixing weight, the image information and label information of the original image are mixed respectively, and then the mixed image information and label information are combined to obtain a mixed image, which can improve the flexibility of image mixing, provide more input samples for the neural network, and facilitate the neural network to mine more knowledge during the training process, thereby improving the accuracy of the neural network; determining the consistency loss of the original feature map and the mixed feature map is equivalent to self-distilling the feature map, which can enable the neural network to learn more semantic features and guide the neural network to produce consistent outputs for the original image and the mixed image; using the KL divergence loss function to determine the loss between the original probability distribution and the mixed probability distribution is equivalent to self-distilling the probability distribution, which can enable the neural network to have robust mixed prediction capabilities; using the trained neural network as an image recognition model to perform image recognition on the target image can improve image recognition efficiency and recognition accuracy.

[0163] According to the fifth aspect of an embodiment of the present invention, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the first aspect of the embodiment of the present invention and / or the second aspect of the embodiment of the present invention.

[0164] According to a sixth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method provided by the first aspect of the embodiment of the present invention and / or the second aspect of the embodiment of the present invention is implemented.

[0165] Figure 8 An exemplary system architecture 800 is shown to which the network model training method or the network model training apparatus according to the embodiment of the present invention can be applied.

[0166] like Figure 8 As shown, system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. Network 804 is used to provide a medium for communication links between terminal devices 801, 802, 803 and server 805. Network 804 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0167] Users can use terminal devices 801, 802, and 803 to interact with server 805 via network 804 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 801, 802, and 803, such as model training applications, image recognition applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0168] The terminal devices 801 , 802 , and 803 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0169] Server 805 may be a server that provides various services, such as a background management server (for example only) that supports training requests for network models sent from upstream terminal devices 801, 802, and 803. In response to receiving multiple original images, the background management server may mix the multiple original images to obtain a mixed image; perform feature extraction on the multiple original images and the mixed image according to a preset neural network to obtain multiple feature maps; classify the multiple feature maps according to the neural network to obtain a probability distribution; determine the loss between the multiple feature maps and the probability distribution according to a preset loss function, perform backpropagation on the neural network according to the loss, and update the weights of the neural network; and feed back the training status of the network model (for example only) to the terminal device. The background management server can also, in response to receiving the target image, extract features of the target image according to a preset image recognition model to obtain a feature map of the target image, where the image recognition model is obtained using any training method described in the first aspect of the embodiment of the present invention; classify the feature map according to the image recognition model to obtain a probability distribution of the feature map; determine the classification result of the target image according to the probability distribution; and feed back the image recognition situation (only as an example) to the terminal device.

[0170] It should be noted that the network model training method provided in the embodiment of the present invention is generally executed by the server 805, and accordingly, the network model training device is generally set in the server 805. The network model training method provided in the embodiment of the present invention can also be executed by the terminal devices 801, 802, and 803, and accordingly, the network model training device can be set in the terminal devices 801, 802, and 803.

[0171] It should be noted that the image recognition method provided in the embodiment of the present invention is generally performed by the server 805, and accordingly, the image recognition apparatus is generally provided in the server 805. The image recognition method provided in the embodiment of the present invention can also be performed by the terminal devices 801, 802, and 803, and accordingly, the image recognition apparatus can be provided in the terminal devices 801, 802, and 803.

[0172] It should be understood that Figure 8 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0173] Reference below Figure 9 , which shows a schematic structural diagram of a computer system 900 of a terminal device suitable for implementing an embodiment of the present invention. Figure 9 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0174] like Figure 9 As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the system 900 are also stored in the RAM 903. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0175] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, and the like; an output section 907 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 908 including a hard disk and the like; and a communication section 909 including a network interface card such as a LAN card or a modem. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 910 as needed, so that computer programs read therefrom can be installed into the storage section 908 as needed.

[0176] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-mentioned functions defined in the system of the embodiment of the present invention are executed.

[0177] It should be noted that the computer-readable medium shown in the embodiments of the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In embodiments of the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In embodiments of the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer programs according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0179] The modules involved in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be provided in a processor. For example, they may be described as: a processor including a mixing module, a feature module, a probability module, and an update module; or a processor including a feature module, a probability module, and an identification module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, a mixing module may also be described as a "module for mixing multiple original images into a mixed image."

[0180] As another aspect, an embodiment of the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device implements the following method: in response to receiving multiple original images, the multiple original images are mixed to obtain a mixed image; according to a preset neural network, the multiple original images and the mixed image are respectively subjected to feature extraction to obtain multiple feature maps; according to the neural network, the multiple feature maps are respectively classified to obtain a probability distribution; according to a preset loss function, the loss between the multiple feature maps and the probability distribution is determined, and the neural network is back-propagated according to the loss to update the weights of the neural network. Alternatively, the device implements the following method: in response to receiving a target image, extracting features of the target image according to a preset image recognition model to obtain a feature map of the target image, wherein the image recognition model is obtained by using any training method described in the first aspect of the embodiment of the present invention; classifying the feature map according to the image recognition model to obtain a probability distribution of the feature map; and determining the classification result of the target image according to the probability distribution.

[0181] According to the technical solution of the embodiment of the present invention, self-distillation training of the neural network is performed based on the original image and the mixed image, which can improve the training efficiency and training effect of the neural network and improve the accuracy of the trained neural network; according to the mixing weight, the image information and label information of the original image are mixed respectively, and then the mixed image information and label information are combined to obtain a mixed image, which can improve the flexibility of image mixing, provide more input samples for the neural network, and facilitate the neural network to mine more knowledge during the training process, thereby improving the accuracy of the neural network; determining the consistency loss of the original feature map and the mixed feature map is equivalent to self-distilling the feature map, which can enable the neural network to learn more semantic features and guide the neural network to produce consistent outputs for the original image and the mixed image; using the KL divergence loss function to determine the loss between the original probability distribution and the mixed probability distribution is equivalent to self-distilling the probability distribution, which can enable the neural network to have robust mixed prediction capabilities; using the trained neural network as an image recognition model to perform image recognition on the target image can improve image recognition efficiency and recognition accuracy.

[0182] The above specific embodiments do not limit the scope of protection of the embodiments of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of the embodiments of the present invention.

Claims

1. A training method for a network model, characterized in that: include: In response to receiving a plurality of original images, mixing the plurality of original images to obtain a mixed image; According to a preset neural network, feature extraction is performed on the multiple original images and the mixed image respectively to obtain multiple feature maps; Classifying the plurality of feature maps according to the neural network to obtain a probability distribution; According to a preset loss function, the loss between the multiple feature maps and the probability distribution is determined, and the neural network is back-propagated according to the loss to update the weight of the neural network.

2. The method according to claim 1, characterized in that The original image includes: image information and label information; the multiple original images are mixed to obtain a mixed image, including: Mixing the image information of the plurality of original images according to a preset mixing weight; The label information of the plurality of original images is mixed according to the mixing weight, and the mixed image information and the mixed label information are used as a mixed image.

3. The method according to claim 1, characterized in that The plurality of feature maps include: original feature maps and mixed feature maps; and determining the loss of the plurality of feature maps according to a preset loss function includes: Determining the consistency loss between the original feature map and the hybrid feature map according to a loss function; Image information of the original image is obtained, a linear transformation is performed on the consistency loss according to the image information, and the consistency loss after the linear transformation is used as the loss of the multiple feature maps.

4. The method according to claim 1, wherein The probability distribution includes: an original probability distribution and a mixed probability distribution; and determining the loss of the probability distribution according to a preset loss function includes: According to the loss function, the loss between the original probability distribution and the mixed probability distribution is determined, and the loss between the original probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

5. The method according to claim 4, characterized in that The original probability distribution includes: a probability distribution of each original image; and determining the loss between the original probability distribution and the mixed probability distribution according to a loss function, including: Performing linear interpolation on the probability distributions of the multiple original images according to preset probability weights to obtain an interpolated probability distribution; According to the loss function, a loss between the interpolated probability distribution and the mixed probability distribution is determined, and the loss between the interpolated probability distribution and the mixed probability distribution is used as the loss of the probability distribution.

6. The method according to claim 5, characterized in that Before backpropagating the neural network according to the loss, the method further includes: Determine the cross entropy loss for each original probability distribution; According to the preset loss weights, the cross entropy loss, the loss of the feature map and the loss of the probability distribution are weighted and summed to obtain a comprehensive loss, which is used to perform backpropagation on the neural network.

7. A method for image recognition, characterized in that: include: In response to receiving a target image, extracting features of the target image according to a preset image recognition model to obtain a feature map of the target image, wherein the image recognition model is obtained by using the training method according to any one of claims 1 to 6; Classifying the feature map according to the image recognition model to obtain a probability distribution of the feature map; A classification result of the target image is determined according to the probability distribution.

8. A network model training device, characterized in that: include: a mixing module, configured to, in response to receiving a plurality of original images, mix the plurality of original images to obtain a mixed image; A feature module, configured to extract features from the plurality of original images and the mixed image according to a preset neural network to obtain a plurality of feature maps; A probability module, configured to classify the plurality of feature maps according to the neural network to obtain a probability distribution; An updating module is used to determine the loss between the multiple feature maps and the probability distribution according to a preset loss function, perform backpropagation on the neural network according to the loss, and update the weights of the neural network.

9. An image recognition device, characterized in that: include: a feature module, configured to, in response to receiving a target image, extract features of the target image according to a preset image recognition model to obtain a feature map of the target image, wherein the image recognition model is obtained using the training method according to any one of claims 1 to 6; A probability module, configured to classify the feature map according to the image recognition model to obtain a probability distribution of the feature map; A recognition module is used to determine a classification result of the target image according to the probability distribution.

10. An electronic device, characterized in that: include: One or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

11. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.