Model training method and device

By extracting joint features and/or separating features from the semantic segmentation knowledge distillation method, the problem of insufficient expression ability of the model for niche categories and image details is solved, and stronger expression ability and recognition ability of small-scale targets are achieved.

CN114627331BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210223406.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-05-23
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The existing semantic segmentation knowledge distillation method can easily lead to weak expression of niche categories and image details during model training, especially under the influence of main categories, it is difficult for the model to effectively capture the characteristics of small-scale targets.

Method used

The expression of the model is enhanced during knowledge distillation by extracting joint features and/or separating features. The specific method includes converting the feature map output by the subject network into a joint feature of the category, and dividing it into separate features in the height and width dimensions, respectively entering the generalized normalization layer for alignment, to construct the loss function of the student model.

Benefits of technology

It effectively enhances the model's ability to express niche categories and image details, avoids the impact of the main categories on niche categories, improves the ability to express small-scale goals, and avoids the neglect problems that the model is prone to in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627331B_ABST
    Figure CN114627331B_ABST
Patent Text Reader

Abstract

The present invention discloses a model training method and device, and relates to the field of artificial intelligence technology. A specific implementation of the method includes: inputting a plurality of training images pre-labeled with semantic segmentation labels in a sample set into a trained teacher model and a student model to be trained respectively, and determining the probability distribution difference between the prediction result of the student model for the category to which the pixel in the training image belongs and the semantic segmentation label as the first difference; and using the first difference combined with the second difference and / or the third difference to construct the loss function of the student model to train the student model. This implementation can enhance the model's ability to express niche categories and / or image detail information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model training method and device. Background Art

[0002] Semantic segmentation is one of the key issues in today's computer vision field. It is widely used in autonomous driving, virtual reality, intelligent diagnosis and treatment, remote sensing and other fields. It can achieve a complete understanding of the scene by reasoning and predicting the category to which each pixel belongs. Since the semantic segmentation task requires understanding complex scenes at the pixel level, it often requires a larger and more complex model to learn powerful feature representation capabilities to ensure prediction accuracy and make the model have good generalization. Due to the large size of the model and high computational cost, it is easy to cause problems such as high resource usage and slow response speed, and it is not suitable for deployment on terminal devices.

[0003] At present, the above problems can be solved by using the knowledge distillation method, that is, a lightweight student model is trained through a complex teacher model, and the lightweight student model is deployed on the terminal device. In the actual semantic segmentation task, this method has the following problems: First, affected by the main category (the category with a large proportion of pixels, the category refers to the category to which the pixel belongs, that is, the category contained in the label), the model has a weak ability to express the niche category (the category with a small proportion of pixels); second, the model has a weak ability to express local information and detail information in the image, especially when the target of interest only occupies a small area of ​​the image and the background occupies a large area, it is easy to ignore the target. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a model training method and device, which can enhance the model's ability to express niche categories and / or image detail information by extracting joint features and / or separation features during the knowledge distillation process.

[0005] To achieve the above objective, according to one aspect of the present invention, a model training method is provided.

[0006] The model training method of the embodiment of the present invention comprises: inputting a plurality of training images pre-labeled with semantic segmentation labels in a sample set into a trained teacher model and a student model to be trained respectively, and determining the probability distribution difference between the prediction result of the student model on the category to which the pixel in the training image belongs and the semantic segmentation label as a first difference; and, both the student model and the teacher model include a main network and a generalized normalization layer connected after the main network; for the feature maps corresponding to the plurality of training images output by the main network of the student model and the teacher model: converting them into joint features of the category and entering the generalized normalization layer, and / or, The image is segmented into multiple separate features in height and width dimensions based on preset segmentation rules and then enters the generalized normalization layer; wherein the joint feature of each category includes probability data that pixels in the feature maps corresponding to the multiple training images belong to the category; each separate feature includes probability data that pixels in the feature map in the same segmentation space belong to the category; the first difference is combined with the second difference and / or the third difference to construct the loss function of the student model to train the student model; wherein the second difference is determined based on the joint feature of the student model and the teacher model, and the third difference is determined based on the separate features of the student model and the teacher model.

[0007] Optionally, the joint features of any category are determined according to the following steps: obtaining probability data of each pixel belonging to the category in a multi-channel feature map corresponding to the multiple training images output by the corresponding main network; and merging the probability data of each pixel belonging to the category into the joint features of the category.

[0008] Optionally, the separation feature is further formed by dividing the feature map in the channel dimension and aggregating it in the category dimension; any segmentation space formed by segmentation in the channel, height and width dimensions corresponding to the separation feature of any category includes: probability data that the pixels in the segmentation space belong to the category.

[0009] Optionally, when the student model forms a joint feature, the teacher model forms a joint feature; when the teacher model forms a joint feature, the student model forms a joint feature; when the student model forms a separate feature, the teacher model forms a separate feature; when the teacher model forms a separate feature, the student model forms a separate feature; and, the use of the first difference combined with the second difference and / or the third difference to construct the loss function of the student model includes: determining the weighted sum of the first difference and the second difference as the loss function; or, determining the weighted sum of the first difference and the third difference as the loss function; or, determining the weighted sum of the first difference, the second difference and the third difference as the loss function.

[0010] Optionally, after the joint feature of each category enters the generalized normalization layer, normalization is performed internally on the joint feature to form the first normalized feature of the category; and the second difference is determined according to the following steps: calculating the KL divergence of the first normalized features of the student model and the teacher model corresponding to the same category; and determining the average value of the KL divergence of each category as the second difference.

[0011] Optionally, after entering the generalized normalization layer, any separation feature of any segmented space formed by segmentation in the channel, height and width dimensions corresponding to any category is normalized internally to form a second normalized feature of the segmented space and the category; and the third difference is determined according to the following steps: calculating the KL divergence of the second normalized features of the segmented space at the same position and the same category of the student model and the teacher model; and determining the average value of the KL divergence of the segmented space at each position and each category as the third difference.

[0012] Optionally, the feature map of the student model enters a narrow normalization layer for calculation, and the prediction result is determined based on the calculation result of the narrow normalization layer; and the narrow normalization layer includes a Softmax layer with a temperature parameter equal to 1, and the generalized normalization layer includes a Softmax layer with a temperature parameter not equal to 1.

[0013] To achieve the above objective, according to another aspect of the present invention, a model training device is provided.

[0014] The model training device of the embodiment of the present invention may include: a supervised training unit, which is used to: input a plurality of training images pre-labeled with semantic segmentation labels in a sample set into a trained teacher model and a student model to be trained respectively, and determine the probability distribution difference between the prediction result of the student model for the category to which the pixel in the training image belongs and the semantic segmentation label as a first difference; and the student model and the teacher model both include a main network and a generalized normalization layer connected after the main network; for the feature maps corresponding to the plurality of training images output by the main network of the student model and the teacher model: convert them into joint features of the category and enter the generalized normalization layer, and / or , after being segmented into multiple separation features in height and width dimensions based on preset segmentation rules, entering the generalized normalization layer; wherein the joint feature of each category includes probability data that pixels in the feature maps corresponding to the multiple training images belong to the category; each separation feature includes probability data that pixels in the feature map in the same segmentation space belong to the category; a distillation training unit is used to: use the first difference combined with the second difference and / or the third difference to construct the loss function of the student model to train the student model; wherein the second difference is determined based on the joint feature of the student model and the teacher model, and the third difference is determined based on the separation feature of the student model and the teacher model.

[0015] Optionally, the joint features of any category are determined according to the following steps: obtaining probability data of each pixel in the multi-channel feature map corresponding to the multiple training images output by the corresponding main network belonging to the category; merging the probability data of each pixel belonging to the category into the joint features of the category; the separation feature is further formed by segmenting the feature map in the channel dimension and aggregating the category dimension; the separation feature corresponding to any category of any segmented space formed by segmentation in the channel, height and width dimensions includes: the probability data of the pixels in the segmented space belonging to the category; when the student model forms a joint feature, the teacher model forms a joint feature; when the teacher model forms a joint feature, the student model forms a joint feature; when the student model forms a separation feature, the teacher model forms a separation feature; when the teacher model forms a separation feature, the student model forms a separation feature; and the distillation training unit is further used to: determine the weighted sum of the first difference and the second difference as the loss function; or, determine the weighted sum of the first difference and the third difference as the loss function; or, determine the weighted sum of the first difference, the second difference and the third difference as the loss function.

[0016] To achieve the above objective, according to another aspect of the present invention, an electronic device is provided.

[0017] An electronic device of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method provided by the present invention.

[0018] To achieve the above objective, according to yet another aspect of the present invention, a computer-readable storage medium is provided.

[0019] A computer-readable storage medium of the present invention stores a computer program, which, when executed by a processor, implements the model training method provided by the present invention.

[0020] According to the technical solution of the present invention, the embodiments of the above invention have the following advantages or beneficial effects:

[0021] In the model training process based on knowledge distillation, the first path after the main network, and the second path and / or the third path are established for the student model, and the fourth path and / or the fifth path after the main network is established for the teacher model, wherein the second path and the fourth path are used at the same time or not at the same time, and a second difference is formed when used at the same time; the third path and the fifth path are used at the same time or not at the same time, and a third difference is formed when used at the same time; the loss function of the student model can be obtained by the first difference between the output result of the first path and the label combined with the second difference and / or the third difference.

[0022] When using the second and fourth paths, in the student model and the teacher model, the feature map output by the main network can be converted into a joint feature of each category, and then input into the normalization layer and then aligned (alignment refers to the use of cross entropy, KL divergence and other functions to calculate the probability distribution difference). The joint feature is composed of the probability data of each pixel in the feature map corresponding to each image in the sample set for the same category, that is, it represents the cross-image merging of pixel probabilities in the category dimension. In this way, by constructing joint features for different categories, niche categories have independent feature expression channels, which is conducive to avoiding the influence of the main category on the niche category in the same image, and strengthening the expression ability of the student model for the niche category, thereby avoiding the defect of the existing semantic segmentation knowledge distillation method that easily weakens the niche category.

[0023] When using the third path and the fifth path, in the student model and the teacher model, the feature map output by the main network can be divided into multiple separation features in the height and width dimensions according to the consistent segmentation rules, and then input into the normalization layer and then aligned. The above separation features can reflect the local information and detail information of the training image. The alignment method based on the separation features can enable the student model to enhance the expression ability of image details and small-range targets, especially when the target occupies a smaller range in the image and the background occupies a larger range. The above alignment method based on the separation features is likely to make the small-range target have an independent expression channel in the model, thereby minimizing the possibility of being ignored.

[0024] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.

[0026] Figure 1 Schematic diagram of the main steps of the model training method in an embodiment of the present invention;

[0027] Figure 2 is a schematic diagram of the structure of a teacher model and a student model in the training process of an embodiment of the present invention;

[0028] Figure 3 It is a schematic diagram of specific use steps of the teacher model and the student model of an embodiment of the present invention;

[0029] Figure 4 It is a schematic diagram of components of a model training device in an embodiment of the present invention;

[0030] Figure 5 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;

[0031] Figure 6 It is a schematic diagram of the structure of an electronic device used to implement the model training method in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0033] Semantic segmentation is a basic topic in computer vision, which aims to assign a unique category to each pixel in the input image, such as people, sky, grass, vehicles, etc. Since the semantic segmentation task needs to understand complex scenes at the pixel level, it often requires larger-scale complex models to learn powerful feature representation capabilities to ensure prediction accuracy and make the model have good generalization. Due to the large model size and high computational cost, it is easy to cause problems such as high resource usage and slow response speed. At the same time, it is not suitable for deployment on terminal devices. At present, it can be solved through the knowledge distillation method, and a lightweight student model (StudentModel) is trained through a complex and trained teacher model (Teacher Model).

[0034] Knowledge distillation is a model compression method and a training method based on "teacher-student". It distills the knowledge contained in the trained teacher model into the student model. The idea is: first establish a teacher model with a relatively complex structure and large size and a student model with a lightweight, small number of parameters and relatively simple structure. After that, the teacher model completes the learning of training samples labeled with semantic segmentation labels (which can be manually labeled), and the student model simultaneously learns the above training samples and the output results of the teacher model (the loss function is the sum of these two parts), thereby transferring the knowledge expression of the teacher model to the student model, and finally deploying the student model online.

[0035] Knowledge distillation involves the temperature parameter T, which is a parameter of the generalized normalization layer (such as the generalized Softmax function). The generalized Softmax function is as follows:

[0036]

[0037] Among them, z i is the i-th component of the input vector, q i Yes i The output of , j represents the sequence number of any component in the input vector.

[0038] It can be seen that the traditional Softmax function is a special case of T=1. The higher T is, the smoother the probability distribution of the function output is, the greater the entropy of its distribution, the information carried by the negative label (that is, labels other than the positive label) will be amplified, and the model training will pay more attention to the negative label. In the actual training process, the temperature parameter T of the generalized normalization layer in the student model and the teacher model can be increased when the student model learns the output result of the teacher model, and T can be reduced after the training is completed. In an embodiment of the present invention, the narrow normalization layer includes a Softmax layer with a temperature parameter equal to 1 (using the traditional Softmax function), and the generalized normalization layer includes a Softmax layer with a temperature parameter T not equal to 1 (generally using the above generalized Softmax function with T greater than 1), and the two are collectively referred to as normalization layers.

[0039] In the actual semantic segmentation task, this method has the following problems: first, affected by the main categories, the model has a weak ability to express the minority categories; second, the model has a weak ability to express the local information and detail information in the image, especially when the target of interest only occupies a small area of ​​the image and the background occupies a large area, the target is easily ignored. The present invention can solve this problem through the following technical solutions: the first problem is solved by knowledge expression based on joint features, and the second problem is solved by knowledge expression based on separation features.

[0040] It should be pointed out that the embodiments of the present invention and the technical features therein may be combined with each other without conflict.

[0041] Figure 1 Schematic diagram of the main steps of the model training method according to an embodiment of the present invention.

[0042] like Figure 1 As shown, the model training method of the embodiment of the present invention can be specifically performed according to the following steps:

[0043] Step S101: input multiple training images in the sample set that are pre-labeled with semantic segmentation labels into the trained teacher model and the student model to be trained respectively, and determine the probability distribution difference between the student model's prediction result of the category to which the pixels in the training image belong and the semantic segmentation label as the first difference.

[0044] In this step, the sample set refers to a sample set containing multiple training samples (image samples in the field of semantic segmentation). The number of samples contained is represented by B, which can be a batch or multiple batches. The semantic segmentation label can be a true value label (i.e., ground truth) manually annotated or a true value label obtained by other non-artificial methods. The semantic segmentation label of a training image includes the category to which each pixel belongs.

[0045] Figure 2 is a schematic diagram of the structure of the teacher model and the student model in the training process of an embodiment of the present invention, see Figure 2 , the student model includes a main network and multiple output paths connected after the main network, the teacher model includes a main network and at least one output path connected after the main network, the above main network can be the model structure of the student model and the teacher model before the normalization layer, in the field of semantic segmentation, the main network of the student model and the teacher model can be CNN (Convolutional Neural Networks, convolutional neural network) as the main body, or other applicable existing models as the main body. In the student model and the teacher model, the above main network outputs a feature map (i.e., Feature Map) backward, and its size is B*H*W, where B can also represent the number of channels, which is equal to the number of training images in the sample set, H is the height (i.e., the number of pixels in the height dimension), W is the width (i.e., the number of pixels in the width dimension), and H and W are integers not less than 1.

[0046] For example, if the size of the above feature map is 10*64*64, it is equivalent to that the above feature map contains 10 channels, each channel is a 64*64 image, each image includes 64*64 pixels, and the entire feature map includes 10*64*64 pixels. It can be understood that the pixel value of each pixel is the probability data of the pixel belonging to each category (related to the probability that the pixel belongs to a certain category, but may not be between zero and 1), and is formally a vector with a length equal to the number of categories (expressed as C).

[0047] It should be noted that each channel of the above feature map corresponds one-to-one to each training image in the sample set. H and W can be equal to the height and width of the training image, respectively, or they can be smaller than the height and width of the training image to speed up the calculation. If it is the latter, when outputting the prediction result, the H*W image can be restored to the training image size through the interpolation method.

[0048] The above output path includes model structures such as a normalization layer and a model output method based on the model structure. The output path of the student model includes a first path, and the first path includes structures such as a narrow normalization layer and an output layer, and may also include a structure for restoring the image size. During the training process of the student model, the output result based on the first path is the student model's prediction result of the category to which the pixels in the training image belong. In step S101, the probability distribution difference between the output result of the student model based on the first path and the pre-labeled semantic segmentation label can be calculated by known cross entropy, KL divergence, etc., and determined as the first difference. It can be understood that the first difference indicates that the student model learns knowledge expression from real samples. Thereafter, the loss function of the student model can be constructed based on the first difference.

[0049] The output path of the student model further includes a second path and / or a third path, and the output path of the teacher model includes a fourth path and / or a fifth path. The above paths all have a generalized normalization layer. It can be understood that the above paths are used to realize the transfer of knowledge from the teacher model to the student model. In an embodiment of the present invention, there are three selection methods for the above paths: first, the second path and the fourth path are selected at the same time, and the third path and the fifth path are not selected; second, the third path and the fifth path are selected at the same time, and the second path and the fourth path are not selected; third, both the second path and the fourth path are selected, and the third path and the fifth path are selected.

[0050] The second path of the student model and the fourth path of the teacher model both correspond to knowledge expression based on joint features. In these two paths, the above feature maps output by the corresponding main network (for the second path, it is the main network of the student model, and for the fourth path, it is the main network of the teacher model) can be converted into joint features of each category and then enter the generalized normalization layer. The joint features of any category are composed of the probability data that the pixels of all channels in the feature map (that is, all pixels in the feature map) belong to that category.

[0051] That is to say, in the second path and the fourth path, after obtaining the feature map output by the main network, the probability data (i.e., pixel value) of each pixel belonging to the category can be extracted from the feature map for each category respectively and arranged in a fixed order, thereby merging into a vector, which is the joint feature of the category. Continuing with the above example, for a feature map of size 10*64*64, if the number of categories C is 19, the pixel value of each pixel is a vector of length 19 (each component corresponds to category 1 to category 19 in turn). When calculating the joint feature of a certain category (category 5), each pixel of each channel image is traversed in a preset fixed order based on channel, height, and width, and the probability data corresponding to category 5 in the pixel value is extracted. Each probability data is merged according to the fixed arrangement position of the corresponding pixel, thereby forming the joint feature of category 5. The joint features of each category then enter the generalized normalization layer connected to the back, and normalization is performed inside the joint features to form the first normalized features of each category. Finally, the probability distribution difference (called the second difference) between the output result of the student model based on the second path and the output result of the teacher model based on the fourth path can be calculated by methods such as KL divergence, and the second difference can be selected to construct the loss function of the student model. Exemplarily, the KL divergence of the first normalized feature of the same category corresponding to the student model and the teacher model can be calculated, and the average value of the KL divergence of each category is determined as the second difference.

[0052] The calculation process of the generalized normalization layer of the student network for the joint features is as follows:

[0053] U i =(u i,1 ,u i,2 ,…,u i,j ,…,u i,B×H×W )

[0054]

[0055] Among them, U i is the first normalized feature of category i in the student network, u i,j It's U i The jth component in is the jth component of the joint features of category i in the student network, is the kth component of the joint feature of category i in the student network (indicating traversal of each component), T u is the temperature parameter of the generalized normalization layer of the second path.

[0056] The generalized normalization layer of the teacher network calculates the joint features as follows:

[0057] V i =(v i,1 ,v i,2 ,…,v i,j ,…,v i,B×H×W )

[0058]

[0059] Among them, V i is the first normalized feature of category i in the teacher network, v i,j Yes V i The jth component in is the jth component of the joint features of category i in the teacher network, is the kth component of the joint feature of category i in the teacher network (indicating traversal of each component), T v is the temperature parameter of the generalized normalized layer of the fourth path, T v Can be equal to T u , which may not be equal to T u .

[0060] The second difference can be calculated as follows:

[0061]

[0062] Among them, L U denotes the second difference, and KL(||) denotes the KL divergence function.

[0063] If the second difference is selected to construct the loss function of the student model, the cross-image aggregation of the probability data of each pixel in the category dimension can be achieved through the above joint feature-based knowledge expression. In this way, the niche category has an independent feature expression channel (i.e., an independent joint feature), and there are no other categories in its independent channel, which is conducive to avoiding the influence of the main category on the niche category in the same image, strengthening the expression ability of the student model for the niche category, and thus avoiding the defect of the existing semantic segmentation knowledge distillation method that easily weakens the niche category. In particular, in some training images, some niche categories may not appear. Using such training images will seriously affect the prediction of the niche category, and the use of the joint feature knowledge expression of the present invention can solve this problem through the cross-image combination of the niche category.

[0064] The third path of the student model and the fifth path of the teacher model both correspond to knowledge expression based on joint features. In these two paths, the above feature maps output by the corresponding main network (the main network of the student model for the third path and the main network of the teacher model for the fifth path) are divided into multiple separation features in the height and width dimensions based on the preset segmentation rules and then enter the generalized normalization layer. Each separation feature includes the probability data of pixels belonging to the above categories in the same segmentation space (which can be a two-dimensional space of height and width dimensions, or a three-dimensional space of height, width and channel dimensions). The above segmentation rules can be executed in the height and width dimensions, and can also be further executed in the channel and category dimensions. The segmentation in each dimension is an average segmentation. The following describes several segmentation rules:

[0065] First, the division is performed only in the height dimension and the width dimension, such as 2×2 (the height and width are equally divided into 2) or 4×4 (the height and width are equally divided into 4). Taking 2×2 as an example, for a feature map of size 10*64*64, 4 division spaces are formed, and the size of each division space is 10*32*32. Then each division space corresponds to a separation feature, and each separation feature is the probability data of each pixel in the corresponding division space belonging to each category, that is, if the separation feature is flattened, its length is 10*32*32*19, and the total number of separation features is 4.

[0066] Second, the height dimension, width dimension and channel dimension are divided. For example, the height and width are 2×2 as before, and the channel is divided into 10 (that is, divided according to different images). For a feature map with a size of 10*64*64, 2×2×10 segmentation spaces are formed. The size of each segmentation space is 32*32, and each separation feature is the probability data of each pixel in the corresponding segmentation space belonging to each category. That is, if the separation feature is flattened, its length is 32*32*19, and the total number of separation features is 2×2×10.

[0067] Third, the image is split in the height, width and channel dimensions, and aggregated in the category dimension. For example, if the height and width are 2×2 as before, and the channel is split into 10 as before, and aggregated according to each category, then for a feature map of size 10*64*64, 2×2×10 split spaces are formed, and the size of each split space is 32*32. After aggregating to each category, the separation feature corresponding to any split space and any category is the probability data of each pixel in the split space belonging to the category, that is, if the separation feature is flattened, its length is 32*32, and the total number of separation features is 2×2×10×19.

[0068] It should be noted that the student model and the teacher model need to use consistent segmentation rules to ensure data consistency. The above segmentation can also be reused. For example, the height dimension and the width dimension can use 2×2 segmentation on the one hand and 4×4 segmentation on the other hand, and finally merge the data from both aspects. Of course, the student model and the teacher model also need to ensure consistency when reused.

[0069] After each separation feature is formed, it enters the generalized normalization layer connected to the rear to perform normalization inside the separation feature, and each separation feature forms a second normalized feature. Taking the third segmentation rule as an example, any segmentation space formed by segmentation of the channel, height and width dimensions corresponding to the separation feature of any category performs normalization inside the separation feature after entering the generalized normalization layer to form the segmentation space and the second normalized feature of the category. Finally, the probability distribution difference (called the third difference) between the output result of the student model based on the third path and the output result of the teacher model based on the fifth path can be calculated by methods such as KL divergence, and the third difference can be selected to construct the loss function of the student model. Exemplarily, the KL divergence of the second normalized features of the segmentation space and the same category of the student model and the teacher model corresponding to the same position can be calculated first, and then the average value of the KL divergence of the segmentation space at each position and each category is determined as the third difference.

[0070] Taking the third segmentation rule as an example, the calculation process of the generalized normalization layer of the student network for the separation feature is as follows:

[0071]

[0072]

[0073] Where m represents the number of divisions of the width dimension and the height dimension (i.e., divided into m parts). In this example, the width and height of the feature map are divided into the same number; D x is the second normalized feature with sequence number x in the student network, and the maximum value of x is B×C×m 2 ;d x,j Yes Dx The jth component in , the total number of components in a second normalized feature is is the jth component of the separation feature with sequence number x in the student network, is the kth component of the separation feature (indicating traversal of each component), T d is the temperature parameter of the generalized normalization layer of the third path.

[0074] The calculation process of the generalized normalization layer of the teacher network for the separation feature is as follows:

[0075]

[0076]

[0077] Among them, F x is the second normalized feature with sequence number x in the teacher network, f x,j Yes F x The jth component in is the jth component of the separation feature with sequence number x in the teacher network, is the kth component of the separation feature (indicating traversal of each component), T f is the temperature parameter of the generalized normalized layer of the fifth path, T f Can be equal to T d , which may not be equal to T d .

[0078] The third difference L D The calculation can be as follows:

[0079]

[0080] If the third difference is selected to construct the loss function of the student model, the local information and detail information of the training image can be reflected through the above knowledge expression based on the separation feature. For example, if the target of interest only occupies 20% or less in the original image, it may occupy more than 80% of the segmented image after segmentation. Therefore, its target information can be fully reflected in the separation features formed thereafter (equivalent to adding an independent model expression channel for small-range targets), avoiding being overwhelmed by the large-range background, thereby greatly enhancing the student model's ability to express image detail information and local information.

[0081] Step S102: construct a loss function of the student model using the first difference in combination with the second difference and / or the third difference.

[0082] As mentioned above, three selection methods can be used to select the path: first, select the second path and the fourth path at the same time, and do not select the third path and the fifth path; second, select the third path and the fifth path at the same time, and do not select the second path and the fourth path; third, select both the second path and the fourth path, and the third path and the fifth path. Under the first selection method, the loss function can be constructed using the second difference based on the second path and the fourth path and the first difference; under the second selection method, the loss function can be constructed using the third difference based on the third path and the fifth path and the first difference; under the third selection method, the loss function can be constructed using the first difference, the second difference and the third difference. In practical applications, the loss function is generally constructed by weighted sum.

[0083] Figure 3 is a schematic diagram of specific use steps of the teacher model and the student model of the embodiment of the present invention, see Figure 3 In step S301, a supervised method is used to train the teacher model through a sample set. In step S302, knowledge expression based on joint features is added to the student model structure. In step S303, knowledge expression based on separation features is added to the student model structure. One or both of step S301 and step S302 can be selected. In step S304, the temperature parameter T of each generalized normalization layer of the student model and the teacher model is increased for training as needed. In step S305, at the end of training, only one output path of the student model can be used. If the retained output path contains a generalized normalization layer, its temperature parameter needs to be set to a fixed value. In step S306, the lightweight student model is deployed on a device with low computing resources (such as a terminal device). In this way, the requirements of the terminal device for the amount of model parameters can be met, and the deployed student model can have the knowledge expression ability of the teacher model.

[0084] In the technical solution of the embodiment of the present invention, the knowledge expression based on joint features and separation features can be used for knowledge distillation from super-large models to lightweight models in the field of semantic segmentation, and has clear physical meaning and strong interpretability. Specifically, the knowledge expression of joint features aggregates multiple training images from the category dimension to enhance the student model's ability to express niche categories; the knowledge expression of separation features cuts the feature graphs into pieces and aligns them separately to enhance the student model's ability to express local information and detail information in the image, thus solving the long-standing pain points in the field of semantic segmentation knowledge distillation.

[0085] It should be noted that, for the convenience of description, the aforementioned method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the described action sequence, and some steps can actually be performed in other sequences or simultaneously. In addition, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary to implement the present invention.

[0086] In order to better implement the above-mentioned solution of the embodiment of the present invention, relevant devices for implementing the above-mentioned solution are also provided below.

[0087] See also Figure 4 As shown, the model training device 400 provided in this embodiment of the present invention may include: a supervised training unit 401 and a distillation training unit 402.

[0088] The supervised training unit 401 can be used to: input multiple training images pre-labeled with semantic segmentation labels in a sample set into a trained teacher model and a student model to be trained respectively, and determine the probability distribution difference between the prediction result of the student model on the category to which the pixels in the training image belong and the semantic segmentation label as the first difference; and, both the student model and the teacher model include a main network and a generalized normalization layer connected after the main network; for the feature maps output by the main network of the student model and the teacher model and corresponding to the multiple training images: enter the generalized normalization layer after being converted into a joint feature of the category, and / or enter the generalized normalization layer after being segmented into multiple separate features in height and width dimensions based on a preset segmentation rule; wherein the joint feature of each category includes probability data that the pixels in the feature map corresponding to the multiple training images belong to the category; and each separate feature includes probability data that the pixels in the feature map in the same segmentation space belong to the category.

[0089] The distillation training unit 402 can be used to: use the first difference combined with the second difference and / or the third difference to construct the loss function of the student model to train the student model; wherein the second difference is determined based on the joint features of the student model and the teacher model, and the third difference is determined based on the separated features of the student model and the teacher model.

[0090] In an embodiment of the present invention, the joint feature of any category is determined according to the following steps: obtaining probability data of each pixel in the multi-channel feature map corresponding to the multiple training images output by the corresponding main network belonging to the category; merging the probability data of each pixel belonging to the category into the joint feature of the category; the separation feature is further formed by segmenting the feature map in the channel dimension and aggregating the category dimension; the separation feature corresponding to any category of any segmented space formed by segmenting in the channel, height and width dimensions includes: the probability data of the pixels in the segmented space belonging to the category; when the student model forms a joint feature, the teacher model forms a joint feature; when the teacher model forms a joint feature, the student model forms a joint feature; when the student model forms a separation feature, the teacher model forms a separation feature; when the teacher model forms a separation feature, the student model forms a separation feature; and the distillation training unit 402 can be further used to: determine the weighted sum of the first difference and the second difference as the loss function; or, determine the weighted sum of the first difference and the third difference as the loss function; or, determine the weighted sum of the first difference, the second difference and the third difference as the loss function.

[0091] As a preferred solution, after the joint feature of each category enters the generalized normalization layer, normalization is performed internally on the joint feature to form the first normalized feature of the category; and the distillation training unit 402 can be further used to: calculate the KL divergence of the first normalized features of the student model and the teacher model corresponding to the same category; and determine the average value of the KL divergence of each category as the second difference.

[0092] Preferably, after entering the generalized normalization layer, any separation feature of any segmented space formed by segmentation in the channel, height and width dimensions corresponding to any category is normalized internally to form a second normalized feature of the segmented space and the category; and the distillation training unit 402 can be further used to: calculate the KL divergence of the second normalized features of the segmented space at the same position and the same category of the student model and the teacher model; and determine the average value of the KL divergence of the segmented space at each position and each category as the third difference.

[0093] In addition, in an embodiment of the present invention, the feature map of the student model enters the narrow normalization layer for calculation, and the prediction result is determined based on the calculation result of the narrow normalization layer; the narrow normalization layer includes a Softmax layer with a temperature parameter equal to 1, and the generalized normalization layer includes a Softmax layer with a temperature parameter not equal to 1.

[0094] According to the technical solution of an embodiment of the present invention, in a model training process based on knowledge distillation, a first path after the main network, and a second path and / or a third path are established for a student model, and a fourth path and / or a fifth path after the main network is established for a teacher model, wherein the second path and the fourth path are used at the same time or not at the same time, and a second difference is formed when used at the same time; the third path and the fifth path are used at the same time or not at the same time, and a third difference is formed when used at the same time; the loss function of the student model can be obtained by combining the first difference between the output result of the first path and the label with the second difference and / or the third difference.

[0095] When using the second and fourth paths, in the student model and the teacher model, the feature map output by the main network can be converted into a joint feature of each category, and then input into the normalization layer and aligned. The joint feature is composed of the probability data of each pixel in the feature map corresponding to each image in the sample set for the same category, that is, it represents the cross-image merging of pixel probabilities in the category dimension. In this way, by constructing joint features for different categories, niche categories have independent feature expression channels, which is conducive to avoiding the influence of the main category on the niche category in the same image, and strengthening the expression ability of the student model for the niche category, thereby avoiding the defect of the existing semantic segmentation knowledge distillation method that easily weakens the niche category.

[0096] When using the third path and the fifth path, in the student model and the teacher model, the feature map output by the main network can be divided into multiple separation features in the height and width dimensions according to the consistent segmentation rules, and then input into the normalization layer and then aligned. The above separation features can reflect the local information and detail information of the training image. The alignment method based on the separation features can enable the student model to enhance the expression ability of image details and small-range targets, especially when the target occupies a smaller range in the image and the background occupies a larger range. The above alignment method based on the separation features is likely to make the small-range target have an independent expression channel in the model, thereby minimizing the possibility of being ignored.

[0097] Figure 5 An exemplary system architecture 500 is shown to which the model training method or model training device according to an embodiment of the present invention can be applied.

[0098] like Figure 5 As shown, the system architecture 500 may include terminal devices 501, 502, 503, a network 504 and a server 505 (this architecture is only an example, and the components included in the specific architecture may be adjusted according to the specific application). The network 504 is used to provide a medium for communication links between the terminal devices 501, 502, 503 and the server 505. The network 504 may include various connection types, such as wired, wireless communication links or optical fiber cables.

[0099] Users can use terminal devices 501, 502, 503 to interact with server 505 through network 504 to receive or send messages, etc. Various client applications can be installed on terminal devices 501, 502, 503, such as applications for training models, etc. (only examples).

[0100] The terminal devices 501 , 502 , and 503 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0101] The server 505 may be a server that provides various services, such as a background server (only an example) that provides support for the application of the training model operated by the user using the terminal devices 501, 502, 503. The background server may process the received model training request and feed back the processing result (such as whether the model training is completed - only an example) to the terminal devices 501, 502, 503.

[0102] It should be noted that the model training method provided in the embodiment of the present invention is generally executed by the server 505 , and accordingly, the model training device is generally arranged in the server 505 .

[0103] It should be understood that Figure 5 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.

[0104] The present invention also provides an electronic device. The electronic device of an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method provided by the present invention.

[0105] Reference below Figure 6 , which shows a schematic diagram of the structure of a computer system 600 of an electronic device suitable for implementing an embodiment of the present invention. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0106] like Figure 6As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0107] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom is installed into the storage section 608 as needed.

[0108] In particular, according to the embodiments disclosed in the present invention, the process described in the main step diagram above can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the main step diagram. In the above embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit 601, the above functions defined in the system of the present invention are executed.

[0109] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0110] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0111] The units involved in the embodiments of the present invention may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes a supervised training unit and a distillation training unit. The names of these units do not constitute a limitation on the units themselves in some cases, for example, the supervised training unit may also be described as a "unit that provides a first difference to the distillation training unit".

[0112] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the steps executed by the device include: inputting a plurality of training images in the sample set that are pre-labeled with semantic segmentation labels into a trained teacher model and a student model to be trained respectively, and determining the probability distribution difference between the prediction result of the student model for the category to which the pixels in the training image belong and the semantic segmentation label as a first difference; and both the student model and the teacher model include a main network and a generalized normalization layer connected after the main network; for the feature maps corresponding to the plurality of training images output by the main network of the student model and the teacher model: converting to a union of the categories The image processing method comprises the following steps: the first difference is combined with the second difference and / or the third difference is used to construct the loss function of the student model to train the student model; the second difference is determined based on the joint feature of the student model and the teacher model, and the third difference is determined based on the separation feature of the student model and the teacher model.

[0113] According to the technical solution of the embodiment of the present invention, the model's ability to express niche categories and / or image detail information can be enhanced by extracting joint features and / or separation features during the knowledge distillation process.

[0114] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A model training method, It is characterized in that include: Inputting a plurality of training images pre-labeled with semantic segmentation labels in a sample set into a trained teacher model and a student model to be trained respectively, and determining the probability distribution difference between the prediction result of the student model on the category to which the pixels in the training image belong and the semantic segmentation label as a first difference; and The student model and the teacher model both include a main network and a generalized normalization layer connected after the main network; for the feature maps corresponding to the multiple training images output by the main network of the student model and the teacher model: after being converted into joint features of the category, they enter the generalized normalization layer, and / or, after being segmented into multiple separate features in height and width dimensions based on a preset segmentation rule, they enter the generalized normalization layer; wherein the joint features of each category include probability data of pixels in the feature maps corresponding to the multiple training images belonging to the category; each separate feature includes probability data of pixels in the feature map in the same segmentation space belonging to the category; the generalized normalization layer includes a Softmax layer whose temperature parameter T is not equal to 1; The student model is trained by constructing a loss function of the student model using the first difference in combination with the second difference and / or the third difference; wherein the second difference is determined based on the joint features of the student model and the teacher model, and the third difference is determined based on the separate features of the student model and the teacher model.

2. The method according to claim 1, It is characterized in that The joint features of any category are determined according to the following steps: Obtaining probability data of each pixel in the multi-channel feature map corresponding to the plurality of training images output by the corresponding subject network belonging to the category; The probability data of each pixel belonging to the category is combined into the joint features of the category.

3. The method according to claim 1, It is characterized in that The separation feature is further formed by segmenting the feature map in the channel dimension and aggregating the feature map in the category dimension; The separation features corresponding to any category of any segmented space formed by segmentation in the channel, height and width dimensions include: probability data that the pixels in the segmented space belong to the category.

4. The method according to claim 1, It is characterized in that In the case where the student model forms a joint feature, the teacher model forms a joint feature; in the case where the teacher model forms a joint feature, the student model forms a joint feature; In the case where the student model forms a separation feature, the teacher model forms a separation feature; in the case where the teacher model forms a separation feature, the student model forms a separation feature; And, the use of the first difference in combination with the second difference and / or the third difference to construct the loss function of the student model includes: determining a weighted sum of the first difference and the second difference as the loss function; or, determining a weighted sum of the first difference and the third difference as the loss function; or, A weighted sum of the first difference, the second difference, and the third difference is determined as the loss function.

5. The method according to claim 2, It is characterized in that After the joint feature of each category enters the generalized normalization layer, normalization is performed inside the joint feature to form a first normalized feature of the category; and the second difference is determined according to the following steps: Calculate the KL divergence of the first normalized features of the student model and the teacher model corresponding to the same category; The average of the KL divergences across categories is determined as the second divergence.

6. The method according to claim 3, It is characterized in that After any segmented space formed by segmentation by channel, height and width dimensions corresponds to a separation feature of any category, normalization is performed inside the separation feature to form a second normalized feature of the segmented space and the category; and the third difference is determined according to the following steps: Calculate the KL divergence of the second normalized features of the student model and the teacher model corresponding to the same position segmentation space and the same category; The average value of the KL divergence of each position segmentation space and each category is determined as the third difference.

7. The method according to any one of claims 1 to 6, It is characterized in that The feature graph of the student model enters a narrow normalization layer for calculation, and the prediction result is determined based on the calculation result of the narrow normalization layer; and Narrow normalization layers include a Softmax layer with a temperature parameter equal to 1.

8. A model training device, It is characterized in that include: A supervised training unit is used to: input multiple training images pre-labeled with semantic segmentation labels in a sample set into a trained teacher model and a student model to be trained respectively, and determine the probability distribution difference between the prediction result of the student model on the category to which the pixels in the training image belong and the semantic segmentation label as a first difference; and the student model and the teacher model both include a main network and a generalized normalization layer connected after the main network; for the feature maps corresponding to the multiple training images output by the main network of the student model and the teacher model: enter the generalized normalization layer after being converted into joint features of the category, and / or enter the generalized normalization layer after being segmented into multiple separate features in height and width dimensions based on preset segmentation rules; wherein the joint features of each category include probability data that the pixels in the feature maps corresponding to the multiple training images belong to the category; each separate feature includes probability data that the pixels in the feature map in the same segmentation space belong to the category; the generalized normalization layer includes a Softmax layer whose temperature parameter T is not equal to 1; A distillation training unit, used to: train the student model by constructing a loss function of the student model using a first difference combined with a second difference and / or a third difference; wherein the second difference is determined based on the joint features of the student model and the teacher model, and the third difference is determined based on the separate features of the student model and the teacher model.

9. The device according to claim 8, It is characterized in that The joint features of any category are determined according to the following steps: obtaining probability data of each pixel belonging to the category in the multi-channel feature map output by the corresponding subject network and corresponding to the plurality of training images; merging the probability data of each pixel belonging to the category into the joint features of the category; The separation feature is further formed by segmenting the feature map in the channel dimension and aggregating the feature map in the category dimension; The separation features corresponding to any category of any segmented space formed by segmentation in the channel, height and width dimensions include: probability data of pixels in the segmented space belonging to the category; In the case where the student model forms a joint feature, the teacher model forms a joint feature; in the case where the teacher model forms a joint feature, the student model forms a joint feature; in the case where the student model forms a separate feature, the teacher model forms a separate feature; in the case where the teacher model forms a separate feature, the student model forms a separate feature; and the distillation training unit is further used to: A weighted sum of the first difference and the second difference is determined as the loss function; or a weighted sum of the first difference and the third difference is determined as the loss function; or a weighted sum of the first difference, the second difference and the third difference is determined as the loss function.

10. An electronic device, It is characterized in that include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

11. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Training method of image classification model and related equipment

    CN110321952A

  • Method and device for compressing a neural network model for machine translation and storage medium

    US20210158126A1