Methods and apparatuses for training an image processing model and for image processing

By introducing a deep learning network attention mechanism with triple loss function, generating anchor examples and masks of positive and negative examples, the problem of difficulty in learning attention mechanism weights in computer vision is solved, and the sensitivity of the image processing model to target features is improved.

CN115641481BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110816801.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-20
Publication Date
2025-07-25
Estimated Expiration
2041-07-20

AI Technical Summary

Technical Problem

In existing computer vision technology, the weight learning of attention mechanism is difficult, and the lack of independent attention loss function leads to poor training of attention mechanisms.

Method used

The deep learning network attention mechanism based on triple loss function is adopted, and attention loss terms are generated by introducing anchor examples and masks of positive and negative examples, and the image processing model is trained to enhance the model's sensitivity to target features.

Benefits of technology

It reduces the difficulty of learning attention weight values, improves the model's sensitivity to target features, and achieves more intuitive attention weight learning, with small calculation amount and simple structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641481B_ABST
    Figure CN115641481B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and apparatuses for training an image processing model and image processing, relating to the field of artificial intelligence technologies, particularly to deep learning and computer vision technologies. The specific implementation solution is as follows: obtaining training samples and an initial image processing model including multiple feature extraction layers; inputting a sample image of the training samples into the initial image processing model to obtain an attention heat map corresponding to a target feature extraction layer, where the attention heat map is generated based on a feature map output by the target feature extraction layer and an attention weight to be trained; training the initial image processing model based on a preset loss function to obtain an image processing model, where the loss function includes a triplet loss generated based on the attention heat map as an anchor example and two masks corresponding to the annotation information of the training samples as positive and negative examples respectively. Thereby enhancing the sensitivity of the trained model to target features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to deep learning and computer vision technologies, and more particularly to methods and apparatuses for training an image processing model and image processing. Background Art

[0002] With the development of deep learning and computer vision technologies, the attention mechanism has also found increasingly widespread applications.

[0003] In order to simulate the selective attention of human vision (by quickly scanning the global image to obtain the target area that needs to be focused on, and then investing more attention resources in this area to obtain more detailed information about the target to be focused on, thereby suppressing other useless information), currently, the attention mechanism in computer vision mainly adopts methods such as channel attention, pixel attention, and multi-stage attention to improve the image processing effect. Summary of the Invention

[0004] A method and an apparatus for training an image processing model and image processing are provided.

[0005] According to a first aspect, a method for training an image processing model is provided. The method includes: obtaining a training sample and an initial image processing model, where the training sample includes a sample image and corresponding annotation information, and the initial image processing model includes a plurality of feature extraction layers; inputting the sample image into the initial image processing model to obtain an attention heat map corresponding to a target feature extraction layer, where the attention heat map is generated based on a feature map output by the target feature extraction layer and an attention weight to be trained; training the initial image processing model based on a preset loss function to obtain an image processing model, where the loss function includes an attention loss term based on triplet loss, and the attention loss term is generated based on an attention heat map serving as an anchor example and two masks corresponding to the annotation information and serving as positive and negative examples respectively.

[0006] According to a second aspect, a method for processing an image is provided. The method includes: obtaining an image to be processed; inputting the image to be processed into a pre-trained image processing model to generate an image processing result, where the image processing model is obtained by the method for training an image processing model described in any implementation manner of the first aspect.

[0007] According to a third aspect, there is provided an apparatus for training an image processing model, the apparatus including: an acquisition unit configured to acquire training samples and an initial image processing model, wherein the training samples include sample images and corresponding annotation information, and the initial image processing model includes a plurality of feature extraction layers; a generation unit configured to input the sample images into the initial image processing model to obtain an attention heat map corresponding to a target feature extraction layer, wherein the attention heat map is generated based on a feature map output by the target feature extraction layer and an attention weight to be trained; a training unit configured to train the initial image processing model based on a preset loss function to obtain an image processing model, wherein the loss function includes an attention loss term based on triplet loss, and the attention loss term is generated based on the attention heat map as an anchor example and two masks corresponding to the annotation information as positive and negative examples respectively.

[0008] According to a fourth aspect, there is provided an apparatus for processing images, the apparatus including: an acquisition unit configured to acquire an image to be processed; a processing unit configured to input the image to be processed into a pre-trained image processing model to generate an image processing result, wherein the image processing model is obtained by the method for training an image processing model described in any implementation manner of the first aspect.

[0009] According to a fifth aspect, there is provided an electronic device, the electronic device including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect or the second aspect.

[0010] According to a sixth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions for enabling a computer to execute the method described in any implementation manner of the first aspect or the second aspect.

[0011] According to a seventh aspect, there is provided a computer program product including a computer program which, when executed by a processor, implements the method described in any implementation manner of the first aspect or the second aspect.

[0012] According to an eighth aspect, there is provided an autonomous vehicle including the electronic device described in the fifth aspect.

[0013] According to the technology of the present disclosure, by proposing a deep learning network attention mechanism based on a triplet loss function, the region that the feature map attention mechanism of the input sample image focuses on should fall more within the region of the mask as a positive example, so that the feature map has a higher activation amount for the region where the object in the mask is located, that is, the sensitivity of the trained image processing model to the target features is enhanced. And, by introducing an attention weight loss into the loss function of model training, a guided learning of the attention weight of the attention mechanism is achieved. Compared with the training method of the traditional attention mechanism without an independent attention loss function, this solution reduces the difficulty of learning the attention weight value and has a more intuitive meaning. Moreover, since the attention mechanism proposed in the present disclosure only needs to introduce weight parameters into the loss function and does not add additional structures, it has the advantages of simple implementation, small computational amount, and intuitive principle.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0016] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0017] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;

[0018] Figure 3 is a schematic diagram of an application scenario of a method for training an image processing model that can implement the embodiments of the present disclosure;

[0019] Figure 4 is a schematic diagram of a device for training an image processing model according to an embodiment of the present disclosure;

[0020] Figure 5 is a schematic diagram of a device for processing images according to an embodiment of the present disclosure;

[0021] Figure 6 is a block diagram of an electronic device for implementing the method for training an image processing model or for processing images according to the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0023] Figure 1 FIG. 100 is a schematic diagram showing a first embodiment of the present disclosure. The method for training an image processing model includes the following steps:

[0024] S101, obtaining training samples and an initial image processing model.

[0025] In this embodiment, the execution subject of the method for training an image processing model can obtain training samples and an initial image processing model in various ways. Among them, the above-mentioned training samples may include sample images and corresponding annotation information. Among them, the above-mentioned initial image processing model may include various deep learning networks for image processing. It may include but is not limited to at least one of the following: object detection model, semantic segmentation model, instance segmentation model, etc. The above-mentioned initial image processing model may include multiple feature extraction layers. The structures of the above-mentioned multiple feature extraction layers are usually similar and linearly connected, that is, the output feature map of the shallow feature extraction layer is usually used as the input of the next layer (i.e., the deep) feature extraction layer.

[0026] It should be noted that the above-mentioned execution subject may also obtain a training sample set composed of multiple training samples, which is not limited herein.

[0027] S102, inputting the sample image into the initial image processing model to obtain an attention heat map corresponding to the target feature extraction layer.

[0028] In this embodiment, the above-mentioned execution subject may input the sample image of the above-mentioned training samples obtained in the above-mentioned step S101 into the above-mentioned initial image processing model, so as to obtain an attention heat map corresponding to the target feature extraction layer. Among them, the above-mentioned attention heat map may be generated based on the feature map output by the above-mentioned target feature extraction layer and the attention weights to be trained.

[0029] In this embodiment, the above-mentioned execution entity can generate the above-mentioned attention heat map based on the feature map output by the above-mentioned target feature extraction layer and the attention weights to be trained in various ways. As an example, the above-mentioned execution entity can multiply the feature map output by the above-mentioned target feature extraction layer element by element with a corresponding attention mask (mask) of the same size and then add it to the above-mentioned feature map to generate an attention heat map. Among them, each element in the above-mentioned attention mask can be used to represent the attention weight to be trained corresponding to the pixel point.

[0030] It should be noted that the above-mentioned target feature extraction layer can be any feature extraction layer pre-specified from among the multiple feature extraction layers included in the above-mentioned initial image processing model. As another example, the above-mentioned target feature extraction layer can be a feature extraction layer determined according to rules, such as the deepest (i.e., the closest to the output layer) feature extraction layer in the above-mentioned initial image processing model.

[0031] It should be noted that when there are multiple input training samples, the above-mentioned generated attention heat map corresponding to the target feature extraction layer can also correspond to the input training samples.

[0032] S103. Train the initial image processing model based on a preset loss function to obtain an image processing model.

[0033] In this embodiment, based on a preset loss function, the above-mentioned execution entity can use a machine learning method to train the above-mentioned initial image processing model to obtain an image processing model. Among them, the above-mentioned loss function can include an attention loss term based on triplet loss. The above-mentioned attention loss term can be generated based on the attention heat map as the anchor example and two masks corresponding to the above-mentioned annotation information as the positive and negative examples respectively. As an example, the above-mentioned two masks corresponding to the above-mentioned annotation information as the positive and negative examples can be generated according to the annotation information for representing the positive training samples and the annotation information for representing the negative training samples respectively. For example, the mask as the positive example can include a matrix in which the elements within the range indicated by the annotation information are set to 1 and the corresponding elements in other ranges are set to 0; the mask as the negative example can include a matrix of the same size as the above-mentioned mask as the positive example, where the element values in the above-mentioned matrix are all slightly greater than 0 (for example, 0.01).

[0034] In this embodiment, as an example, the above-mentioned attention loss term can be represented by the following formula (1):

[0035] L k =H k ×M k -H k ×M' k (1)

[0036] Among them, the above-mentioned Lk The triplet loss that can be used to characterize the k-th layer feature extraction layer (such as the target feature extraction layer). The above H k The attention heat map of the k-th layer feature extraction layer corresponding to the generated sample image of the current input that can be used to characterize. The above M k and M' k can be used to characterize two masks corresponding to the above annotation information as positive and negative examples respectively. The above × can be used to characterize the multiplication of corresponding elements in the matrix.

[0037] The method provided by the above embodiments of the present disclosure, by proposing a deep learning network attention mechanism based on the triplet loss function, enables the region focused by the feature map attention mechanism corresponding to the input sample image to fall more within the region of the mask as a positive example, so that the feature map has a higher activation amount for the region where the object in the mask is located, that is, the sensitivity of the trained image processing model to the target feature is enhanced. And, by introducing the attention weight loss into the loss function of model training, the guided learning of the attention weight of the attention mechanism is realized. Compared with the traditional attention mechanism training method without an independent attention loss function, this solution reduces the difficulty of learning the attention weight value and has a more intuitive meaning. Moreover, since the attention mechanism proposed by the present disclosure only needs to introduce weight parameters in the loss function and does not add additional structures, it has the advantages of simple implementation, small computational amount, and intuitive principle.

[0038] In some optional implementation manners of this embodiment, the above annotation information may include information for annotating the position of the target object in the above sample image. The above target object may be related to the use of the image processing model, for example, it may be used to characterize an object to be recognized or detected. As an example, the above target object may be a face, and the above annotation information may be the display position of the face in the image (for example, it may be represented by a detection frame). Based on this, the two masks corresponding to the annotation information can be generated through the following steps:

[0039] The first step is to determine the position of the target object in the feature map output by the target feature extraction layer according to the annotation information.

[0040] In these implementation manners, according to the annotation information, the above execution subject can determine the position of the target object in the feature map output by the target feature extraction layer. As an example, the above execution subject can perform a proportional transformation on the above annotation information according to the size relationship between the feature map output by the above target feature extraction layer and the sample image of the training sample, so as to present in the feature map output by the above target feature extraction layer. Thus, the position of the above target object in the feature map output by the target feature extraction layer can be determined.

[0041] The second step is to generate the first mask according to the position.

[0042] In these implementation manners, the above-mentioned execution entity may generate a first mask in various ways according to the position determined in the above-mentioned first step.

[0043] As an example, the above-mentioned first mask may be represented by the following formula (2):

[0044]

[0045] Among them, the above-mentioned can be used to represent the element at the (i, j) position of the first mask of the feature map output by the k-th layer feature extraction layer (such as the target feature extraction layer). The above-mentioned S can be used to represent the range covered by the position determined in the above-mentioned first step.

[0046] It should be noted that for the above-mentioned initial image processing model including multiple feature extraction layers, the value of the above-mentioned k can start from the deepest feature extraction layer of the network and increase forward. It is also possible to select an intermediate feature extraction layer as the target extraction layer according to actual applications, which is not limited here.

[0047] The third step is to generate a second mask based on the inversion of the elements in the first mask.

[0048] In these implementation manners, the above-mentioned execution entity may perform element inversion based on the first mask generated in the above-mentioned second step to generate a second mask. As an example, the above-mentioned execution entity may set the value of the element whose original value is 0 in the above-mentioned first mask to 1, and then set the value of the element whose original value is 1 to 0 to generate a second mask.

[0049] Based on the above optional implementation manners, this solution can generate two masks corresponding to the annotation information as positive and negative examples with relatively large differences based on the annotation information of the training samples, thereby amplifying the differences between the corresponding samples and helping the model improve the training effect.

[0050] In some optional implementation manners of this embodiment, the above-mentioned execution entity may input the sample image into the initial image processing model through the following steps to obtain an attention heat map corresponding to the target feature extraction layer:

[0051] The first step is to input the sample image into the initial image processing model to obtain the feature map output by the target feature extraction layer.

[0052] The second step is to fuse the product of each channel in the feature map and the corresponding attention weight to be trained with the feature map to generate a new feature map.

[0053] In these implementation manners, based on the product of each channel in the feature map output in the above-mentioned first step and the corresponding attention weight to be trained, the above-mentioned execution entity can fuse it with the feature map in various ways to generate a new feature map.

[0054] As an example, when adopting the channel attention mechanism, the above-mentioned new feature map can be represented by the following formula (3):

[0055]

[0056] Among them, the above-mentioned can be used to represent the c-th channel of the new feature map corresponding to the feature map output by the k-th layer feature extraction layer (such as the target feature extraction layer). The above-mentioned can be used to represent the c-th channel of the feature map output by the k-th layer feature extraction layer (such as the target feature extraction layer). The above-mentioned can be used to represent the attention weight value corresponding to the c-th channel of the above-mentioned target feature extraction layer. The above-mentioned p can be used to represent the total number of channels of the feature map output by the above-mentioned target feature extraction layer.

[0057] The third step is to generate an attention heat map based on the new feature map and the attention weight to be trained.

[0058] In these implementation manners, based on the new feature map generated in the above-mentioned second step and the attention weight to be trained, the above-mentioned execution entity can generate an attention heat map in various ways.

[0059] As an example, when adopting the channel attention mechanism, the attention heat map corresponding to the above-mentioned new feature map can be represented by the following formula (4):

[0060]

[0061] Among them, the above-mentioned H k 、 The meaning of p can be the same as the previous description and will not be elaborated here.

[0062] Therefore, the attention heat map corresponding to the above-mentioned target feature extraction layer is usually the same as the length and width of the feature map output by the above-mentioned target feature extraction layer. By analogy with Class Activation Mapping (CAM), the above-mentioned attention heat map can reflect the attention situation of the above-mentioned target feature extraction layer to different positions of the input sample image.

[0063] Based on the above optional implementation manners, this solution can apply the channel domain attention mechanism to the attention heat map, thereby enriching the generation method of the attention heat map, and further enriching the training method of the image processing model based on the triplet loss.

[0064] In some alternative implementation manners of this embodiment, the above attention loss term may be determined based on attention loss sub - terms corresponding to respective training samples in the batch to which the above training samples belong. The above attention loss sub - terms may be determined according to the difference between the first fusion result and the second fusion result and a preset compensation parameter. The above first fusion result may be calculated based on the above attention heat map and a mask as a positive example. The above second fusion result is calculated based on the above attention heat map and a mask as a negative example.

[0065] In these implementation manners, by way of example, the above attention loss term may be represented by the following formula (5):

[0066]

[0067] Wherein, the above L k , M k and M' k may have the same meanings as those described above, and will not be elaborated here. The above can be used to represent the attention heat map of the k - th feature extraction layer of the sample image corresponding to the i - th input in the above - generated batch. The above α k can be used to represent the above - preset compensation parameter. Among them, the above compensation parameter is usually greater than 0. The above {} + can be used to represent taking a positive value. The above m can be used to represent the number of training samples in the batch to which the above training samples belong.

[0068] Based on the above - mentioned alternative implementation manners, this solution can apply the above - mentioned model training method based on triplet loss to the batch training process, thereby improving the applicability of the method for training an image processing model.

[0069] In some alternative implementation manners of this embodiment, the above loss function may include statistical values of attention loss terms obtained by respectively using multiple feature extraction layers as target feature extraction layers. The above loss function may also include a difference term between the output result corresponding to the input training sample output by the above initial image processing model and the annotation information.

[0070] In these implementation manners, the above statistical values may include but are not limited to at least one of the following: average value, maximum value, minimum value.

[0071] Based on the above optional implementation manners, the present solution can reflect the overall situation of the attention weights of multiple feature extraction layers included in the above initial image processing model (such as the average value of the attention loss terms corresponding to the multiple feature extraction layers included) through the above statistical values, so as to optimize the overall image processing model. Moreover, since the difference term between the annotation information corresponding to the sample image and the final output result of the initial image processing model is still retained in the loss function for model training, it will not have an obvious adverse effect on the model training effect, thus ensuring the model training effect.

[0072] Continue to refer to Figure 2 , Figure 2 which is a schematic diagram 200 according to the second embodiment of the present disclosure. The method for processing an image includes the following steps:

[0073] S201, obtain an image to be processed.

[0074] In this embodiment, the execution subject of the method for processing an image can obtain the image to be processed in various ways. Among them, the above image to be processed can include various images that can be processed by a deep learning network, which is not limited herein.

[0075] In this embodiment, as an example, the above image to be processed can be an image containing a person. As another example, the above image to be processed can be a road condition image captured by an autonomous driving vehicle. The above execution subject can obtain the above image to be processed from a local or communicatively connected electronic device.

[0076] S202, input the image to be processed into a pre-trained image processing model to generate an image processing result.

[0077] In this embodiment, the above execution subject can input the image to be processed obtained in the above step S201 into a pre-trained image processing model to generate an image processing result corresponding to the image to be processed. Among them, the above image processing model can be trained by the method for training an image processing model described in any of the implementation manners in the foregoing embodiments. The above image processing result can correspond to the above image processing model. As an example, when the above image processing model is a face recognition model, the above image processing result can be used to characterize the information of the person shown in the face image. As another example, when the above image processing model is a lane detection model, the above image processing result can be used to indicate the position where the lane line is displayed in the image.

[0078] From Figure 2As can be seen, the process 200 of the method for processing images in this embodiment embodies the steps of using the image processing model trained by the method of training the image processing model to perform corresponding processing on the image. Thus, the solution described in this embodiment can improve the sensitivity of the image to target feature recognition and further improve the effect of image processing because it uses an image processing model trained by introducing an improved attention mechanism based on triplet loss.

[0079] Continue to refer to Figure 3 , Figure 3 which is a schematic diagram of an application scenario of the method for training an image processing model according to an embodiment of the present disclosure. In the Figure 3 application scenario, the execution entity (such as a server) for training the image processing model can first obtain a training sample 301 and an initial image processing model 302. Among them, the above training sample 301 may include a sample image 3011 and corresponding annotation information 3012. Then, the above execution entity can input the above sample image 3011 into the above initial image processing model 302 to obtain an attention heat map 303 corresponding to the target feature extraction layer. After that, based on a preset loss function, the above execution entity can use a machine learning method to adjust the weights of the above initial image processing model and the attention weights to obtain an image processing model. Among them, the attention loss term included in the above loss function is generated based on the attention heat map 303 as the anchor example and the masks 304 and 305 as the positive and negative examples respectively.

[0080] Currently, one of the existing technologies usually introduces an attention mechanism by means of channel domain attention or attention masks. However, since there is no term including attention loss in the loss function, it is impossible to guide the learning of the weights of the attention mechanism, resulting in a difficult learning process for the attention weights. The method provided in the above embodiment of the present disclosure proposes a deep learning network attention mechanism based on a triplet loss function, so that the region focused by the feature map attention mechanism corresponding to the input sample image should fall more within the region of the mask as the positive example, so that the feature map has a higher activation amount for the region where the object in the mask is located, that is, the sensitivity of the trained image processing model to the target feature is enhanced. And, by introducing the attention weight loss into the loss function of model training, the guided learning of the attention weights of the attention mechanism is realized. Compared with the traditional training method of the attention mechanism without an independent attention loss function, this solution reduces the difficulty of learning the attention weight value and has a more intuitive meaning. Moreover, since the attention mechanism proposed by the present disclosure only introduces weight parameters in the loss function and does not add additional structures, it has the advantages of simple implementation, small computational amount, and intuitive principle.

[0081] Further refer to Figure 4, as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an apparatus for training an image processing model. This apparatus embodiment corresponds to Figure 1 the method embodiment shown, and this apparatus can be specifically applied to various electronic devices.

[0082] As shown in Figure 4 , the apparatus 400 for training an image processing model provided in this embodiment includes an acquisition unit 401, a generation unit 402, and a training unit 403. Among them, the acquisition unit 401 is configured to acquire training samples and an initial image processing model, where the training samples include sample images and corresponding annotation information, and the initial image processing model includes multiple feature extraction layers; the generation unit 402 is configured to input the sample image into the initial image processing model to obtain an attention heat map corresponding to the target feature extraction layer, where the attention heat map is generated based on the feature map output by the target feature extraction layer and the attention weights to be trained; the training unit 403 is configured to train the initial image processing model based on a preset loss function to obtain an image processing model, where the loss function includes an attention loss term based on triplet loss, and the attention loss term is generated based on the attention heat map as the anchor example and two masks corresponding to the annotation information as the positive and negative examples respectively.

[0083] In this embodiment, in the apparatus 400 for training an image processing model: the specific processing of the acquisition unit 401, the generation unit 402, and the training unit 403 and the technical effects brought by them can respectively refer to Figure 1 the relevant descriptions of steps S101, S102, and S103 in the corresponding embodiments, which will not be elaborated here.

[0084] In some optional implementation manners of this embodiment, the above annotation information may include information for annotating the position of the target object in the sample image, and the two masks corresponding to the annotation information may be generated through the following steps: determining the position of the target object in the feature map output by the target feature extraction layer according to the annotation information; generating a first mask according to the position; generating a second mask based on the inversion of the elements in the first mask.

[0085] In some optional implementation manners of this embodiment, the above generation unit 402 may be further configured to: input the sample image into the initial image processing model to obtain the feature map output by the target feature extraction layer; generate a new feature map by fusing the product of each channel in the feature map and the corresponding attention weights to be trained with the feature map; generate an attention heat map based on the new feature map and the attention weights to be trained.

[0086] In some alternative implementation manners of this embodiment, the above attention loss term may be determined based on attention loss sub - terms corresponding to each training sample in the batch to which the above training sample belongs. The above attention loss sub - term may be determined according to the difference between the first fusion result and the second fusion result and a preset compensation parameter. The above first fusion result may be calculated based on the above attention heat map and a mask used as a positive example. The above second fusion result may be calculated based on the above attention heat map and a mask used as a negative example.

[0087] In some alternative implementation manners of this embodiment, the above loss function may include statistical values of attention loss terms obtained by using multiple feature extraction layers as target feature extraction layers respectively; and the above loss function may further include a difference term between the output result corresponding to the input training sample output by the initial image processing model and the annotation information.

[0088] The device provided in the above embodiment of the present disclosure trains the initial image processing model through the triple loss generated by the training unit 403 based on the attention heat map generated by the generation unit 402 and the masks used as positive and negative examples corresponding to the sample image acquired by the acquisition unit 401, so that the region focused by the feature map attention mechanism corresponding to the input sample image should fall more within the region of the mask used as a positive example, so that the feature map has a higher activation amount for the region where the object in the mask is located, that is, the sensitivity of the trained image processing model to the target feature is enhanced. And, by introducing the attention weight loss into the loss function of the model training, the guided learning of the attention weight of the attention mechanism is realized. Compared with the traditional training method of the attention mechanism without an independent attention loss function, this solution reduces the difficulty of learning the attention weight value, and the meaning is more intuitive. Moreover, since the attention mechanism proposed in the present disclosure only needs to introduce weight parameters in the loss function and does not add additional structures, it has the advantages of simple implementation, small computational amount, and intuitive principle.

[0089] Further referring to Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for processing images. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0090] As Figure 5As shown in the figure, the apparatus 500 for processing images provided in this embodiment includes an acquisition unit 501 and a processing unit 502. Among them, the acquisition unit 501 is configured to acquire an image to be processed; the processing unit 502 is configured to input the image to be processed into a pre-trained image processing model to generate an image processing result, where the image processing model is obtained by the method for training an image processing model described in the foregoing embodiment.

[0091] In this embodiment, in the apparatus 500 for processing images: For the specific processing of the acquisition unit 501 and the processing unit 502 and the technical effects brought by them, reference can be made to Figure 2 the relevant descriptions of steps S201 and S202 in the corresponding embodiment, which will not be elaborated here.

[0092] The apparatus provided in the above embodiment of the present disclosure uses the image processing model obtained by training the processing unit 502 using the method for training an image processing model to perform corresponding processing on the image acquired by the acquisition unit 501. Since an image processing model trained by introducing an improved attention mechanism based on triplet loss is used, the sensitivity of the image to target feature recognition can be improved, and thus the effect of image processing can be enhanced.

[0093] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.

[0094] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, a computer program product, and an autonomous vehicle.

[0095] Figure 6 The schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0096] The autonomous vehicle provided by the present disclosure may include the above-mentioned electronic device as Figure 6 shown.

[0097] As Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 602 or computer programs loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0098] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0099] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for training an image processing model or the method for processing images. For example, in some embodiments, the method for training an image processing model or the method for processing images can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for training an image processing model or the method for processing images described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for training an image processing model or the method for processing images in any other appropriate manner (e.g., by means of firmware).

[0100] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0101] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0104] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0105] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0106] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0107] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for training an image processing model, comprising: Obtaining a training sample and an initial image processing model, wherein the training sample includes a sample image and corresponding annotation information, and the initial image processing model includes a plurality of feature extraction layers; Inputting the sample image into the initial image processing model to obtain an attention heat map corresponding to a target feature extraction layer, wherein the attention heat map is generated based on a feature map output by the target feature extraction layer and an attention weight to be trained; Training the initial image processing model based on a preset loss function to obtain an image processing model, wherein the loss function includes an attention loss term based on triplet loss, and the attention loss term is generated based on the attention heat map as an anchor example and two masks corresponding to the annotation information as positive and negative examples respectively; The annotation information includes information for annotating the position of a target object in the sample image, and the two masks corresponding to the annotation information are generated through the following steps: Determining the position of the target object in the feature map output by the target feature extraction layer according to the annotation information; Generating a first mask according to the position; Generating a second mask based on the inversion of the elements in the first mask.

2. The method according to claim 1, wherein, The step of inputting the sample image into the initial image processing model to obtain an attention heat map corresponding to a target feature extraction layer includes: Inputting the sample image into the initial image processing model to obtain the feature map output by the target feature extraction layer; Fusing the product of each channel in the feature map and the corresponding attention weight to be trained with the feature map to generate a new feature map; Generating the attention heat map based on the new feature map and the attention weight to be trained.

3. The method according to claim 1 or 2, wherein The attention loss term is determined based on attention loss sub - terms corresponding to each training sample in the batch to which the training sample belongs, and the attention loss sub - term is determined according to the difference between a first fusion result and a second fusion result and a preset compensation parameter. The first fusion result is calculated based on the attention heat map and the mask as a positive example, and the second fusion result is calculated based on the attention heat map and the mask as a negative example.

4. The method according to claim 1 or 2, wherein The loss function includes the statistical value of the attention loss terms obtained when the plurality of feature extraction layers are respectively used as the target feature extraction layer; and The loss function further includes the difference term between the output result corresponding to the input training sample output by the initial image processing model and the annotation information.

5. A method for image processing, comprising: Obtaining an image to be processed; Inputting the image to be processed into a pre - trained image processing model to generate an image processing result, wherein the image processing model is trained according to the method of any one of claims 1 - 4.

6. An apparatus for training an image processing model, comprising: An obtaining unit configured to obtain a training sample and an initial image processing model, wherein the training sample includes a sample image and corresponding annotation information, and the initial image processing model includes a plurality of feature extraction layers; A generation unit configured to input the sample image into the initial image processing model to obtain an attention heat map corresponding to the target feature extraction layer, where the attention heat map is generated based on the feature map output by the target feature extraction layer and the attention weights to be trained; A training unit configured to train the initial image processing model based on a preset loss function to obtain an image processing model, where the loss function includes an attention loss term based on triplet loss, and the attention loss term is generated based on the attention heat map as the anchor example and two masks corresponding to the annotation information as the positive and negative examples respectively; The annotation information includes information for annotating the position of the target object in the sample image, and the two masks corresponding to the annotation information are generated through the following steps: Determine the position of the target object in the feature map output by the target feature extraction layer according to the annotation information; Generate a first mask according to the position; Generate a second mask based on the inversion of the elements in the first mask.

7. The apparatus according to claim 6, wherein The generation unit is further configured to: Input the sample image into the initial image processing model to obtain the feature map output by the target feature extraction layer; Fuse the product of each channel in the feature map and the corresponding attention weights to be trained with the feature map to generate a new feature map; Generate the attention heat map based on the new feature map and the attention weights to be trained.

8. The device according to claim 6 or 7, wherein The attention loss term is determined based on the attention loss sub-terms corresponding to each training sample in the batch to which the training sample belongs, and the attention loss sub-term is determined according to the difference between the first fusion result and the second fusion result and a preset compensation parameter. The first fusion result is calculated based on the attention heat map and the mask as the positive example, and the second fusion result is calculated based on the attention heat map and the mask as the negative example.

9. The device according to one of claims 6 or 7, wherein The loss function includes the statistical value of the attention loss terms obtained by using the multiple feature extraction layers as the target feature extraction layer respectively; and The loss function further includes the difference term between the output result corresponding to the input training sample output by the initial image processing model and the annotation information.

10. An apparatus for image processing, comprising: An image acquisition unit configured to acquire an image to be processed; A processing unit configured to input the image to be processed into a pre-trained image processing model to generate an image processing result, where the image processing model is trained according to the method of any one of claims 1-4.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method of any one of claims 1-5.

13. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.

14. An autonomous vehicle comprising the electronic device according to claim 11.

Citation Information

Patent Citations

  • Target detection model training method and device and target detection method and device

    CN112200862A

  • Performing attribute-aware based tasks via an attention-controlled neural network

    US20190258925A1