Training method of deep learning model, image processing method, device and equipment

By combining instance-level CUTMIX operations and adversarial image generation with hybrid training of deep learning models, the problems of insufficient model accuracy and adversarial capabilities are solved, and the accuracy and generalization ability of the model in target object recognition are improved.

CN116758373BActive Publication Date: 2025-11-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310722402.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-11-04
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

Existing deep learning models suffer from insufficient accuracy, generalization ability, and adversarial capabilities in smart city applications. They are particularly difficult to effectively identify target objects when the amount of data is small, and image-level data augmentation methods have limited effectiveness.

Method used

By combining instance-level CUTMIX operations and gradient-based adversarial image generation with hybrid training of deep learning models, enhanced and adversarial images are generated and trained together to improve the model's recognition accuracy and adversarial capabilities.

Benefits of technology

It improves the accuracy, generalization ability, and adversarial capability of deep learning models in target object recognition, and enhances the model's recognition ability in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758373B_ABST
    Figure CN116758373B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method of a deep learning model, relates to the technical field of artificial intelligence, in particular to deep learning and computer vision technology, and can be applied in the smart city scenario. The specific implementation scheme is as follows: determining at least one image block containing a target object from each original image in a plurality of original images for a batch training, obtaining an image block set; for each original image, generating an enhanced image of the original image according to the image block set; for each original image, adding interference to the original image according to an output result of the deep learning model for the original image, obtaining an adversarial image of the original image; and training the deep learning model using the enhanced image and the adversarial image of each of the plurality of original images, obtaining a trained deep learning model. The present disclosure also provides an image processing method and device, an electronic device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to deep learning, specifically deep learning and computer vision technology, which can be applied in smart city scenarios. More specifically, this disclosure provides a method for training a deep learning model, an image processing method, an apparatus, an electronic device, and a storage medium. Background Technology

[0002] Artificial intelligence is playing an increasingly important role in the construction of smart cities. For example, computer vision technology in artificial intelligence is used in scenarios such as case identification and violation detection, which can improve the efficiency and enhance the security of urban management. Summary of the Invention

[0003] This disclosure provides a method for training a deep learning model, an image processing method, an apparatus, a device, and a storage medium.

[0004] According to a first aspect, a method for training a deep learning model is provided, the method comprising: determining at least one image patch containing a target object from each of a plurality of original images used for batch training, thereby obtaining an image patch set; generating an augmented image of the original image based on the image patch set for each original image; adding perturbation to the original image based on the output of the deep learning model for the original image for each original image, thereby obtaining an adversarial image of the original image; and training the deep learning model using the augmented image and adversarial image of each of the plurality of original images, thereby obtaining a trained deep learning model.

[0005] According to the second aspect, an image processing method is provided, the method comprising: acquiring an image to be processed; and inputting the image to be processed into a deep learning model to obtain a processing result of the image to be processed; wherein the deep learning model is trained according to the training method of the aforementioned deep learning model.

[0006] According to a third aspect, a training apparatus for a deep learning model is provided, the apparatus comprising: an image patch determination module for determining at least one image patch containing a target object from each of a plurality of original images used for batch training, thereby obtaining an image patch set; an augmented image generation module for generating an augmented image of the original image for each original image based on the image patch set; an adversarial image generation module for adding perturbation to the original image for each original image based on the output of the deep learning model for the original image, thereby obtaining an adversarial image of the original image; and a training module for training the deep learning model using the respective augmented images and adversarial images of the plurality of original images, thereby obtaining a trained deep learning model.

[0007] According to a fourth aspect, an image processing apparatus is provided, the apparatus comprising: an acquisition module for acquiring an image to be processed; and a processing module for inputting the image to be processed into a deep learning model to obtain a processing result of the image to be processed; wherein the deep learning model is trained using the aforementioned deep learning model training apparatus.

[0008] According to a fifth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to the present disclosure.

[0009] According to a sixth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods provided in this disclosure.

[0010] According to a seventh aspect, a computer program product is provided, comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method provided in this disclosure when executed by a processor.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This is an exemplary system architecture diagram illustrating a training method for a deep learning model and an image processing method applicable to an embodiment of this disclosure;

[0014] Figure 2 This is a flowchart of a training method for a deep learning model according to an embodiment of the present disclosure;

[0015] Figure 3 This is a schematic diagram of a training method for a deep learning model according to an embodiment of the present disclosure;

[0016] Figure 4A This is a schematic diagram of the processing module containing the BN layer and the corresponding feature distribution information in the related technology;

[0017] Figure 4B This is a schematic diagram of a processing module including a primary BN sublayer and an auxiliary BN sublayer, and corresponding feature distribution information, according to an embodiment of the present disclosure.

[0018] Figure 5 This is a schematic diagram of a training method for a deep learning model according to an embodiment of the present disclosure;

[0019] Figure 6 This is a flowchart of an image processing method according to an embodiment of the present disclosure;

[0020] Figure 7 This is a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure;

[0021] Figure 8 This is a block diagram of an image processing apparatus according to an embodiment of the present disclosure;

[0022] Figure 9 This is a block diagram of an electronic device comprising at least one of a deep learning model training method and an image processing method according to an embodiment of the present disclosure. Detailed Implementation

[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0024] Smart cities adhere to the development strategy of "platform + ecosystem". Starting from the overall perspective that the city is a "living organism", it gives full play to the advantages of data, technology, ecosystem and security, etc., with the goal of solving many problems faced in the process of urban development. It provides intelligent products and solutions in the fields of urban insight, urban governance, industrial development and public services, and comprehensively helps the digital transformation and upgrading of cities.

[0025] For AI models that are actually put into production, their performance can be evaluated from three perspectives: accuracy, generalization ability, and adversarial capability.

[0026] Regarding model accuracy, a model can generally be deployed only if its prediction accuracy reaches a certain standard.

[0027] Regarding the generalization ability of a model, the model should have good generalization ability for predicting unknown samples and good predictive accuracy for unseen scenes and objects. However, in many cases, some models, in order to achieve better recognition ability in certain scenarios, overemphasize the recognition ability of specific scenarios, which can lead to overfitting and loss of the expected generalization ability.

[0028] Regarding the adversarial nature of a model, it should be insensitive to attacks from malicious data and should not be fooled into producing incorrect predictions. For example, in facial recognition scenarios, some models with poor adversarial capabilities may allow individuals wearing masks to pass facial recognition, causing adverse effects.

[0029] Currently, the overall performance of the model can be improved in the following ways.

[0030] One approach is to modify the model's structure, such as adding a module that provides global model information or modifying the structure of individual modules, to improve the model's robustness. However, this approach is not adaptable to the specific model and scenario, has a long development cycle, and the benefits are not guaranteed. For example, modifications that work on model A may fail on model B, or may not be applicable to model B at all.

[0031] One approach is to leverage large models or transfer learning to learn the weights of other pre-trained models to improve the robustness of the model. However, developing large models requires much more data, longer model update times, and is expensive.

[0032] One approach is to use data augmentation to create as much sample information as possible to improve the robustness of the model.

[0033] Data augmentation can be adapted to both of the first two methods. In other words, making full use of effective data can improve the capabilities of general operators, and can then be combined with modifying the model structure and training large models to improve the overall performance of the model.

[0034] However, many business scenarios face the problem of insufficient data volume. On the one hand, in current urban scenarios, there are many application areas where the number of certain cases is too small. For example, in terms of security behavior, due to the lack of legal awareness, it is difficult to collect cases of violations. On the other hand, many identification behaviors also have a large data requirement. Therefore, it is necessary to make the best use of data.

[0035] One data augmentation method can enhance the original image by cropping, flipping, and brightness transformation. However, these are all operations performed on the image itself, i.e., image-level data augmentation. Image-level data augmentation does not increase the number of target objects in the image; therefore, image-level augmentation methods have limited effect on improving the model's ability to identify target objects.

[0036] Furthermore, image-level cropping may remove instances, and for object detection tasks, methods such as flipping and random cropping can cause the detection boxes of the objects to change together, resulting in different detection boxes being generated for the same input, which has a negative impact on model training.

[0037] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0038] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0039] Figure 1 This is a schematic diagram of an exemplary system architecture according to an embodiment of the present disclosure, illustrating a training method for a deep learning model and an image processing method that can be applied. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0040] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0041] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices, including but not limited to smartphones, tablets, laptops, etc.

[0042] The training method for the deep learning model provided in this disclosure can generally be executed by the server 105. Accordingly, the training device for the deep learning model provided in this disclosure can generally be located in the server 105.

[0043] The image processing training method provided in this disclosure can generally be executed by terminal devices 101, 102, and 103. Correspondingly, the image processing apparatus provided in this disclosure can generally be disposed in terminal devices 101, 102, and 103.

[0044] Figure 2 This is a flowchart of a training method for a deep learning model according to an embodiment of the present disclosure.

[0045] like Figure 2 As shown, the training method 200 of this deep learning model includes operations S210 to S240.

[0046] In operation S210, at least one image patch containing the target object is determined from each of the multiple original images used for batch training, resulting in a set of image patches.

[0047] Multiple original images can be original images within a batch. Each original image can include at least one target object. The target object can be an animal such as a cat or dog, an object such as a table or book, or an object such as a pedestrian or vehicle.

[0048] According to embodiments of this disclosure, for each original image, the boundary of at least one target object is determined from the original image, and at least one image block is cropped from the original image based on the boundary; and an image block set is determined based on at least one image block corresponding to each original image.

[0049] For example, for each original image, the boundary of the target object can be identified from the original image. The target object can then be cropped from the original image according to the boundary, resulting in an image patch containing the target object. A target object can also be called an instance, and an image patch is an instance-level image.

[0050] Each original image can be cropped into at least one image block. Each original image in a batch corresponds to at least one image block. The at least one image block corresponding to each original image is combined together to form an image block set.

[0051] In operation S220, for each original image, an enhanced image of the original image is generated based on the set of image patches.

[0052] For example, for each original image, a certain number (e.g., 2) of image blocks can be randomly selected from the image block set, and the selected image blocks can be pasted into the original image to obtain the enhanced image of the original image.

[0053] It can be understood that the enhanced image contains the target object from the original image itself, as well as the target objects from other original images within the same batch. This allows target objects from multiple original images within the batch to be pasted into a wider background, increasing the diversity of target objects. Consequently, the model learns more features of the target objects, improving its ability to recognize them.

[0054] Operations S210 to S220 described above can be referred to as object-level CUTMIX operations (or instance-level CUTMIX operations). CUTMIX refers to cropping a portion of an image and then overlaying the cropped portion onto another image; this is also a form of image enhancement. However, the cropping region in common CUTMIX operations is random. In this embodiment, instance-level CUTMIX crops an image patch containing the target object from one image and pastes the cropped patch onto another image. Furthermore, this embodiment performs instance-level CUTMIX on multiple original images within a batch, allowing the target objects in that batch to be pasted into more backgrounds, increasing the diversity of target objects and improving the model's ability to recognize target objects.

[0055] In operation S230, an adversarial image of the original image is generated based on the output of the deep learning model for the original image.

[0056] To enhance the adversarial nature of the model, adversarial images corresponding to the original images can be introduced into the training process.

[0057] For example, adversarial images can be generated by adding perturbations to the original image, and the labels (including the category and location of the target object) in the adversarial image are consistent with those in the original image. Because adversarial images contain perturbation information, models with poor adversarial capabilities may misidentify the target object in the adversarial image (wrong category or wrong location).

[0058] Therefore, the purpose of introducing adversarial images into the training is to enable the model to correctly identify target objects not only in the original image but also in adversarial images containing interference information. In other words, even if interference is added to the original image, the model can still correctly identify the target object, thus making the model highly adversarial.

[0059] In this embodiment, the adversarial image is generated based on the gradient of the original image. For example, the original image is input into a deep learning model for forward inference to obtain an output corresponding to the original image, which may include the category of the target object. Then, the loss of the original image is calculated based on its label (true category), the gradient is calculated using the loss, and perturbations are added to the original image based on the gradient.

[0060] This step inputs the original image into the deep learning model for forward inference and calculates the gradient based on the output. It's important to note that the gradient generated in this step is not used for backpropagation, meaning it doesn't update the model's parameters. Instead, it's used to generate adversarial images.

[0061] Since gradient information reflects the direction of the model's incorrect prediction of the original image, this embodiment uses the gradient information corresponding to the original image to generate adversarial images. Using adversarial images to train the model can increase the difficulty for the model to identify target objects in the adversarial images, thereby improving the training effect of the model.

[0062] It should be noted that for object detection tasks, the output of the original image includes the category and the bounding box. However, before being input into the deep learning model, the original image may undergo some cropping and rotation operations to adapt to the size requirements of the model. The classification result of the original image will not change with the cropping and rotation operations, but the bounding box will. Therefore, to avoid the problem of inaccurate loss caused by changes in the bounding box, the gradient can be calculated using only the loss of the category recognition result, and the gradient can be used to generate adversarial images. This ensures that the adversarial images are generated with accurate gradient information, thereby improving the adversarial nature of the model.

[0063] In operation S240, the deep learning model is trained using the augmented and adversarial images of multiple original images to obtain the trained deep learning model.

[0064] For example, augmented images, obtained through instance-level data augmentation, contain richer information about the target object, enabling the model to learn more features of that object. Adversarial images, generated using gradient information, increase the difficulty for the model to identify the target object. Therefore, using augmented and adversarial images for mixed training of deep learning models can improve their training performance. For instance, it can improve at least one of the following: accuracy, generalization ability, and adversarial capability, thereby enhancing the accuracy of the deep learning model in identifying target objects in images.

[0065] The embodiments of this disclosure determine image patches containing target objects from each of a plurality of original images, perform instance-level data augmentation on the original images using the image patches to obtain augmented images, add perturbations to the original images using the output of a deep learning model to obtain adversarial images, and perform mixed training using augmented images and adversarial images to improve the training effect of the deep learning model, thereby improving the accuracy of the deep learning model in recognizing target objects in images.

[0066] Figure 3 This is a schematic diagram of a training method for a deep learning model according to an embodiment of the present disclosure.

[0067] like Figure 3As shown, this embodiment represents one round of training for a deep learning model. The original image 301 can be multiple original images used for a batch of training; the number of original images 301 in this batch can be two, four, or other quantities. Before training begins, the original image 301 undergoes two-branch processing, which generates the enhanced image 303 and the adversarial image 312, respectively. It can be understood that each round of training includes the generation of the enhanced image 303 and the adversarial image 312.

[0068] The generation of enhanced image 303 is explained below.

[0069] For each original image 301 in the current batch, image blocks containing the target object can be cropped from each original image 301, and the image blocks cropped from each original image 301 are combined into an image block set 302. The image blocks in this image block set 302 are all instance-level image blocks.

[0070] According to embodiments of this disclosure, a predetermined number of image blocks are determined from a set of image blocks, wherein the predetermined number is greater than or equal to a minimum number of image blocks and less than or equal to a maximum number of image blocks, and the number of image blocks includes the number of image blocks in each of the multiple original images; the remaining image blocks from the predetermined number of image blocks, excluding the image blocks from the original images, are pasted into the original images to obtain a fused image; and the edges of the image blocks in the fused image are smoothed to obtain an enhanced image of the original image.

[0071] For each original image 301, a predetermined number of image blocks are randomly selected from the image block set 302. This predetermined number can range from [min, max], where min represents the minimum number of image blocks and max represents the maximum number of image blocks. The number of image blocks includes the number of image blocks (target objects) in each original image within the batch.

[0072] For example, if a batch contains four original images, with the first image containing two target objects, the second containing three, the third containing one, and the fourth containing two, then the minimum number of image patches (min) = 1, and the maximum number of image patches (max) = 3. The predetermined number of patches can be set within the range of [1, 3]. This ensures that the number of target objects in the multiple enhanced images is not significantly different, allowing the model to perform balanced recognition across all enhanced images within a single batch.

[0073] A predetermined number of image blocks randomly selected from the image block set 302 may contain the target object itself in the current original image 301; that is, the selected image blocks may originate from the current original image 301 itself. Therefore, image blocks originating from the current original image 301 itself can be removed from the predetermined number of image blocks, and the remaining image blocks can be pasted into the current original image 301 to obtain a fused image. The pasted location can be any location other than the target object in the current original image 301 itself; that is, the pasted image blocks do not cover the target object in the current original image 301 itself.

[0074] For the fused image, the edges of image patches can be smoothed, for example, by averaging the edge pixels of the image patches, making the fused image look more like a single image. The smoothed fused image can then be used as an enhanced image 303 of the original image 301. This enhanced image 303 is an instance-level data augmentation image, enabling the target objects of the original image 301 within a batch to be pasted into more backgrounds, increasing the diversity of target objects and improving the model's ability to recognize target objects.

[0075] The generation of adversarial image 312 is explained below.

[0076] According to embodiments of this disclosure, the loss of the original image is determined based on the output of the deep learning model for the original image; gradient information corresponding to the original image is determined based on the loss; and perturbation is added to the original image based on the gradient information to obtain an adversarial image.

[0077] The adversarial image 312 is also generated based on the original image 301. Specifically, for each original image 301, the original image 301 is input into the deep learning model 310 to obtain the output result. The gradient 311 corresponding to the original image 301 is calculated based on the output result, and the gradient 311 is used to add perturbation to the original image 301 to obtain the adversarial image 312.

[0078] Since gradient information reflects the direction of the model's incorrect prediction of the original image, this embodiment uses the gradient information 311 corresponding to the original image 301 to generate an adversarial image 312. Using the adversarial image 312 to train the deep learning model 310 can increase the difficulty of the deep learning model 310 in recognizing the target object in the adversarial image 312, thereby improving the training effect of the deep learning model 310.

[0079] The output corresponding to the original image 301 can include the category and location information (e.g., detection boxes) of the target object in the original image 301. Since the classification result of the original image 301 will not change with cropping or rotation operations, but the bounding box will change with cropping or rotation operations, in order to avoid the problem of inaccurate loss caused by changes in the detection box, the gradient 311 can be calculated using only the loss of the category recognition result to ensure the accuracy of the gradient information.

[0080] Next, the deep learning model 310 is trained using the enhanced image 303 and the adversarial image 312.

[0081] According to embodiments of this disclosure, an enhanced image and an adversarial image are stitched together to obtain a stitched image; the stitched image is input into a deep learning model to obtain the output results of the enhanced image and the adversarial image; the loss of the enhanced image and the loss of the adversarial image are determined based on the output results of the enhanced image and the adversarial image; and the parameters of the deep learning model are adjusted based on the loss of the enhanced image and the loss of the adversarial image.

[0082] For example, a batch may contain four original images, which can generate four enhanced images 303 and four adversarial images 312. These four enhanced images 303 and four adversarial images 312 are then concatenated to obtain a stitched image. It can be understood that the first four images in this stitched image are the enhanced images 303, and the last four are the adversarial images 312. Inputting this stitched image into a deep learning model 310 for training allows the deep learning model 310 to learn the features of both the enhanced images 303 and the adversarial images 312 during the same training process, thus improving the deep learning model 310's feature learning ability.

[0083] For example, inputting the stitched image into the deep learning model 310 yields the output result 304, which includes the output results of each enhanced image 303 and each adversarial image 312. For each enhanced image 303, the loss of the enhanced image 303 can be determined based on its output result, and the loss of the adversarial image 312 can be determined based on its output result. Based on the losses of the enhanced image 303 and the adversarial image 312, the overall loss can be determined, and then the parameters of the deep learning model 310 can be adjusted based on the overall loss to obtain the deep learning model trained in this round.

[0084] The output of enhanced image 303 can include the category and location information of the target object in enhanced image 303, and the output of adversarial image 312 can include the category and location information of the target object in adversarial image 312. Therefore, the loss of enhanced image 303 can include the loss of category recognition result and the loss of location recognition result, and the loss of adversarial image 312 can also include the loss of category recognition result and the loss of location recognition result. The total loss includes the category recognition result loss and the location recognition result loss of enhanced image 303, as well as the category recognition result loss and the location recognition result loss of adversarial image 312.

[0085] This embodiment utilizes enhanced images and adversarial images to perform hybrid training on the deep learning model, which can improve the training effect of the deep learning model and thus improve the accuracy of the deep learning model in recognizing target objects in images.

[0086] The generation of adversarial images will be explained in detail below.

[0087] The following formula (1) can be used to add perturbation to the original image using gradient.

[0088]

[0089] Here, the sign function is a flag function; it is 1 if the value of the sign function is greater than 0, and 0 otherwise. L() represents the loss function, X is the original image, and y is the loss function. true The annotation results for X, To find the gradient function, ∈ represents the preset parameters, and θ represents the parameters of the deep learning model. X adv This represents the interference image obtained after adding interference.

[0090] Since the pixel values ​​of the interference image change after the interference is added, for example, exceeding 255, the pixel values ​​of the interference image can be adjusted again to bring the pixel values ​​into a preset range, for example, the preset range is [0, 255].

[0091] The above operations of adding interference and adjusting pixel values ​​can be repeated to continuously modify the original image. After multiple repetitions, an adversarial image is obtained. Each operation of adding interference and adjusting pixel values ​​can be called an update. The original image can be updated multiple times to obtain the adversarial image.

[0092] According to embodiments of this disclosure, interference is added to the original image based on gradient information to obtain the first interference image, and the pixel values ​​of the first interference image are adjusted to a preset range to obtain the first adjusted image; interference is added to the nth adjusted image based on gradient information to obtain the (n+1)th interference image, and the pixel values ​​of the (n+1)th interference image are adjusted to a preset range to obtain the (n+1)th adjusted image, where n = 1, ..., N-2, and N is an integer greater than 2; interference is added to the (N-1)th adjusted image based on gradient information to obtain the Nth interference image, which serves as the adversarial image.

[0093] The above update process can be represented by the following formula (2).

[0094]

[0095] in, This represents the nth interfering image. This represents the (n+1)th interfering image. The Clip operation is used to adjust the pixel values ​​of the interfering image to a preset range. ∈ represents the preset parameters, and θ represents the parameters of the deep learning model.

[0096] For example, generating an adversarial image requires updating N times using the formula (2) above, where N is an integer greater than 1, such as N = 10. In the first update, the original image is processed according to the formula (1) above to obtain the first interfering image. The pixel values ​​of the first interfering image are adjusted to a preset range to obtain the first adjusted image. In the second update, n = 1, and the second adjusted image is obtained using the formula (2) above. And so on. After obtaining the Nth interfering image in the Nth (n = N-1) update, the Nth interfering image is determined as the adversarial image. That is, in the last update, the Clip operation is not performed because the Clip operation will reduce the accuracy of the interfering image. The Nth interfering image can be determined as the adversarial image without adding the Clip operation in the last update, preserving the accuracy of the adversarial image, which is convenient for the deep learning model to recognize the adversarial image.

[0097] The following section details the use of augmented and adversarial images for hybrid training of deep learning models.

[0098] The adversarial image and the augmented image have different data distributions. The loss function for training the adversarial image and the augmented image together is shown in the following formula (3).

[0099]

[0100] Where L(θ, x, y) true ) represents the loss for image enhancement, max ∈ L(θ, x+∈, y)true ) represents the loss for the adversarial image, and E represents the combined operation of the loss for the augmented image and the loss for the adversarial image, such as averaging or weighted summation. ∈ represents preset parameters, and θ represents the parameters of the deep learning model.

[0101] This loss essentially comprises two parts: the loss for the enhanced image and the loss for the adversarial image. Hybrid training essentially involves combined training under two different distributions. To decompose this hybrid distribution into two simpler distributions, this application improves the structure of the deep learning model.

[0102] According to embodiments of this disclosure, a deep learning model includes at least one statistical layer, each statistical layer including an original statistical sublayer and an auxiliary statistical sublayer; inputting a stitched image into the deep learning model to obtain the output results of an enhanced image and an adversarial image includes: for each statistical layer, inputting the features of the stitched image into the statistical layer, splitting the features of the stitched image into features of the enhanced image and features of the adversarial image; inputting the features of the enhanced image into the original statistical sublayer to obtain feature distribution information of the enhanced image; inputting the features of the adversarial image into the auxiliary statistical sublayer to obtain feature distribution information of the adversarial image; stitching the feature distribution information of the enhanced image and the feature distribution information of the adversarial image together to obtain the output result of the statistical layer; and determining the output result of the enhanced image and the output result of the adversarial image based on the output result of the statistical layer.

[0103] Specifically, deep learning models can generally be divided into a backbone network, a neck network, and a head network. The backbone network is the main network used for feature extraction. The head network is the detection head used to predict the category and location of the target object. The neck network is located between the backbone network and the head network and is used to adjust the features from the backbone network to better adapt to the task requirements. Each of the backbone network, neck network, and head network contains at least one BN (Batch Normalization) layer. The BN layer is used to statistically analyze the distribution of features and can also be called a statistical layer.

[0104] In this embodiment, at each position in the deep learning model where a BN (Batch Normalization) layer is required, an auxiliary BN layer is inserted in parallel with the original BN layer.

[0105] The structural improvements in this implementation will be described in detail below.

[0106] Figure 4A This is a schematic diagram of the processing module containing the BN layer and the corresponding feature distribution information in related technologies.

[0107] like Figure 4A As shown, the processing module 41 can be part of any of the Backbone network, Neck network, and Head network. The processing module 41 includes convolutional layers, BN layers (original BN layers), and ReLU, where ReLU is the activation function. Image enhancement (X) clean ) and adversarial images (X) adv After the stitched image is input into the deep learning model, it passes through the convolutional layer to obtain the stitched features, and the stitched features are input into the BN layer.

[0108] As shown in feature distribution information 410, curve 411 is the enhanced image (X). clean The characteristic distribution curve of the adversarial image (X), curve 412 is the characteristic distribution curve of the adversarial image (X). adv The characteristic distribution curve of the enhanced image (X). clean ) and adversarial images (X) adv The concatenated features of the image (X) are input into an original BN layer, which will enhance the image (X). clean Feature distribution information and adversarial images (X) adv The feature distribution information is fused so that the fused feature distribution information conforms to curve 413.

[0109] It is understandable that image enhancement (X) clean An augmented image (X) is a real image, and its feature distribution information conforms to the natural laws of the real world. An adversarial image, however, contains interference, and its feature distribution information is simulated and does not conform to the feature distribution laws of the real world. clean ) and adversarial images (X) adv The concatenated features of the images are input into the same BN layer (original BN layer). This original BN layer merges the real and simulated distributions into curve 413, which is not the real distribution either. Therefore, training with an original BN layer structure will cause newly arriving images to be statistically assigned to the incorrect distribution curve 413, resulting in model recognition errors.

[0110] Figure 4B This is a schematic diagram of a processing module including a primary BN sublayer and an auxiliary BN sublayer, and corresponding feature distribution information, according to an embodiment of the present disclosure.

[0111] like Figure 4B As shown, the processing module 42 can be part of any of the Backbone network, Neck network, and Head network. The processing module 42 includes convolutional layers, BN layers, and ReLU, with the BN layers including original BN sub-layers and auxiliary BN sub-layers. Image enhancement (X) clean ) and adversarial images (X) advAfter the stitched image is input into the deep learning model, it passes through a convolutional layer to obtain stitched features, which are then input into a batch normalization (BN) layer. The BN layer decomposes the stitched features into an augmented image (X). clean Features of adversarial images (X) adv The features of ) will enhance the image (X) clean The features of the adversarial image (X) are input into the original BN sublayer, and the adversarial image (X) is used as input. adv The feature input is used to assist the Batch Normalization (BN) sublayer. This will enhance the image (X). clean Features of adversarial images (X) adv The features are input into different BN layers for processing.

[0112] As shown in feature distribution information 420, the enhanced image (X) obtained after statistical analysis of the original BN sublayer... clean The feature distribution information of ) follows curve 421, and the adversarial image (X) statistically analyzed by the auxiliary BN sublayer adv The feature distribution information follows curve 422. This embodiment will not generate feature distribution information that follows curve 423 (obtained by fusing curves 421 and 422).

[0113] Enhanced image (X) output from the original BN sublayer clean The feature distribution information and the adversarial image (X) output by the auxiliary BN sublayer adv The feature distribution information is concatenated and then input into ReLU for further processing.

[0114] In this embodiment, an auxiliary BN sublayer is inserted in parallel with the original BN sublayer to enhance the image (X). clean The features of the adversarial image (X) are input into the original BN sublayer, and the adversarial image (X) is used as input. adv The feature input of the auxiliary BN sublayer is used to statistically enhance the image (X). clean The feature distribution information of ) assists the statistical adversarial image (X) of the BN sublayer. adv The deep learning model can collect feature distribution information from both real and simulated images. This expands the knowledge domain of the deep learning model. For newly arrived real images, the original Batch Normalization (BN) sublayer can be selected for feature statistics, avoiding the inclusion of features in incorrect distribution curves and ensuring the accuracy of the deep learning model in identifying target objects.

[0115] Figure 5 This is a schematic diagram of a training method for a deep learning model according to an embodiment of the present disclosure.

[0116] like Figure 5As shown, the original images in a batch may include original image 510 and original image 520. Original image 510 includes target object A1 and target object A2, and original image 520 includes target object B1 and target object B2.

[0117] Instance-level data augmentation of the original images 510 and 520 yields augmented images 511 and 521. For example, pasting an image patch containing target object B1 from the original image 520 into the original image 510 results in augmented image 511. Similarly, pasting an image patch containing target object A2 from the original image 510 into the original image 520 results in augmented image 521. Augmented images 511 and 521 contain a greater variety of target objects.

[0118] Adding perturbation information to the original image 510 yields an adversarial image 512, which can be generated based on the gradient information corresponding to the original image 510. Adding perturbation to the original image 520 yields an adversarial image 522, which can also be generated based on the gradient information corresponding to the original image 520. The adversarial image 512 contains the gradient information of the original image 510, and the adversarial image 522 contains the gradient information of the original image 520.

[0119] Next, the enhanced images 511 and 521, along with the adversarial images 512 and 522, are stitched together to obtain a stitched image. This stitched image is then input into a deep learning model to obtain stitched features. The deep learning model includes at least one batch normalization (BN) layer 530, and each BN layer includes an original BN sublayer and an auxiliary BN sublayer.

[0120] In response to the concatenated feature input to the BN layer 530 of the deep learning model, the features of the enhanced images 511 and 521 are input into the original BN sublayer to obtain the feature statistics of the enhanced images. The features of the adversarial images 512 and 522 are input into the auxiliary BN sublayer to obtain the feature statistics of the adversarial images. The feature statistics of the enhanced images output from the original BN sublayer and the feature statistics of the adversarial images output from the auxiliary BN sublayer are then concatenated and input into the next processing layer until the final output of the deep learning model is obtained.

[0121] The output results may include the category recognition results and location recognition results of enhanced image 511, enhanced image 521, adversarial image 512, and adversarial image 522.

[0122] Based on the differences between the recognition results of the enhanced images 511 and 521 and the adversarial images 512 and 522 and their actual labels, the losses of the enhanced images 511 and 521 and the adversarial images 512 and 522 can be calculated.

[0123] According to embodiments of this disclosure, adjusting the parameters of a deep learning model based on the loss of the enhanced image and the loss of the adversarial image includes: weighting the loss of the enhanced image and the loss of the adversarial image to obtain an overall loss; and adjusting the parameters of the deep learning model based on the overall loss.

[0124] For example, the weights of the losses for enhanced images 511 and 521 are the same, and the weights of the losses for adversarial images 512 and 522 are the same. The overall loss can be calculated using the following formula (4).

[0125]

[0126] in, Let J(θ, x, y) represent the overall loss, α be the weight of the loss for the enhanced image, and (1-α) be the weight of the loss for the adversarial image. J(θ, x, y) represents the loss for the enhanced image, which could be the sum of the losses for enhanced images 511 and 521. This represents the loss for the adversarial images, which could be the sum of the losses for adversarial images 512 and 522. ∈ represents preset parameters, and θ represents the parameters of the deep learning model.

[0127] For example, the value of α can be 0.9, meaning the weight of the loss for enhancing the image is 0.9, and the weight of the loss for adversarial images is 0.1.

[0128] This embodiment utilizes augmented and adversarial images for hybrid training, and weights the losses from augmented and adversarial images to obtain an overall loss. The parameters of the deep learning model are adjusted based on the overall loss, which can improve the recognition performance and adversarial nature of the trained deep learning model, thereby improving the accuracy of the deep learning model in recognizing target objects in images.

[0129] Figure 6 This is a flowchart of an image processing method according to an embodiment of the present disclosure.

[0130] like Figure 6 As shown, the image processing method 600 includes operations S610 to S620.

[0131] The S610 is used to acquire the image to be processed.

[0132] When operating the S620, the image to be processed is input into the deep learning model to obtain the processing result of the image.

[0133] A deep learning model can be a deep learning model trained using the training methods described above.

[0134] The image to be processed can be an image containing a target object, which can include animals such as cats and dogs, objects such as tables and books, as well as pedestrians and vehicles. Inputting the image to be processed into a trained deep learning model yields at least one of the following: the category and location information of the target object in the image. The location information can be a bounding box for the target object.

[0135] This embodiment uses a trained deep learning model to identify the image to be processed, which can improve the accuracy of identifying the category and location information of the target object in the image to be processed.

[0136] According to embodiments of this disclosure, the deep learning model includes at least one statistical layer, and each statistical layer includes an original statistical sublayer and an auxiliary statistical sublayer. Operation S620 includes: for each statistical layer, in response to the input of features of the image to be processed into the statistical layer, inputting the features of the image to be processed into the original statistical sublayer to obtain feature distribution information of the image to be processed, which is used as the output of the statistical layer; and determining the output of the image to be processed based on the output of the statistical layer.

[0137] The Batch Normalization (BN) layer of the deep learning model in this embodiment includes an original BN sublayer and an auxiliary BN sublayer. Features of the augmented image are input into the original BN sublayer. During training, features of the adversarial image are input into the auxiliary BN sublayer. The original BN sublayer statistically analyzes the feature distribution information of the augmented image, while the auxiliary BN sublayer statistically analyzes the feature distribution information of the adversarial image. Therefore, the deep learning model can statistically analyze the feature distribution information of both real and simulated images, thus expanding the knowledge domain of the deep learning model.

[0138] During the model's use, the image to be processed is a real image. Therefore, the features of the image to be processed can be input into the original BN sublayer for feature statistics, ensuring the accuracy of the deep learning model in recognizing the target object.

[0139] Figure 7 This is a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure.

[0140] like Figure 7 As shown, the training device 700 for the deep learning model includes an image patch determination module 701, an enhanced image generation module 702, an adversarial image generation module 703, and a training module 704.

[0141] The image patch determination module 701 is used to determine at least one image patch containing the target object from each of a plurality of original images used for batch training, thereby obtaining an image patch set.

[0142] The enhanced image generation module 702 is used to generate an enhanced image of the original image based on a set of image patches for each original image.

[0143] The adversarial image generation module 703 is used to add perturbations to each original image based on the output of the deep learning model for the original image, so as to obtain an adversarial image of the original image.

[0144] The training module 704 is used to train a deep learning model using the augmented and adversarial images of multiple original images to obtain a trained deep learning model.

[0145] The enhanced image generation module 702 includes an image block determination unit, a pasting unit, and a smoothing processing unit.

[0146] The image block determination unit is used to determine a predetermined number of image blocks from the image block set, wherein the predetermined number is greater than or equal to the minimum number of image blocks and less than or equal to the maximum number of image blocks, and the number of image blocks includes the number of image blocks of each of the multiple original images.

[0147] The pasting unit is used to paste the remaining image blocks (excluding those from the original image) from a predetermined number of image blocks into the original image to obtain a fused image.

[0148] The smoothing unit is used to smooth the edges of image patches in the fused image to obtain an enhanced image of the original image.

[0149] The image block determination module 701 includes a cropping unit and an image block set determination unit.

[0150] The cropping unit is used to determine the boundaries of at least one target object from the original image for each original image, and to crop at least one image block from the original image based on the boundaries.

[0151] The image block set determination unit is used to determine the image block set based on at least one image block corresponding to each original image.

[0152] The adversarial image generation module 703 includes a first loss determination unit, a gradient determination unit, and an adversarial image generation unit.

[0153] The first loss determination unit is used to determine the loss of the original image based on the output of the deep learning model for the original image.

[0154] The gradient determination unit is used to determine the gradient information corresponding to the original image based on the loss.

[0155] The adversarial image generation unit is used to add perturbation to the original image based on gradient information to obtain an adversarial image.

[0156] The adversarial image generation unit includes a first update subunit, a second update subunit, and a third update subunit.

[0157] The first update subunit is used to add interference to the original image based on gradient information to obtain the first interference image, and adjust the pixel values ​​of the first interference image to a preset range to obtain the first adjusted image.

[0158] The second update subunit is used to add interference to the nth adjusted image according to the gradient information to obtain the (n+1)th interference image, and adjust the pixel values ​​of the (n+1)th interference image to a preset range to obtain the (n+1)th adjusted image, where n = 1, ..., N-2, and N is an integer greater than 2.

[0159] The third update subunit is used to add perturbation to the Nth,1st adjusted image based on gradient information to obtain the Nth perturbation image, which serves as the adversarial image.

[0160] The training module 704 includes a stitching unit, a processing unit, a second loss determination unit, and an adjustment unit.

[0161] The stitching unit is used to stitch the enhanced image and the adversarial image together to obtain the stitched image.

[0162] The processing unit is used to input the stitched image into the deep learning model to obtain the output results of the enhanced image and the adversarial image.

[0163] The second loss determination unit is used to determine the loss of the enhanced image and the loss of the adversarial image based on the output results of the enhanced image and the output results of the adversarial image.

[0164] The adjustment unit is used to adjust the parameters of the deep learning model based on the loss of the augmented image and the loss of the adversarial image.

[0165] According to embodiments of this disclosure, the deep learning model includes at least one statistical layer, each statistical layer including an original statistical sublayer and an auxiliary statistical sublayer. The processing unit includes a splitting subunit, a first statistical subunit, a second statistical subunit, a concatenation subunit, and an output result determination subunit.

[0166] The splitting subunit is used to split the features of the stitched image into features of the augmented image and features of the adversarial image in response to the feature input statistical layer of the stitched image.

[0167] The first statistical sub-unit is used to input the features of the enhanced image into the original statistical sub-layer to obtain the feature distribution information of the enhanced image.

[0168] The second statistical sub-unit is used to input the features of the adversarial image into the auxiliary statistical sub-layer to obtain the feature distribution information of the adversarial image.

[0169] The splicing subunit is used to splice the feature distribution information of the enhanced image and the feature distribution information of the adversarial image to obtain the output of the statistical layer.

[0170] The output result determination subunit is used to determine the output result of the enhanced image and the output result of the adversarial image based on the output result of the statistical layer.

[0171] The second loss determination unit is used to weight the loss of the augmented image and the loss of the adversarial image to obtain the overall loss; and to adjust the parameters of the deep learning model based on the overall loss.

[0172] The output of the original image includes the category of the target object in the original image; the output of the enhanced image includes at least one of the category and location information of the target object in the enhanced image; and the output of the adversarial image includes at least one of the category and location information of the target object in the adversarial image.

[0173] Figure 8 This is a block diagram of an image processing apparatus according to an embodiment of the present disclosure.

[0174] like Figure 8 As shown, the image processing device 800 includes an acquisition module 801 and a processing module 802.

[0175] The acquisition module 801 is used to acquire the image to be processed.

[0176] The processing module 802 is used to input the image to be processed into the deep learning model to obtain the processing result of the image to be processed.

[0177] The deep learning model is trained using the training device described above.

[0178] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0179] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0180] like Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0181] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0182] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as at least one of deep learning model training methods and image processing methods. For example, in some embodiments, at least one of the deep learning model training methods and image processing methods can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of at least one of the deep learning model training methods and image processing methods described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured by any other suitable means (e.g., by means of firmware) to perform at least one of a deep learning model training method and an image processing method.

[0183] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0184] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0185] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0186] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0187] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0188] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0189] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0190] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training a deep learning model, comprising: From each of the multiple original images used for batch training, at least one image patch containing the target object is determined to obtain a set of image patches; For each original image, an enhanced image of the original image is generated based on the set of image patches; For each original image, based on the output of the deep learning model for the original image, perturbation is added to the original image to obtain an adversarial image of the original image; as well as The deep learning model is trained using the enhanced and adversarial images of each of the multiple original images to obtain a trained deep learning model. Specifically, the step of adding perturbations to each original image based on the output of the deep learning model to obtain an adversarial image of the original image includes: for each original image, Based on the output of the deep learning model for the original image, determine the loss of the original image; Based on the loss, determine the gradient information corresponding to the original image; and The adversarial image is obtained by adding perturbation to the original image based on the gradient information.

2. The method according to claim 1, wherein, The step of generating an enhanced image of the original image based on the set of image patches for each original image includes: for each original image, A predetermined number of image blocks are determined from the set of image blocks, wherein the predetermined number is greater than or equal to the minimum number of image blocks and less than or equal to the maximum number of image blocks, and the number of image blocks includes the number of image blocks of each of the plurality of original images; The remaining image blocks from the predetermined number of image blocks, excluding those from the original image, are pasted into the original image to obtain a fused image; and The edges of the image blocks in the fused image are smoothed to obtain the enhanced image of the original image.

3. The method according to claim 1 or 2, wherein, The step of determining at least one image patch containing the target object from each of a plurality of original images to obtain a set of image patches includes: For each original image, determine the boundary of at least one target object from the original image, and crop at least one image patch from the original image based on the boundary; and The set of image blocks is determined based on at least one image block corresponding to each original image.

4. The method according to claim 1, wherein, The step of adding perturbation to the original image according to the gradient to obtain the adversarial image includes: Based on the gradient information, interference is added to the original image to obtain the first interference image, and the pixel values ​​of the first interference image are adjusted to a preset range to obtain the first adjusted image; According to the gradient information, interference is added to the nth adjusted image to obtain the (n+1)th interference image, and the pixel values ​​of the (n+1)th interference image are adjusted to the preset range to obtain the (n+1)th adjusted image, where n=1, ..., N-2, and N is an integer greater than 2; Based on the gradient information, interference is added to the (N-1)th adjusted image to obtain the Nth interference image, which serves as the adversarial image.

5. The method according to claim 1, wherein, The step of training the deep learning model using the enhanced and adversarial images of each of the multiple original images to obtain the trained deep learning model includes: The enhanced image and the adversarial image are stitched together to obtain a stitched image; The stitched image is input into the deep learning model to obtain the output of the enhanced image and the output of the adversarial image; Based on the output results of the enhanced image and the adversarial image, determine the loss of the enhanced image and the loss of the adversarial image; and The parameters of the deep learning model are adjusted based on the loss of the enhanced image and the loss of the adversarial image.

6. The method according to claim 5, wherein, The deep learning model includes at least one statistical layer, and each statistical layer includes an original statistical sublayer and an auxiliary statistical sublayer; the step of inputting the stitched image into the deep learning model to obtain the output result of the enhanced image and the output result of the adversarial image includes: for each statistical layer, In response to the feature input of the stitched image into the statistical layer, the features of the stitched image are split into features of the enhanced image and features of the adversarial image; The features of the enhanced image are input into the original statistical sublayer to obtain the feature distribution information of the enhanced image; The features of the adversarial image are input into the auxiliary statistical sublayer to obtain the feature distribution information of the adversarial image; The feature distribution information of the enhanced image and the feature distribution information of the adversarial image are concatenated to obtain the output result of the statistical layer; and Based on the output of the statistical layer, the output of the enhanced image and the output of the adversarial image are determined.

7. The method according to claim 5, wherein, The step of adjusting the parameters of the deep learning model based on the loss of the enhanced image and the loss of the adversarial image includes: The loss of the enhanced image and the loss of the adversarial image are weighted to obtain the overall loss; and The parameters of the deep learning model are adjusted based on the overall loss.

8. The method according to any one of claims 5 to 7, wherein, The output of the original image includes the category of the target object in the original image, the output of the enhanced image includes at least one of the category and location information of the target object in the enhanced image, and the output of the adversarial image includes at least one of the category and location information of the target object in the adversarial image.

9. An image processing method, comprising: Obtain the image to be processed; as well as The image to be processed is input into a deep learning model to obtain the processing result of the image to be processed; The deep learning model is trained using the method described in any one of claims 1 to 8.

10. The method according to claim 9, wherein, The deep learning model includes at least one statistical layer, and each statistical layer includes an original statistical sublayer and an auxiliary statistical sublayer; the process of inputting the image to be processed into the deep learning model to obtain the processing result of the image to be processed includes: for each statistical layer, In response to the input of the features of the image to be processed into the statistical layer, the features of the image to be processed are input into the original statistical sub-layer to obtain the feature distribution information of the image to be processed, which is used as the output result of the statistical layer. The output result of the image to be processed is determined based on the output result of the statistical layer.

11. The method according to claim 9 or 10, wherein, The processing result of the image to be processed includes at least one of the category and location information of the target object in the image to be processed.

12. A training device for a deep learning model, comprising: The image patch determination module is used to determine at least one image patch containing the target object from each of a plurality of original images used for batch training, thereby obtaining a set of image patches; An enhanced image generation module is used to generate an enhanced image of each original image based on the set of image blocks. An adversarial image generation module is used to add perturbations to each original image based on the output of the deep learning model for the original image, thereby obtaining an adversarial image of the original image. as well as The training module is used to train the deep learning model using the augmented and adversarial images of the multiple original images respectively, so as to obtain the trained deep learning model. The adversarial image generation module includes: The first loss determination unit is used to determine the loss of the original image based on the output of the deep learning model for the original image; A gradient determination unit is configured to determine gradient information corresponding to the original image based on the loss; and An adversarial image generation unit is used to add perturbation to the original image based on the gradient information to obtain the adversarial image.

13. The apparatus according to claim 12, wherein, The enhanced image generation module includes: An image block determination unit is configured to determine a predetermined number of image blocks from the image block set, wherein the predetermined number is greater than or equal to a minimum number of image blocks and less than or equal to a maximum number of image blocks, and the number of image blocks includes the number of image blocks of each of the plurality of original images; A pasting unit is configured to paste the remaining image blocks (excluding those from the original image) from the predetermined number of image blocks into the original image to obtain a fused image; and A smoothing processing unit is used to smooth the edges of image blocks in the fused image to obtain an enhanced image of the original image.

14. The apparatus according to claim 12 or 13, wherein, The image patch set determination module includes: A cropping unit is configured to, for each original image, determine the boundary of at least one target object from the original image, and crop at least one image patch from the original image based on the boundary; and An image block set determination unit is used to determine the image block set based on at least one image block corresponding to each original image.

15. The apparatus according to claim 12, wherein, The adversarial image generation unit includes: The first update subunit is used to add interference to the original image according to the gradient information to obtain the first interference image, and adjust the pixel values ​​of the first interference image to a preset range to obtain the first adjusted image; The second update subunit is used to add interference to the nth adjusted image according to the gradient information to obtain the (n+1)th interference image, and adjust the pixel values ​​of the (n+1)th interference image to the preset range to obtain the (n+1)th adjusted image, where n=1, ..., N-2, and N is an integer greater than 2; The third update subunit is used to add interference to the (N-1)th adjusted image according to the gradient information to obtain the Nth interference image, which serves as the adversarial image.

16. The apparatus according to claim 12, wherein, The training module includes: A stitching unit is used to stitch the enhanced image and the adversarial image together to obtain a stitched image; The processing unit is used to input the stitched image into the deep learning model to obtain the output result of the enhanced image and the output result of the adversarial image; The second loss determination unit is configured to determine the loss of the enhanced image and the loss of the adversarial image based on the output results of the enhanced image and the output results of the adversarial image; and An adjustment unit is used to adjust the parameters of the deep learning model based on the loss of the enhanced image and the loss of the adversarial image.

17. The apparatus according to claim 16, wherein, The deep learning model includes at least one statistical layer, and each statistical layer includes an original statistical sublayer and an auxiliary statistical sublayer; the processing unit includes: A splitting subunit is used to split the features of the stitched image into features of the enhanced image and features of the adversarial image in response to the feature input of the stitched image into the statistical layer. The first statistical subunit is used to input the features of the enhanced image into the original statistical sublayer to obtain the feature distribution information of the enhanced image; The second statistical subunit is used to input the features of the adversarial image into the auxiliary statistical sublayer to obtain the feature distribution information of the adversarial image. A concatenation subunit is used to concatenate the feature distribution information of the enhanced image and the feature distribution information of the adversarial image to obtain the output result of the statistical layer; and The output result determination subunit is used to determine the output result of the enhanced image and the output result of the adversarial image based on the output result of the statistical layer.

18. The apparatus according to claim 16, wherein, The second loss determination unit is used to perform weighted processing on the loss of the enhanced image and the loss of the adversarial image to obtain the overall loss; and to adjust the parameters of the deep learning model according to the overall loss.

19. The apparatus according to any one of claims 16 to 18, wherein, The output of the original image includes the category of the target object in the original image, the output of the enhanced image includes at least one of the category and location information of the target object in the enhanced image, and the output of the adversarial image includes at least one of the category and location information of the target object in the adversarial image.

20. An image processing apparatus, comprising: The acquisition module is used to acquire the image to be processed; as well as The processing module is used to input the image to be processed into a deep learning model to obtain the processing result of the image to be processed; The deep learning model is trained using the apparatus according to any one of claims 12 to 19.

21. The apparatus according to claim 20, wherein, The deep learning model includes at least one statistical layer, and each statistical layer includes an original statistical sublayer and an auxiliary statistical sublayer. The processing module is configured to, in response to the input of the features of the image to be processed into the statistical layer, input the features of the image to be processed into the original statistical sublayer to obtain the feature distribution information of the image to be processed, which is used as the output result of the statistical layer; and determine the output result of the image to be processed based on the output result of the statistical layer.

22. The apparatus according to claim 20 or 21, wherein, The processing result of the image to be processed includes at least one of the category and location information of the target object in the image to be processed.

23. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 11.

25. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Remote sensing image target sample enhancement method

    CN114758123A