Method, apparatus, computer readable medium and program product for generating image

By segmenting the target object and fusing it with multiple background images, multiple fusion images are generated, which solves the problem of limited number of training samples and improves the recognition and generalization capabilities of the model.

CN120013973APending Publication Date: 2025-05-16SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510042221.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, when using deep learning to identify pictures involving inappropriate content, the model recognition performance does not meet the standards due to the limited number of training samples available.

Method used

By segmenting the target object, obtaining the object image, and fusing it with multiple background images, multiple fused images containing the target object are generated, thereby augmenting the sample data.

Benefits of technology

It improves the model's ability to identify target objects, and can accurately identify target objects under different backgrounds and angles, which enhances the generalization ability of the model and the robustness and accuracy of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013973A_ABST
    Figure CN120013973A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for generating an image, electronic equipment, a computer readable medium and a computer program product. The method comprises the following steps: performing segmentation processing on a plurality of target images to obtain object images corresponding to one or more target objects contained in the target images; for each target image, acquiring a plurality of matched background images based on a target object contained in the target image; and fusing the target object and each background image to generate a plurality of new images containing the target object. According to the method, the target object contained in the sample is segmented, the segmented target object is fused with the plurality of matched background images, and the conditions of the target object in various real environment backgrounds are simulated, so that the sample image generation efficiency is improved, and a limited number of sample images are greatly expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, computer-readable medium and computer program product for generating an image. Background Art

[0002] When deep learning technology is used to identify images with inappropriate content, the model often fails to achieve the desired learning effect due to the limited number of available training samples, resulting in substandard recognition performance. Currently, the following two data augmentation techniques are mainly used to address this challenge: one is to perform operations such as oversampling, rotating, flipping, and adjusting the lighting of the image; the other is to use image generation models such as generative adversarial networks (GANs) to expand the data set.

[0003] However, the existing data enhancement technology is essentially a simple transformation of the existing sample set, which cannot effectively help the model learn more samples. Although image generation models such as generative adversarial networks can generate more accurate images, their operation also requires a large amount of sample data, which is in conflict with the current situation of insufficient sample numbers. Summary of the invention

[0004] Multiple aspects of the present application provide a method, an apparatus, an electronic device, a computer-readable medium, and a computer program product for generating an image.

[0005] In one aspect of the present application, a method for generating an image is provided, wherein the method comprises:

[0006] Segmenting the target image to obtain an object image corresponding to the target object, wherein the target object is contained in the target image;

[0007] Acquire multiple background images corresponding to the target object;

[0008] The object image and the multiple background images are fused respectively to generate multiple fused images containing the target object.

[0009] In one aspect of the present application, a device for generating an image is provided, wherein the device comprises:

[0010] A device for performing segmentation processing on a target image to obtain an object image corresponding to a target object, wherein the target object is contained in the target image;

[0011] Means for acquiring a plurality of background images corresponding to a target object;

[0012] A device for fusing the object image with the multiple background images respectively to generate multiple fused images containing the target object.

[0013] Another aspect of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of an embodiment of the present application.

[0014] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.

[0015] Another aspect of the present application provides a computer program product, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.

[0016] In the solution provided in the embodiment of the present application, by segmenting the target object contained in the sample and fusing the segmented object image with multiple matching background images, the situation of the target object in various real environment backgrounds is simulated, thereby improving the efficiency of generating sample images, and achieving a large number of expansions of a limited number of sample images, thereby meeting the demand for generating a large number of training samples from a small number of image samples; by using edge segmentation and data enhancement technology in generating sample images input to the model, the model's recognition ability of the target object is improved, and the target object can be accurately identified under different backgrounds and angles, thereby improving the generalization ability of the model, and improving the robustness and accuracy of the model in performing tasks such as classification, segmentation or target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0018] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0019] Figure 1 A schematic diagram of a process for generating an image provided by an embodiment of the present application is shown;

[0020] Figure 2(a) to Figure 2(f) A schematic diagram showing an exemplary image according to an embodiment of the present application;

[0021] Figure 3(a) to Figure 3(e) A schematic diagram of an exemplary method flow according to an embodiment of the present application is shown;

[0022] Figure 4 A schematic diagram of the structure of a device for generating an image provided by an embodiment of the present application is shown;

[0023] Figure 5 A schematic diagram of the structure of a device suitable for implementing the solution in the embodiment of the present application is shown.

[0024] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0026] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPU), input / output interface, network interface and memory.

[0027] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0028] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer program instructions, data structures, modules of programs or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0029] Figure 1A schematic flow chart of a method for generating an image provided by an embodiment of the present application is shown. The method at least includes step S101, step S102 and step S103.

[0030] In actual scenarios, the execution subject of the method can be a network device, or an application running on a network device, wherein the network device includes but is not limited to a network host, a single network server, a plurality of network server sets, or a collection of computers based on cloud computing, and can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, wherein cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.

[0031] Since it is often costly and time-consuming to obtain a large number of high-quality, diverse sample images, the method of the embodiment of the present application can expand the sample images used to train the model. The method of the embodiment of the present application can expand a limited number of sample images by accurately segmenting the target objects from the original image and integrating these target objects into various backgrounds to obtain fused images integrated into various backgrounds.

[0032] Reference Figure 1 In step S101, the target image is segmented to obtain an object image corresponding to the target object, and the target object is included in the target image.

[0033] The target image is a sample image used for training the model. For example, in a risk prediction scenario, sample images containing predetermined risk elements that need to be identified can be obtained as target images.

[0034] The target object is an object separated from the background in the sample image, and the separated object is the object that the model needs to recognize and understand. The target object includes but is not limited to specific people, animals, objects, etc.

[0035] The segmentation process is used to segment the target object and the background in the target image. The segmentation result information obtained by the segmentation process may be a grayscale image, the grayscale image includes a foreground area and a background area, and the foreground area corresponds to the target object.

[0036] Those skilled in the art should be familiar with the fact that segmentation processing can be performed using a variety of image segmentation algorithms. Those skilled in the art can select a suitable image segmentation algorithm to perform segmentation processing on the target image based on actual needs, which will not be described in detail here.

[0037] According to one embodiment, for each target image, a segmentation process is performed using the Labelme tool to obtain a corresponding label map, wherein each pixel in the label map is marked as belonging to the target object or the background. Then, based on the obtained label map, a computer vision library, such as OpenCV, is used to extract the target object.

[0038] Among them, Labelme is an open source image annotation tool that allows users to draw shapes such as rectangles and polygons on images to mark the target objects in the image. With Labelme, users can interactively segment images and mark the exact location of the target objects.

[0039] Among them, OpenCV provides a variety of image processing and analysis tools, which can use label graphs to locate and extract target objects. For example, the contour of the target object can be identified by finding connected components, and then the pixel values ​​of these areas can be extracted to obtain the object image of the target object.

[0040] According to one embodiment, in step S101, edge detection is performed on the target object in the target image, for example, edge detection algorithms such as Canny are used to determine the edge contour of the target object. Then, according to the edge detection result, the object image corresponding to the target object is segmented along the edge contour. The segmentation method by performing edge detection on the target object helps the model to better understand the shape and contour of the target object, so as to more accurately identify the target in practical applications.

[0041] According to one embodiment, the method performs rotation processing on the object image to obtain object images of the target object corresponding to multiple angles.

[0042] It should be noted that when using Labelme to rotate an object image, it is necessary to ensure that the object image (img) of the target object and the label image (label) of the corresponding pixel position are rotated at the same angle. This is because the label image contains the precise location information of the target object. If only the img image is rotated but the label image is not rotated, the original image extracted from the label image will not be able to correctly match the rotated object image of the target object.

[0043] The method of this embodiment increases the diversity of samples by rotating the segmented target object. This data enhancement method can help the model learn the appearance of the target object at different angles, thereby improving the model's ability to recognize the target object. Even when the target object changes at multiple angles, the model can still accurately identify the target object.

[0044] Continue to refer to Figure 1To explain, in step S102, a plurality of background images corresponding to the target object are obtained.

[0045] According to one embodiment, step S102 includes step S1021 and step S1022.

[0046] In step S1021, a background scene matching the target object is determined.

[0047] Optionally, the method may determine the background scene matching the target object based on a preset correspondence between the target object and the background elements. For example, based on the correspondence between the target object and the background elements preset by relevant personnel, if the target object is a ship, the matching background scene is the sea, a lake or a port.

[0048] Optionally, according to the present embodiment, the target object corresponds to a specific business scenario, and the method determines a background scene matching the target object based on the business scenario. For example, if the business scenario corresponding to the target object is to identify whether there are scratches on a batch of batteries, then the selected background scene is the actual application of the battery.

[0049] Optionally, the method determines the background scene matching the target object by performing scene analysis on the target image. Specifically, the method can use machine learning methods, such as neural networks, support vector machines, random forests, etc., to analyze and predict scenes where the target object may appear. This approach more accurately predicts the background scene matching the target object by identifying patterns and trends from a large amount of data.

[0050] In step S1022, a background image corresponding to the background scene is obtained.

[0051] Specifically, based on the background scenes obtained by scene analysis, background images matching these background scenes are collected or created.

[0052] Among them, the technical personnel in this field should be familiar with that there are many ways to collect or create background images that match these background scenes, for example, searching for images that match the required background scenes in a specific image database based on the background scenes obtained by scene analysis. Alternatively, creating the required background images through drawing software, etc., will not be repeated here.

[0053] Continue to refer to Figure 1 To explain, in step S103, the object image and the plurality of background images are fused respectively to generate a plurality of fused images including the target object.

[0054] According to one embodiment, the method uses a Poisson fusion algorithm to fuse the object image with the multiple background images respectively to generate multiple fused images containing the target object. For example, for N target objects and M background images, the Poisson fusion algorithm can generate N*M different fused images.

[0055] Wherein, the method synchronously generates segmentation result information or target detection information of the target object during the process of fusion processing using the Poisson fusion algorithm.

[0056] Among them, the Poisson fusion algorithm seamlessly blends the object image into the background image while maintaining the consistency of illumination, color and texture in the fusion area. The seamless splicing principle of the Poisson fusion algorithm is achieved by optimizing the Laplace equation, which requires that the fused object image and the background image have the same gradient in the fusion area to achieve consistency in light and shadow texture, thereby achieving a "seamless" fusion effect.

[0057] Optionally, the method performs adjustment processing on the object image or the multiple background images in step S103, and performs fusion processing on the adjusted object image and the multiple background images respectively, so as to increase the diversity of changes of the target object.

[0058] The adjustment process includes but is not limited to at least one of the following:

[0059] 1) Adjust the position of the object image in the background image; specifically, by defining an area in the background image, the target object can appear at any position in the area;

[0060] 2) In the fusion area of ​​the object image and the background image, set and adjust the weight of the object image and the weight of the background image. For example, set the weight of the object image to w and the weight of the background image to 1-w. Since the target object needs to cover a specific area of ​​the background image, the fusion effect can be optimized by adjusting the weight;

[0061] 3) Adjust the rotation angle of the object image to introduce more variations.

[0062] According to the first example of this application, referring to Figure 2(a) to Figure 2(e) An exemplary image is shown in . A target image containing a target object and a label map corresponding to pixel positions are prepared so as to extract the target object from the background through segmentation processing. Among them, the target image is shown in FIG2(a), and the label map is shown in FIG2(b). Next, a background image for fusion is prepared as shown in FIG2(c), and then, the object image of the target object is fused with the background image by performing a Poisson fusion operation to obtain a fused image as shown in FIG2(d), and a label map corresponding to the fused image is generated as shown in FIG2(e).

[0063] According to one embodiment, the method further includes step S104 and step S105.

[0064] In step S104, a cleaning process is performed on the plurality of fused images to select fused images that meet the sample standard as sample images.

[0065] The method performs cleaning processing on a plurality of fused images to select fused images that meet the sample standard as sample images, including:

[0066] 1) Extract the same area from the target image and the fused image respectively, the area is located at the junction of the object image and the background; then, obtain the average pixel values ​​of the two areas and compare them. If the difference between the average pixel values ​​of the two areas is less than the preset threshold, the corresponding fused image is discarded. Among them, if the difference between the average pixel values ​​of the two areas is less than the preset threshold, it indicates that the object image and the background in the fused image are over-fused and do not meet the sample standard, and the fused image should be discarded.

[0067] 2) Perform dilation processing on the target image and the target image in the fused image respectively to obtain the dilation zone area; obtain the average pixel value of the dilation zone area in the target image and the fused image respectively, and if the difference between the average pixel values ​​of the two dilation zone areas is greater than a preset threshold, the corresponding fused image is discarded. Among them, if the difference between the average pixel values ​​of the two dilation zone areas is greater than the preset threshold, it indicates that the boundary obtained when the object image and the background are fused in the fused image is too obvious and does not meet the sample standard, and the fused image should be discarded.

[0068] For example, the target image and the target image in the fused image are dilated to obtain a gray dilated band as shown in FIG2(f).

[0069] In step S105, a plurality of sample images that have been cleaned are input into the model for training.

[0070] Among them, the model can be used to perform various tasks such as classification, segmentation, target detection, etc. Since the sample image is generated by fusing the object image of the target object with the background image, the process simultaneously generates the segmentation information or target detection information of the target object, which can adapt to the classification, segmentation and target detection tasks of the deep learning model.

[0071] After the fusion map is fed into the deep learning model, the model optimizes its weights and parameters by learning the features in the image.

[0072] The following describes the classification task, segmentation task, and target detection task performed by the model when Labelme is used for segmentation in the embodiment of the present application:

[0073] 1) Classification task: If the subsequent deep learning model is used for classification tasks, then only the fusion map needs to be used. The goal of the classification task is to identify the object category in the image, so the model needs to learn features that distinguish different categories.

[0074] 2) Segmentation task: If the model is used for segmentation tasks, then in addition to the fused image, the original label map is also required. The goal of the segmentation task is to accurately identify which object each pixel in the image belongs to, so accurate pixel-level annotation information is required.

[0075] 3) Object detection task: For the object detection task, it is necessary to fuse the image and the BoundingBox of the target object. The BoundingBox refers to the coordinates of the minimum enclosing rectangle of the target element in the label image, which provides the model with the location information of the target object in the image.

[0076] Optionally, the fused image can be further subjected to traditional data augmentation operations in the deep learning model, such as adjusting lighting, rotation, cropping, etc. These operations can increase the diversity of samples, prevent the model from overfitting, and improve the model's ability to recognize target objects under different conditions.

[0077] The process of the method in the embodiment of the present application is explained below with reference to an example.

[0078] Referring to FIG3(a), labellme is used to segment the target image shown in the figure, and the illegal elements contained in the target image are obtained as the target object. Among them, the object image of the target object extracted by labellme and the object images of the target object at different angles are obtained by generating a rotation random factor (i.e., +-a angle rotation), as shown in FIG3(a). Then, referring to FIG3(b), multiple background images matching the target object are obtained. Then, referring to FIG3(c), the object images of the target object at different angles are fused with multiple background images respectively by the Poisson fusion algorithm to obtain multiple fused images (represented as Figures Ae to Ce). Then, referring to FIG3(d), multiple fused images are cleaned and images with poor fusion effects are eliminated. The fused images obtained after cleaning are shown in FIG3(e). These fused images can be used as sample images input to the model for training.

[0079] According to the method of the embodiment of the present application, by segmenting the target object contained in the sample and fusing the segmented object image of the target object with multiple matching background images, the situation of the target object in various real environment backgrounds is simulated, thereby improving the efficiency of generating sample images, and achieving a large number of expansions of a limited number of sample images, thereby meeting the demand for generating a large number of training samples from a small number of image samples; by using edge segmentation and data enhancement technology in generating sample images input to the model, the model's recognition ability of the target object is improved, and the target object can be accurately identified under different backgrounds and angles, thereby improving the generalization ability of the model, and improving the robustness and accuracy of the model in performing tasks such as classification, segmentation or target detection.

[0080] Figure 4 A schematic structural diagram of a device for generating an image provided in an embodiment of the present application is shown.

[0081] The device includes: a device for segmenting a target image to obtain an object image corresponding to the target object (hereinafter referred to as "image segmentation device 101"), a device for acquiring multiple background images corresponding to the target object (hereinafter referred to as "background acquisition device 102"), and a device for fusing the object image with the multiple background images respectively to generate multiple fused images containing the target object (hereinafter referred to as "fusion processing device 103").

[0082] Reference Figure 4 The image segmentation device 101 performs segmentation processing on the target image to obtain an object image corresponding to the target object, and the target object is included in the target image.

[0083] The target image is a sample image used for training the model. For example, in a risk prediction scenario, sample images containing predetermined risk elements that need to be identified can be obtained as target images.

[0084] The target object is an object separated from the background in the sample image, and the separated object is the object that the model needs to recognize and understand. The target object includes but is not limited to specific people, animals, objects, etc.

[0085] The segmentation process is used to segment the target object and the background in the target image. The segmentation result information obtained by the segmentation process may be a grayscale image, the grayscale image includes a foreground area and a background area, and the foreground area corresponds to the target object.

[0086] Those skilled in the art should be familiar with the fact that segmentation processing can be performed using a variety of image segmentation algorithms. Those skilled in the art can select a suitable image segmentation algorithm to perform segmentation processing on the target image based on actual needs, which will not be described in detail here.

[0087] According to one embodiment, for each target image, the image segmentation device 101 uses the Labelme tool to perform segmentation processing to obtain a corresponding label map. Each pixel in the label map is marked as belonging to the target object or the background. Then, based on the obtained label map, a computer vision library such as OpenCV is used to extract the target object.

[0088] Among them, Labelme is an open source image annotation tool that allows users to draw shapes such as rectangles and polygons on images to mark the target objects in the image. With Labelme, users can interactively segment images and mark the exact location of the target objects.

[0089] Among them, OpenCV provides a variety of image processing and analysis tools, which can use label graphs to locate and extract target objects. For example, the contour of the target object can be identified by finding connected components, and then the pixel values ​​of these areas can be extracted to obtain the object image of the target object.

[0090] According to one embodiment, the image segmentation device 101 performs edge detection on the target object in the target image, for example, using an edge detection algorithm such as Canny to determine the edge contour of the target object. Then, according to the edge detection result, the object image corresponding to the target object is segmented along the edge contour. The method of segmenting by performing edge detection on the target object helps the model to better understand the shape and contour of the target object, so as to more accurately identify the target in practical applications.

[0091] According to one embodiment, the device performs rotation processing on the object image to obtain object images of the target object corresponding to multiple angles.

[0092] It should be noted that when using Labelme to rotate an object image, it is necessary to ensure that the object image (img) of the target object and the label image (label) of the corresponding pixel position are rotated at the same angle. This is because the label image contains the precise location information of the target object. If only the img image is rotated but the label image is not rotated, the original image extracted from the label image will not be able to correctly match the rotated object image of the target object.

[0093] The method of this embodiment increases the diversity of samples by rotating the segmented target object. This data enhancement method can help the model learn the appearance of the target object at different angles, thereby improving the model's ability to recognize the target object. Even when the target object changes at multiple angles, the model can still accurately identify the target object.

[0094] Continue to refer to Figure 4 To explain, the background acquisition device 102 acquires a plurality of background images corresponding to the target object.

[0095] According to one embodiment, the background acquisition device 102 includes a scene determination device and a scene image acquisition device.

[0096] The scene determining device determines a background scene that matches the target object.

[0097] Optionally, the scene determination device may determine the background scene that matches the target object based on a preset correspondence between the target object and the background elements. For example, based on the correspondence between the target object and the background elements preset by relevant personnel, if the target object is a ship, the matching background scene is the sea, a lake or a port.

[0098] Optionally, according to this embodiment, the target object corresponds to a specific business scenario, and the scenario determination device determines a background scene matching the target object based on the business scenario. For example, if the business scenario corresponding to the target object is to identify whether there are scratches on a batch of batteries, then the selected background scene is the actual application of the battery.

[0099] Optionally, the scene determination device determines the background scene matching the target object by performing scene analysis on the target image. Specifically, the method can use machine learning methods, such as neural networks, support vector machines, random forests, etc., to analyze and predict scenes where the target object may appear. This method can more accurately predict the background scene matching the target object by identifying patterns and trends from a large amount of data.

[0100] The scene image acquisition device acquires a background image corresponding to the background scene.

[0101] Specifically, the scene image acquisition device collects or creates background images matching these background scenes based on the background scenes obtained by scene analysis.

[0102] Among them, the technical personnel in this field should be familiar with that the scene image acquisition device can collect or create background images matching these background scenes in a variety of ways, for example, searching for images matching the required background scenes in a specific image database based on the background scenes obtained by scene analysis. Alternatively, the required background images can be created through drawing software, etc., which will not be described in detail here.

[0103] Continue to refer to Figure 4 To explain, the fusion processing device 103 performs fusion processing on the object image and the multiple background images respectively to generate multiple fused images including the target object.

[0104] According to one embodiment, the fusion processing device 103 uses a Poisson fusion algorithm to fuse the object image with the multiple background images respectively to generate multiple fused images containing the target object. For example, for N target objects and M background images, N*M different fused images can be generated by the Poisson fusion algorithm.

[0105] The fusion processing device 103 synchronously generates segmentation result information or target detection information of the target object during the fusion processing process using the Poisson fusion algorithm.

[0106] Among them, the Poisson fusion algorithm seamlessly blends the object image into the background image while maintaining the consistency of illumination, color and texture in the fusion area. The seamless splicing principle of the Poisson fusion algorithm is achieved by optimizing the Laplace equation, which requires that the fused object image and the background image have the same gradient in the fusion area to achieve consistency in light and shadow texture, thereby achieving a "seamless" fusion effect.

[0107] Optionally, the fusion processing device 103 performs adjustment processing on the object image or the multiple background images, and performs fusion processing on the adjusted object image and the multiple background images respectively, so as to increase the diversity of changes of the target object.

[0108] The adjustment process includes but is not limited to at least one of the following:

[0109] 1) Adjust the position of the object image in the background image; specifically, by defining an area in the background image, the target object can appear at any position in the area;

[0110] 2) In the fusion area of ​​the object image and the background image, set and adjust the weight of the object image and the weight of the background image. For example, set the weight of the object image to w and the weight of the background image to 1-w. Since the target object needs to cover a specific area of ​​the background image, the fusion effect can be optimized by adjusting the weight;

[0111] 3) Adjust the rotation angle of the object image to introduce more variations.

[0112] According to the first example of this application, referring to Figure 2(a) to Figure 2(e) An exemplary image is shown in . A target image containing a target object and a label map corresponding to pixel positions are prepared so as to extract the target object from the background through segmentation processing. Among them, the target image is shown in FIG2(a), and the label map is shown in FIG2(b). Next, a background image for fusion is prepared as shown in FIG2(c), and then, the object image of the target object is fused with the background image by performing a Poisson fusion operation to obtain a fused image as shown in FIG2(d), and a label map corresponding to the fused image is generated as shown in FIG2(e).

[0113] According to one embodiment, the device further comprises a cleaning treatment device and a model input device.

[0114] The cleaning processing device performs cleaning processing on the plurality of fused images to select fused images that meet the sample standard as sample images.

[0115] The device performs cleaning processing on a plurality of fused images to select fused images that meet the sample standard as sample images, including:

[0116] 1) Extract the same area from the target image and the fused image respectively, the area is located at the junction of the object image and the background; then, obtain the average pixel values ​​of the two areas and compare them. If the difference between the average pixel values ​​of the two areas is less than the preset threshold, the corresponding fused image is discarded. Among them, if the difference between the average pixel values ​​of the two areas is less than the preset threshold, it indicates that the object image and the background in the fused image are over-fused and do not meet the sample standard, and the fused image should be discarded.

[0117] 2) Perform dilation processing on the target image and the target image in the fused image respectively to obtain the dilation zone area; obtain the average pixel value of the dilation zone area in the target image and the fused image respectively, and if the difference between the average pixel values ​​of the two dilation zone areas is greater than a preset threshold, the corresponding fused image is discarded. Among them, if the difference between the average pixel values ​​of the two dilation zone areas is greater than the preset threshold, it indicates that the boundary obtained when the object image and the background are fused in the fused image is too obvious and does not meet the sample standard, and the fused image should be discarded.

[0118] For example, the target image and the target image in the fused image are dilated to obtain a gray dilated band as shown in FIG2(f).

[0119] The model input device inputs a plurality of sample images that have been cleaned into the model for training.

[0120] Among them, the model can be used to perform various tasks such as classification, segmentation, target detection, etc. Since the sample image is generated by fusing the object image of the target object with the background image, the process simultaneously generates the segmentation information or target detection information of the target object, which can adapt to the classification, segmentation and target detection tasks of the deep learning model.

[0121] After the fusion map is fed into the deep learning model, the model optimizes its weights and parameters by learning the features in the image.

[0122] The following describes the classification task, segmentation task, and target detection task performed by the model when Labelme is used for segmentation in the embodiment of the present application:

[0123] 1) Classification task: If the subsequent deep learning model is used for classification tasks, then only the fusion map needs to be used. The goal of the classification task is to identify the object category in the image, so the model needs to learn features that distinguish different categories.

[0124] 2) Segmentation task: If the model is used for segmentation tasks, then in addition to the fused image, the original label map is also required. The goal of the segmentation task is to accurately identify which object each pixel in the image belongs to, so accurate pixel-level annotation information is required.

[0125] 3) Object detection task: For the object detection task, it is necessary to fuse the image and the BoundingBox of the target object. The BoundingBox refers to the coordinates of the minimum enclosing rectangle of the target element in the label image, which provides the model with the location information of the target object in the image.

[0126] Optionally, the fused image can be further subjected to traditional data augmentation operations in the deep learning model, such as adjusting lighting, rotation, cropping, etc. These operations can increase the diversity of samples, prevent the model from overfitting, and improve the model's ability to recognize target objects under different conditions.

[0127] According to the device of the embodiment of the present application, by segmenting the target object contained in the sample and fusing the segmented object image of the target object with multiple matching background images, the situation of the target object in various real environment backgrounds is simulated, thereby improving the efficiency of generating sample images, and achieving a large number of expansions of a limited number of sample images, thereby meeting the demand for generating a large number of training samples from a small number of image samples; by using edge segmentation and data enhancement technology in generating sample images input to the model, the model's recognition ability of the target object is improved, and the target can be accurately identified under different backgrounds and angles, thereby improving the generalization ability of the model, and improving the robustness and accuracy of the model in performing tasks such as classification, segmentation or target detection.

[0128] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application, and the method corresponding to the electronic device may be the method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.

[0129] The electronic device may be a user device, or a device formed by integrating a user device and a network device through a network, or may be an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablet computers, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, which can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.

[0130] Figure 5 The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown, and the device 1200 includes a central processing unit (CPU, Central Processing Unit) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM, Random Access Memory) 1203. In RAM1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O, Input / Output) interface 1205 is also connected to bus 1204.

[0131] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker, etc.; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet.

[0132] In particular, the methods and / or embodiments in the embodiments of the present application may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above functions defined in the method of the present application are executed.

[0133] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application described above.

[0134] Specifically, the present embodiment may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device, or device.

[0135] Computer readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0136] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0137] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0138] The flow chart or block diagram in the accompanying drawings shows the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0139] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0140] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0141] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0143] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0145] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

Claims

1. A method for generating an image, wherein: The method comprises: Segmenting the target image to obtain an object image corresponding to the target object, wherein the target object is contained in the target image; Acquire multiple background images corresponding to the target object; The object image and the multiple background images are fused respectively to generate multiple fused images containing the target object.

2. The method according to claim 1, wherein: The method further comprises: Cleaning the multiple fused images to select fused images that meet the sample standard as sample images; Input multiple sample images into the model for training.

3. The method according to claim 1 or 2, wherein: The segmenting of the target image to obtain an object image corresponding to the target object further comprises: The object image is rotated to obtain object images of the target object corresponding to multiple angles.

4. The method according to claim 1 or 2, wherein: The fusing the object image with the multiple background images to generate multiple fused images containing the target object comprises: The object image and the multiple background images are fused separately using a Poisson fusion algorithm to generate multiple fused images containing the target object.

5. The method according to claim 2, wherein: The cleaning process of the plurality of fused images to select the fused images meeting the sample standard as the sample images includes: extracting the same area from the target image and the fused image respectively, the area being located at the junction of the object image and the background; The average pixel values ​​of the two regions are obtained, and if the difference between the average pixel values ​​of the two regions is less than a preset threshold, the corresponding fused image is discarded.

6. The method according to claim 2, wherein: The cleaning process of the plurality of fused images to select fused images meeting the sample standard as sample images includes: Dilation processing is performed on the target image and the target image in the fused image respectively to obtain the dilation zone area; The average pixel values ​​of the expansion band areas in the target image and the fused image are obtained respectively. If the difference between the average pixel values ​​of the two expansion band areas is greater than a preset threshold, the corresponding fused image is discarded.

7. The method according to claim 1 or 2, wherein: The fusing the object image with the multiple background images comprises: Performing adjustment processing on the object image or the multiple background images, and fusing the adjusted object image with the multiple background images respectively; The adjustment process includes: Adjust the position of the object image in the background image; In the fusion area of ​​the object image and the background image, set and adjust the weight of the target object and the weight of the background image; Adjust the rotation angle of the object image.

8. The method according to claim 1 or 2, wherein: The step of performing segmentation processing on a plurality of target images to obtain an object graph corresponding to one or more target objects contained in the target images comprises: Perform edge detection on the target object in the target image; According to the edge detection result, the object image corresponding to the target object is segmented along the edge contour.

9. The method according to claim 1 or 2, wherein: The obtaining of multiple background images corresponding to the target object comprises: By performing scene analysis on the target image, a background scene matching the target object is determined; A background image corresponding to the background scene is obtained.

10. A device for generating an image, wherein: The device comprises: A device for performing segmentation processing on a target image to obtain an object image corresponding to a target object, wherein the target object is contained in the target image; Means for acquiring a plurality of background images corresponding to a target object; A device for fusing the object image with the multiple background images respectively to generate multiple fused images containing the target object.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

12. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.