An image augmentation method based on background semantic generation and orientation component replacement
Through an image amplification method based on background semantic generation and orientation component replacement, the problem of difficult to balance category fidelity and sample diversity in the prior art is solved, efficient data set diversification enhancement is achieved, and the generalization ability and accuracy of the model are improved.
Patent Information
- Application Number
- CN202510210936.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing image mixing methods are difficult to enhance sample diversity while maintaining high class fidelity, resulting in insufficient model overfitting and generalization capabilities.
The image amplification method based on background semantic generation and orientation component replacement is adopted to collect background information through visual language models, diffusion models generate new samples with semantic control, and ensure the category fidelity and diversity of the new samples through component segmentation and replacement techniques.
This achieves a significant increase in sample diversity without sacrificing category fidelity, thereby improving the generalization ability, accuracy and robustness of deep learning models.
Smart Images

Figure CN119693725B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual data set enhancement, and in particular to an image augmentation method based on background semantic generation and orientation component replacement. Background Art
[0002] With the development of technology, deep neural networks have achieved improvements and success in various tasks. Compared with the rapid growth of deep neural network model size, the expansion of data scale has lagged significantly. Existing fine-grained visual classification datasets are usually small and lack diversity, which may lead to model overfitting. Data augmentation techniques are essential to alleviate overfitting by increasing the diversity or size of the dataset.
[0003] At present, methods based on image mixing are particularly popular due to their simplicity and effectiveness, and are mainly divided into two categories: label mixing and label preservation. Label mixing technology generates new samples by combining images and their labels, but this method may lead to label ambiguity, reduced category fidelity, and neglect of salient areas, resulting in loss of contextual information. Some studies have tried to solve the above problems by focusing on salient areas, but these methods often sacrifice the integrity of the image, and still cannot guarantee category fidelity due to inaccurate positioning and lack of contextual semantics. Another type of label mixing method is the label preservation enhancement technique, which maintains category fidelity to a certain extent, but lacks sufficient sample diversity. There are also studies that try to introduce diffusion models to increase diversity, but the generated samples are limited to style changes and cannot provide sufficient semantic or structural changes in diversity.
[0004] Therefore, the present invention proposes an image augmentation method based on background semantic generation and orientation component replacement, which can not only enhance sample diversity but also maintain high category fidelity to achieve diversified enhancement of data sets, thereby improving the generalization ability, accuracy and robustness of deep learning models. Summary of the invention
[0005] In view of this, the present invention provides an image augmentation method based on background semantic generation and orientation component replacement, so as to solve the technical problem that the balance between category fidelity and sample diversity cannot be achieved in the current image mixing method.
[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides an image augmentation method based on background semantic generation and orientation component replacement, comprising:
[0008] Use the visual language model to collect the background information of each original sample in the original image dataset and add a direction label to each original sample;
[0009] Generate multiple new samples by using a diffusion model, in which the foreground of the original sample is kept unchanged and the background is controlled by the semantics of the background information;
[0010] The foreground of the original sample and the new sample is segmented based on the component segmentation technology, and the foreground components in the new sample are replaced with corresponding foreground components in the original sample of the same category according to the direction label to obtain an augmented sample data set.
[0011] Furthermore, the visual language model is used to collect the background information of each original sample in the original image dataset, including:
[0012] Use the visual language model to traverse each original sample in the original image dataset to obtain the background semantic label of each original sample;
[0013] The background semantic labels of all original samples are stored to obtain a background semantic label dataset.
[0014] Furthermore, a plurality of new samples are generated by using a diffusion model, in which the foreground of the original sample is kept unchanged and the background is controlled by the semantics of the background information, including:
[0015] For the foreground of each original sample in the original image dataset, randomly select any background semantic label from the background semantic label dataset;
[0016] Establishing a relational connection between the background semantic label and the category to which the foreground belongs, and obtaining an image generation instruction;
[0017] Obtaining new samples with controllable background semantics according to the image generation instructions through a diffusion model;
[0018] Traversing the original image data set, generating new samples corresponding to the foreground of each original sample, and obtaining a new sample data set;
[0019] The foreground of each new sample has the same direction label as the foreground of the corresponding original sample.
[0020] Furthermore, a direction label is added to each original sample, including:
[0021] The direction labels include left, right and forward, which are used to indicate the direction of the original sample foreground;
[0022] A direction dictionary is used to save label pairs consisting of the original sample and its direction label.
[0023] Furthermore, the original sample and the new sample are segmented into components based on the component segmentation technology, including:
[0024] Using a semantic segmentation model to segment the foreground of the original sample at a component level to obtain multiple components of the foreground;
[0025] The bounding box of each component is determined to obtain a set of bounding boxes of the foreground components.
[0026] Furthermore, according to the direction label, the foreground component in the new sample is replaced with the corresponding foreground component in the original sample of the same category to obtain an augmented sample data set, including:
[0027] Select original samples from the original image dataset that match the foreground category of the new sample and have the same direction label;
[0028] The components of the foreground of the original sample are adjusted to match the shape of the to-be-replaced part of the foreground of the new sample, and the to-be-replaced parts of the foreground of the new sample are replaced with the adjusted foreground components of the original sample, thereby updating the new sample;
[0029] Each component in the foreground of the new sample is traversed, a replacement operation is performed on each component, and the new sample is iterated to obtain an augmented sample.
[0030] Further, the components of the original sample foreground are adjusted to match the shape of the to-be-replaced part of the new sample foreground, including:
[0031] According to the bounding box of the foreground part of the new sample to be replaced, the bounding box of the corresponding foreground part of the original sample is affine transformed or key point matched to make the outline and shape of the original sample and the new sample parts consistent.
[0032] In a second aspect, the present invention further provides an image augmentation system based on background semantic generation and orientation component replacement, comprising:
[0033] The semantic analysis module is used to collect the background information of each original sample in the original image dataset using the visual language model and add a direction label to each original sample;
[0034] An amplification module, used for generating a plurality of new samples by using a diffusion model, in which the foreground image of the original sample is kept unchanged and the background is controlled by the semantics of the background information;
[0035] The component replacement module is used to perform component segmentation on the original sample and the new sample based on the component segmentation technology, and replace the components in the new sample with corresponding components of the original sample of the same category according to the direction label to obtain an augmented sample data set.
[0036] In a third aspect, the present invention further provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the image augmentation method based on background semantic generation and orientation component replacement as described in any of the above technical solutions is implemented.
[0037] In a fourth aspect, the present invention further provides a computer-readable storage medium, and when the computer program is executed by a processor, it implements the image augmentation method based on background semantic generation and orientation component replacement as described in any of the above technical solutions.
[0038] Compared with the prior art, the present invention provides the following advantages:
[0039] 1. This method uses a diffusion model to achieve semantically controlled background generation while keeping the foreground unchanged, ensuring high category fidelity while bringing a more diverse sample distribution.
[0040] 2. The directional semantics of foreground objects in the image dataset are expanded through the visual language model, and more accurate and reasonable candidate objects are screened out for replacement at the foreground component level, making the enhanced data image target structure more reasonable.
[0041] 3. Combining directional semantics with component-level segmentation breaks through the limitations of saliency detection, accurately replaces category target components of the same direction, and ensures that the structure of the image after mixing is reasonable.
[0042] In summary, the diffusion model generates new samples with unchanged foreground and semantically controlled background. Through the technical means of direction labeling and component-level segmentation, refined component-level replacement and more reasonable and accurate structural layout are achieved, which diversifies and enhances the dataset while maintaining label consistency, improving the generalization ability, accuracy and robustness of the model. This method can automatically generate enhanced samples that meet semantic requirements and have high quality without a lot of manual intervention, which helps to improve the performance of visual models. After extensive experiments, including fine-grained classification, transfer learning, robustness testing and weakly supervised object localization (WSOL), it is proved that our proposed method is superior to the existing state-of-the-art image enhancement technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of the process of the image augmentation method based on background semantic generation and orientation component replacement provided by the present invention;
[0044] Figure 2 A schematic diagram of the generation result of the diffusion module provided by the present invention;
[0045] Figure 3 A display diagram of the direction label provided by the present invention;
[0046] Figure 4 A schematic diagram of component segmentation provided by the present invention;
[0047] Figure 5 A schematic diagram of a component replacement process provided by the present invention;
[0048] Figure 6 A graphical flow chart of the present invention;
[0049] FIG7( a) is a schematic diagram of t-SNE feature visualization provided by the present invention;
[0050] FIG7( b ) is a schematic diagram showing changes in the blind area in the feature space provided by the present invention;
[0051] Figure 8 A schematic diagram showing the comparison between the method of the present invention and other image mixing algorithms;
[0052] Fig. 9 A schematic diagram of the structure of an image augmentation system based on background semantic generation and orientation component replacement provided by the present invention;
[0053] Fig.10 This is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0054] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0055] Before describing the embodiments of the present invention, the inventive concept of the present invention is first introduced.
[0056] Image mixing-based methods are popular because of their simplicity and effectiveness in alleviating model overfitting problems. They are mainly divided into label mixing methods and label preservation methods. Among them:
[0057] Label mixing methods such as Mixup generate composite images by linear interpolation of randomly selected image pairs. CutMix pastes random patches of one image onto another. Manifold Mixup introduces randomness to enhance representation during training by interpolating the hidden state of the network, randomly selecting sample pairs, hidden layers, and interpolation coefficients. However, the above-mentioned random mixing label mixing methods may lose salient areas in the image. Therefore, label mixing techniques based on salient areas have been proposed. For example, SaliencyMix uses saliency maps to focus on the most important areas in the image to ensure the integrity of the overall image. PuzzleMix considers both salient areas and local statistical information during image mixing.
[0058] Due to the potential for label ambiguity and reduced class fidelity when images of different categories are mixed, label-preserving augmentation strategies are gaining attention. Label-preserving data augmentation methods, such as AugMix, apply a random combination of color and contrast transformations to images for data augmentation. DiffuseMix combines the original image with a stylized image and overlays a fractal image for data augmentation. InPS sets a threshold based on the saliency spectrum to achieve intra-class region exchange. These methods ensure high class fidelity to a certain extent, but they fall short in terms of sample diversity.
[0059] At present, diffusion models can generate accurate images according to instructions and have been widely used in image recognition tasks. In order to achieve high category fidelity, the present invention introduces an invariant foreground diffusion model to generate new samples based on background semantics to enhance sample diversity. On the other hand, through component-level semantic segmentation, the various components of the image target are accurately located, and the foreground direction semantics are combined with the target components to achieve controllable directional replacement of target components within the category, which exceeds the limitations of traditional saliency detection methods and can ensure that the structure of the mixed image is reasonable, the regional features are diverse, and the component semantics are accurate.
[0060] See also Figure 1 This embodiment provides an image augmentation method based on background semantic generation and orientation component replacement, including:
[0061] Step S101: using a visual language model to collect background information of each original sample in the original image dataset, and adding a direction label to each original sample;
[0062] Step S102: generating a plurality of new samples by using a diffusion model, in which the foreground of the original sample is kept unchanged and the background is controlled by the semantics of the background information;
[0063] Step S103: performing component segmentation on the foreground of the original sample and the new sample based on the component segmentation technology, and replacing the foreground components in the new sample with corresponding foreground components in the original sample of the same category according to the direction label to obtain an augmented sample data set.
[0064] The method of this embodiment, first, uses a diffusion model to generate new samples based on background semantic information, which can create diverse background changes while keeping the original foreground content unchanged, effectively expanding the diversity of the training data set, thereby improving the generalization ability of the model. Diverse backgrounds help the model to better learn the relationship between the foreground and the background, thereby improving the ability to recognize targets in different background environments. Secondly, adding a direction label to the foreground of each original sample can provide semantic guidance for the processing of the foreground, help replace components in subsequent processing, ensure that the generated enhanced samples can still meet the orientation and posture of the actual object, and further improve the quality of the enhanced samples. Finally, the component segmentation and replacement method ensures that these components can be identified and separated under different backgrounds during the training process, thereby more accurately capturing the detailed features of the foreground and improving the recognition accuracy. The effect of data enhancement is achieved by flexibly replacing components.
[0065] As a preferred embodiment, in step S101, the background information of each original sample in the original image dataset is collected by using a visual language model, including:
[0066] Use the visual language model to traverse each original sample in the original image dataset to obtain the background semantic label of each original sample;
[0067] The background semantic labels of all original samples are stored to obtain a background semantic label dataset.
[0068] Specifically, for example, the original image dataset of birds may contain backgrounds such as grass, sky, ocean, etc. By collecting these background keywords, they can be input into the diffusion model to generate a replacement background for the original image.
[0069] As a preferred embodiment, in step S102, a plurality of new samples are generated by using a diffusion model, in which the foreground of the original sample is kept unchanged and the background is controlled by the semantics of the background information, including:
[0070] For the foreground of each original sample in the original image dataset, randomly select any background semantic label from the background semantic label dataset;
[0071] Establishing a relational connection between the background semantic label and the category to which the foreground belongs, and obtaining an image generation instruction;
[0072] Obtaining new samples with controllable background semantics according to the image generation instructions through a diffusion model;
[0073] Traversing the original image data set, generating new samples corresponding to the foreground of each original sample, and obtaining a new sample data set;
[0074] The foreground of each new sample has the same direction label as the foreground of the corresponding original sample.
[0075] like Figure 2 As shown, Figure 2 The visualization results of multi-scene background semantic diffusion are shown. Figure 2 In the figure, the original sample is the bird on the branch in the upper left corner of the picture, and the foreground is the bird in the lower left corner. The diffusion model generates four new samples with the foreground unchanged and the background replaced according to the background semantic labels of grass or snow.
[0076] In some embodiments, in order to generate samples with a wider diversity, background lighting conditions, weather conditions, time periods (such as daytime, dusk, night), etc. can also be controlled to generate samples in different environments. In addition, the artistic style of the background can be controlled, for example, by adding artistic effects (such as oil painting style, watercolor style, etc.), or using style transfer technology to transfer the styles of different backgrounds, to further improve the diversity of samples.
[0077] As a preferred embodiment, adding a direction label to each original sample includes:
[0078] The direction labels include left, right and forward, which are used to indicate the direction of the original sample foreground;
[0079] A direction dictionary is used to save label pairs consisting of the original sample and its direction label.
[0080] It should be noted that the diffusion model is used to fuse the foreground and background to generate new sample images. The generated new samples are image samples with unchanged foreground and changed background scenes. The direction label is used to mark the foreground orientation. The direction label can be used to select replacement parts with the same orientation in the subsequent processing process. For example: if the bird in the picture of the part to be replaced is facing left, then select another bird facing left in the original image data set for component replacement, so that the replaced structure can be more reasonable.
[0081] The process of generating a semantic control background is described below through a specific embodiment.
[0082] Step 1: Use the large model inference framework vLLM to traverse each original sample image I in the original image dataset D i , get image I i Background semantic label TB i , the background semantic labels TB of all images in the dataset D i Stored in the set BG, the elements in the set BG are recorded as BG j .
[0083] Step 2: traverse each image I in the original dataset againi , randomly select BG from the set BG j With image I i The prospect belongs to the category FC i Concatenate to get the image generation instructions for the diffusion model: , where ○ represents a connection relationship.
[0084] Step 3: Use the diffusion model G(·) to generate the image I according to the diffusion instruction Pg i Corresponding sample I i ', and finally generate a dataset D' with controllable background semantics corresponding to the original dataset D.
[0085] To ensure that the overall structure is more reasonable after the parts are replaced, vLLM is used to mark the direction of the objects in the image. In the specific processing process, a list O = [O1, O2, O3] is defined to represent the direction of the image foreground, where O1, O2 and O3 represent left, right and front, respectively. Use vLLM to traverse each real image I in the original dataset i To get its direction label Finally, the direction dictionary IO is used to save the label pairs of the original image and its foreground direction {I i :O k}.like Figure 3 As shown, Figure 3 Three different orientation labels are shown: left, center, and right, with the vehicle as the foreground.
[0086] As a preferred embodiment, in step S103, component segmentation is performed on the original sample and the new sample based on the component segmentation technology, including:
[0087] Using a semantic segmentation model to segment the foreground of the original sample at a component level to obtain multiple components of the foreground;
[0088] The bounding box of each component is determined to obtain a set of bounding boxes of the foreground components.
[0089] like Figure 4 As shown, Figure 4 A schematic diagram of part-level segmentation is shown. Figure 4 The foreground (vehicle) in the figure is divided into front doors, rear doors, wheels, windshield and other parts.
[0090] As a preferred embodiment, replacing the foreground component in the new sample with the corresponding foreground component in the original sample of the same category according to the direction label to obtain an augmented sample data set includes:
[0091] Select original samples from the original image dataset that match the foreground category of the new sample and have the same direction label;
[0092] The components of the foreground of the original sample are adjusted to match the shape of the to-be-replaced part of the foreground of the new sample, and the to-be-replaced parts of the foreground of the new sample are replaced with the adjusted foreground components of the original sample, thereby updating the new sample;
[0093] Each component in the foreground of the new sample is traversed, and the new sample is iterated to obtain an augmented sample.
[0094] The process of component replacement is described below using a specific embodiment.
[0095] First, the semantic segmentation model S(·) is used to segment each image I in the original image dataset D. i Segment the foreground at the component level. Assume I i The foreground contains n parts, P = [P1, P2, ....., Pn], and the bounding box BBox corresponding to each Pj will be saved to B i In the format: B i = [BBox i1 ,......, BBox in For example, for an original image sample I with a bird as the foreground i , component P j It can represent the head, wings or tail of a bird, etc. j The corresponding BBox is the coordinate of the bounding box of the bird's head. The content stored in Bi is the list of bounding boxes of each component corresponding to the image Ii.
[0096] Second, use function A:I i = M(I i '), find the new sample I i ' The corresponding original sample I i . The image will be enlarged Initialized to I i '. Traverse Each component in , and extract P from the bounding box set B j BBox ij .
[0097] Since the diffusion model used in the present invention has the characteristics of unchanged foreground object position and structure, BBox ij Also applies to I i 'Part P in j . Find the image I i The corresponding direction label list O k , and use function B:I m = C(I i',D, FC i ,O k ), randomly select a sample from the original dataset D that matches the generated sample I i 'Same category, FC i and direction O k The real image I m ;
[0098] Find image I in set B m Part P j BBox mj Using the function Adjust the real sample I m area2 to match the resulting image The shape of the area to be replaced in area1 is replaced, and area1 is replaced to iteratively update the result image . The real image I m BBox in mj Replace with BBox in ij , iteratively update the result image .
[0099] Please refer to Figure 5 , Figure 5 The entire process of the above segmentation and replacement is shown accordingly.
[0100] It should be noted that the illustrations of the present invention are based on examples of specific data sets, but the application of this method is centered on component structure and is applicable to all objects with component structures. It achieves data enhancement effects by flexibly replacing components and is not limited to these data sets. It has wide applicability and extensibility.
[0101] In some embodiments, adjusting the component of the original sample foreground to match the shape of the to-be-replaced portion of the new sample foreground includes:
[0102] According to the bounding box of the foreground part of the new sample to be replaced, the bounding box of the corresponding foreground part of the original sample is affine transformed or key point matched to make the outline and shape of the original sample and the new sample parts consistent.
[0103] Specifically, affine transformations include translation, scaling, rotation, and shearing, which can be used independently or in combination. In addition to affine transformations, adversarial training methods such as generative adversarial networks (GANs) can also be used to optimize the effect of component replacement. By introducing a discriminant network to determine whether the replaced components conform to natural visual laws, the quality of replacement can be improved. In some embodiments, the components to be replaced can also be intelligently selected based on the image content and the requirements of the generation task (such as visual consistency, semantic matching, style matching, etc.).
[0104] like Figure 6 As shown, Figure 6 The overall process of this embodiment is presented in a graphical flow chart.
[0105] In order to verify the actual use effect of the present invention, the effect of the method is demonstrated below using actual experimental data. For the sake of simplicity, PartMix is used to refer to the method of the present invention.
[0106] Assume that the random variable X b ~D b represents the set of feature vectors in the original data set, where D b is the distribution of feature vectors in the feature space. The sample points of the original data set are represented as , probability density function P b (x) reflects the distribution density of these sample points. The total coverage of the feature vectors in the original data set can be expressed as:
[0107]
[0108] This represents the area occupied by these sample points in the feature space.
[0109] Likewise, X g ~D g represents the set of feature vectors from the PartMix dataset, D g Represents the corresponding feature distribution. The sample points of the enhanced dataset are represented as , probability density function P g (x) describes their distribution density. The total coverage of the PartMix dataset is:
[0110]
[0111] It reflects the expanded area occupied by the feature vector in the enhanced dataset.
[0112] As shown in Figure 7(a), the pink area shows the PartMix sample X g The distribution after dimensionality reduction in the feature space, while the blue area represents the original data set X b The relationship between their total coverage is:
[0113]
[0114] The above analysis shows that the samples output by PartMix have a wider coverage and denser distribution in the feature space. However, due to the non-uniform probability density distribution in the original dataset, X b ~D bSome areas in the image show insufficient feature coverage, forming "blind areas", as shown in Figure 7(b) (areas 1 and 2). The definition of blind areas is:
[0115] ,
[0116] PartMix significantly reduces the blind areas in the feature space through background semantic generation and controllable directional part replacement. For example, Figure 7(b) shows two typical blind areas in the original dataset: Area 1 is due to insufficient sample coverage density. After introducing the new samples generated by PartMix, the coverage density of this area has increased significantly (satisfying ), successfully transforming the blind area into a covered area. Similarly, area 2, which was not covered in the original dataset, was effectively covered by the enhancement strategy, significantly reducing the overall scope of the blind area. This strategy of expanding the coverage of the feature space enables the model to learn a wider distribution of features, greatly enhancing its generalization ability and robustness in diverse scenarios.
[0117] In order to more intuitively illustrate the effect of the present invention, Figure 8 As shown, Figure 8 The figure shows the comparison between PartMix and other image mixing algorithms. In the figure, F stands for fidelity and D stands for diversity. The left side of the figure shows the comparative visualization of existing label mixing and label-preserving enhancement methods, and the right side shows the effect of the PartMix method. While ensuring high fidelity of categories, PartMix greatly enriches the diversity of samples, thereby improving the generalization ability of the model.
[0118] This embodiment also provides a system, such as Fig. 9 As shown, the image augmentation system 900 based on background semantic generation and orientation component replacement includes:
[0119] Semantic analysis module 901, used to collect background information of each original sample in the original image data set by using a visual language model, and add a direction label to each original sample;
[0120] An amplification module 902 is used to generate a plurality of new samples by using a diffusion model, in which the foreground image of the original sample is kept unchanged and the background is controlled by the semantics of the background information;
[0121] The component replacement module 903 is used to perform component segmentation on the original sample and the new sample based on the component segmentation technology, and replace the components in the new sample with corresponding components of the original sample of the same category according to the direction label to obtain an augmented sample data set.
[0122] like Fig.10As shown, the above-mentioned image augmentation method based on background semantic generation and orientation component replacement, the present invention also provides an electronic device 1000 accordingly, the electronic device can be a computing device such as a mobile terminal, a desktop computer, a notebook, a palm computer and a server. The electronic device includes a processor 1001, a memory 1002 and a display 1003.
[0123] In some embodiments, the memory 1002 may be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the memory 1002 may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory 1002 may also include both an internal storage unit of the computer device and an external storage device. The memory 1002 is used to store application software and various data installed on the computer device, such as program codes installed on the computer device. The memory 1002 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a program 1004 of an image augmentation method based on background semantic generation and orientation component replacement is stored on the memory 1002, and the program 1004 of an image augmentation method based on background semantic generation and orientation component replacement can be executed by the processor 1001, thereby realizing an image augmentation method based on background semantic generation and orientation component replacement in each embodiment of the present invention.
[0124] In some embodiments, the processor 1001 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 1002, such as executing an image augmentation method program based on background semantic generation and orientation component replacement.
[0125] In some embodiments, the display 1003 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 1003 is used to display information on the computer device and to display a visual user interface. The components 1001-1003 of the computer device communicate with each other through a system bus.
[0126] This embodiment also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the image augmentation method based on background semantic generation and orientation component replacement as described in any of the above technical solutions is implemented.
[0127] The computer-readable storage medium and computing device provided according to the above-mentioned embodiments of the present invention can be implemented with reference to the specific description of the image augmentation method based on background semantic generation and orientation component replacement as described above according to the present invention, and have similar beneficial effects as the image augmentation method based on background semantic generation and orientation component replacement as described above, which will not be repeated here.
[0128] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. An image augmentation method based on background semantic generation and orientation component replacement, characterized in that: include: Use the visual language model to collect the background information of each original sample in the original image dataset and add a direction label to each original sample; Generate multiple new samples by using a diffusion model, in which the foreground of the original sample is kept unchanged and the background is controlled by the semantics of the background information; The foreground of the original sample and the new sample is segmented into components based on the component segmentation technology, including: using a semantic segmentation model to segment the foreground of the original sample at the component level to obtain multiple components of the foreground; determining the bounding box of each component to obtain a set of bounding boxes of each foreground component; Replacing the foreground component in the new sample with the corresponding foreground component in the original sample of the same category according to the direction label to obtain an augmented sample data set includes: An original sample matching the foreground category of the new sample and having the same direction label is selected from the original image data set; the component of the foreground of the original sample is adjusted to match the shape of the to-be-replaced part of the foreground of the new sample, and the to-be-replaced component of the foreground of the new sample is replaced with the adjusted foreground component of the original sample to update the new sample; each component in the foreground of the new sample is traversed, a replacement operation is performed on each component, and the new sample is iteratively updated to obtain an augmented sample.
2. The image augmentation method based on background semantic generation and orientation component replacement according to claim 1, characterized in that: The visual language model is used to collect background information of each original sample in the original image dataset, including: Use the visual language model to traverse each original sample in the original image dataset to obtain the background semantic label of each original sample; The background semantic labels of all original samples are stored to obtain a background semantic label dataset.
3. The image augmentation method based on background semantic generation and orientation component replacement according to claim 2, characterized in that: Generating a plurality of new samples by using a diffusion model, in which the foreground of the original sample is kept unchanged and the background is controlled by the semantics of the background information, includes: For the foreground of each original sample in the original image dataset, randomly select any background semantic label from the background semantic label dataset; Establishing a relational connection between the background semantic label and the category to which the foreground belongs, and obtaining an image generation instruction; Obtaining new samples with controllable background semantics according to the image generation instructions through a diffusion model; Traversing the original image data set, generating new samples corresponding to the foreground of each original sample, and obtaining a new sample data set; The foreground of each new sample has the same direction label as the foreground of the corresponding original sample.
4. The image augmentation method based on background semantic generation and orientation component replacement according to claim 1, characterized in that: Add direction labels to each original sample, including: The direction labels include left, right and forward, which are used to indicate the direction of the original sample foreground; A direction dictionary is used to save label pairs consisting of the original sample and its direction label.
5. The image augmentation method based on background semantic generation and orientation component replacement according to claim 1, characterized in that: The components of the original sample foreground are adjusted to match the shape of the to-be-replaced parts of the new sample foreground, including: According to the bounding box of the foreground part of the new sample to be replaced, the bounding box of the corresponding foreground part of the original sample is affine transformed or key point matched to make the outline and shape of the original sample and the new sample parts consistent.
6. An image augmentation system based on background semantic generation and orientation component replacement, characterized in that: include: The semantic analysis module is used to collect the background information of each original sample in the original image dataset using the visual language model and add a direction label to each original sample; An amplification module, used for generating a plurality of new samples by using a diffusion model, in which the foreground image of the original sample is kept unchanged and the background is controlled by the semantics of the background information; The component replacement module is used to perform component segmentation on the original sample and the new sample based on the component segmentation technology, including: using a semantic segmentation model to perform component-level segmentation on the foreground of the original sample to obtain multiple components of the foreground; determining the bounding box of each component to obtain a set of bounding boxes of the foreground components; replacing the components in the new sample with corresponding components of the original sample of the same category according to the direction label to obtain an augmented sample data set, including: selecting an original sample that matches the foreground category of the new sample and has the same direction label from the original image data set; adjusting the components of the foreground of the original sample to match the shape of the to-be-replaced part of the foreground of the new sample, and replacing the to-be-replaced components of the foreground of the new sample with the adjusted foreground components of the original sample to update the new sample; traversing each component in the foreground of the new sample, performing a replacement operation on each component, iteratively updating the new sample, and obtaining an augmented sample.
7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the image augmentation method based on background semantic generation and orientation component replacement as described in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that: When the computer program in the computer-readable storage medium is executed by the processor, the image augmentation method based on background semantic generation and orientation component replacement as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Training sample image enhancement method and system based on neural network
CN114241256A
Children image synthesis method based on streetscape
CN114913274A