Method, apparatus, storage medium and electronic device for generating image data with annotations

By combining the background image generation model and the foreground object generation model, border coordinate segmentation and fusion technology are used to generate efficient and high-quality annotated image data, solving the problem of low manual annotation efficiency.

CN114723646BActive Publication Date: 2025-07-22BEIJING YUDA ORIENTAL SOFTWARE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210179838.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-07-22
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The method of manually labeling image data is inefficient and it is difficult to efficiently generate high-quality model training data.

Method used

By inputting the random matrix into the training background image to generate a model, obtaining the border coordinates and segmenting the background image, combining the foreground object generation model to generate the target sub-image, and fusing the border coordinates as label information to generate annotated image data.

Benefits of technology

The efficiency of image data annotation is improved, and the generated image data is of higher quality, avoiding the inefficiency problem of manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114723646B_ABST
    Figure CN114723646B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, storage medium, and electronic device for generating annotated image data to solve problems existing in related technologies. The method for generating annotated image data includes: inputting a first random matrix into a trained background image generation model to obtain a background image and border coordinates output by the background image generation model, where the border coordinates are used to represent the position of a first sub-image of a foreground object to be generated in the background image; segmenting the background image according to the border coordinates to obtain a first sub-image and a second sub-image; inputting a second random matrix and the first sub-image into a trained foreground object generation model to obtain a target sub-image output by the foreground object generation model, where the target sub-image includes a target foreground object; using at least the border coordinates as annotation information for the target foreground object in the target sub-image, and fusing the annotated target sub-image with the second sub-image to obtain annotated image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image generation, and more particularly, to a method, apparatus, storage medium, and electronic device for generating annotated image data. Background Art

[0002] In the field of artificial intelligence, to train a model with relatively high quality, a large amount of training data is often required. In related technologies, to ensure the accuracy of training data, model training data is often obtained by manually annotating data. However, the method of manually annotating data is inefficient. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a method, apparatus, storage medium, and electronic device for generating annotated image data to solve the problems existing in related technologies.

[0004] To achieve the above purpose, according to the first aspect of the embodiments of the present disclosure, a method for generating annotated image data is provided. The method includes:

[0005] Input a first random matrix into a trained background image generation model to obtain a background image and a border coordinate output by the background image generation model. The border coordinate is used to represent the position of a first sub-image of a foreground object to be generated in the background image;

[0006] Segment the background image according to the border coordinate to obtain the first sub-image and a second sub-image;

[0007] Input a second random matrix and the first sub-image into a trained foreground object generation model to obtain a target sub-image output by the foreground object generation model. The size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes a target foreground object;

[0008] Use at least the border coordinate as annotation information of the target foreground object in the target sub-image, and fuse the annotated target sub-image with the second sub-image to obtain annotated image data.

[0009] Optionally, the training process of the background image generation model includes:

[0010] Input a random matrix sample into a background image generation model to be trained to obtain a synthetic background image output by the background image generation model to be trained;

[0011] Input the synthetic background image into a foreground recognizer to obtain a first discrimination result output by the foreground recognizer on whether there is a target foreground object in the synthetic background image;

[0012] Input the random matrix sample, the synthesized background image, and the background image sample set into the discriminant model to be trained, so as to obtain a second discrimination result on whether the discriminant model to be trained discriminates the synthesized background image as an image in the background image sample set, and a third discrimination result on whether the discriminant model to be trained discriminates each background image sample in the background image sample set as an image generated by the background image generation model to be trained;

[0013] Calculate loss information according to the first discrimination result, the second discrimination result, and the third discrimination result;

[0014] Adjust the training parameters of the discriminant model to be trained according to the loss information, obtain the updated discriminant model to be trained, and return to execute the step of inputting the random matrix sample, the synthesized background image, and the background image sample set into the discriminant model to be trained to obtain the second discrimination result and the third discrimination result until the training parameters of the discriminant model to be trained are updated N times.

[0015] Optionally, the method further includes:

[0016] After the training parameters of the discriminant model to be trained are updated for the Nth time, update the training parameters of the background image generation model according to the loss information calculated for the (N + 1)th time.

[0017] Optionally, calculating the loss information according to the first discrimination result, the second discrimination result, and the third discrimination result includes:

[0018] Calculate the loss information through the following formula:

[0019]

[0020] where J represents the loss information, m represents the total number of sample samplings, x (i) represents the i-th background image sample, z (i) represents the i-th random matrix sample, G(z (i) ) represents the i-th synthesized background image, D(x (i) ) represents the probability of discriminating the i-th background image sample as an image generated by the background image generation model to be trained, T(G(z (i) )) represents the probability that the foreground object recognizer recognizes a target foreground object in the i-th synthesized background image, and D(G(z (i) )) represents the probability of discriminating the i-th synthesized background image as an image in the background image sample set.

[0021] Optionally, the foreground object generation model includes a network fusion module. The step of inputting the second random matrix and the first sub-image into the trained foreground object generation model to obtain the target sub-image output by the foreground object generation model includes:

[0022] Input the second random matrix and the first sub-image into the network fusion module to obtain a fusion matrix after fusing the image matrices corresponding to the second random matrix and the first sub-image;

[0023] Generate an initial target sub-image based on the fusion matrix, where the size of the initial target sub-image is the same as the size of the first sub-image;

[0024] Fuse the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image to obtain the initial target sub-image after edge fusion, and the target sub-image represents the initial target sub-image after edge fusion.

[0025] Optionally, the step of fusing the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image includes:

[0026] For each first pixel point in the first edge region of the initial target image, calculate the pixel value of the third pixel point at the same position as the first pixel point in the target sub-image according to the pixel value of the first pixel point and the pixel value of the second pixel point at the same position in the second edge region of the first sub-image.

[0027] Optionally, the method further includes:

[0028] Obtain the object type information of the target foreground object;

[0029] The step of using at least the border coordinates as the annotation information of the target foreground object in the target sub-image includes:

[0030] Use the border coordinates and the object type information as the annotation information of the target foreground object in the target sub-image.

[0031] According to the second aspect of the embodiments of the present disclosure, there is provided an annotated image data generation device, the device includes:

[0032] A first input module, configured to input a first random matrix into a trained background image generation model to obtain a background image and border coordinates output by the background image generation model, where the border coordinates are used to represent the position of the first sub-image of the foreground object to be generated in the background image;

[0033] A splitting module, configured to split the background image according to the border coordinates to obtain the first sub-image and the second sub-image;

[0034] A second input module, configured to input a second random matrix and the first sub-image into a trained foreground object generation model to obtain a target sub-image output by the foreground object generation model, where the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes a target foreground object;

[0035] A fusion module, configured to use at least the border coordinates as annotation information of the target foreground object in the target sub-image, and fuse the annotated target sub-image with the second sub-image to obtain annotated image data.

[0036] According to a third aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the first aspects above are implemented.

[0037] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0038] A memory, on which a computer program is stored;

[0039] A processor, configured to execute the computer program in the memory to implement the steps of the method described in any one of the first aspects above.

[0040] By adopting the above technical solutions, at least the following technical effects can be achieved:

[0041] By inputting a first random matrix into a trained background image generation model, a background image and border coordinates output by the background image generation model are obtained. Since the border coordinates can be used to represent the position of the first sub-image of the foreground object to be generated in the background image, the background image can be split according to the border coordinates to obtain the first sub-image and the second sub-image. On this basis, by inputting a second random matrix and the first sub-image into a trained foreground object generation model, a target sub-image output by the foreground object generation model is obtained. Wherein, the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes a target foreground object. At least the border coordinates are used as annotation information of the target foreground object in the target sub-image, and the annotated target sub-image is fused with the second sub-image, so that annotated image data can be obtained. Compared with the method of manually annotating data in the related art, the method of generating annotated image data by the model in the present disclosure is more efficient.

[0042] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. Brief Description of the Drawings

[0043] The accompanying drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the accompanying drawings:

[0044] Figure 1 is a flowchart of a method for generating annotated image data according to an exemplary embodiment of the present disclosure.

[0045] Figure 2 is a block diagram of a device for generating annotated image data according to an exemplary embodiment of the present disclosure.

[0046] Figure 3 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Description of the Embodiments

[0047] The following provides a detailed description of the specific embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure, and are not used to limit the present disclosure.

[0048] In the related art, compared with traditional neural network models, a generative adversarial network (GAN, Generative Adversarial Networks) is an unsupervised network architecture that includes two independent networks, namely a discriminative network and a generative network. Among them, the generative network is used to generate data, and the discriminative network is used to determine whether the data generated by the generative network is real data or fake data (i.e., the data generated by the generative network). During the training process of GAN, the training parameters of the discriminative network and / or the generative network are usually adjusted according to the discrimination result of the discriminative network, so that the discriminative network has a strong discrimination ability (i.e., the ability to correctly distinguish real data and fake data), and the generative network has a strong data generation ability (i.e., the generated data can make the discriminative network identify it as real data).

[0049] The following provides a detailed description of the method, device, storage medium, and electronic device for generating annotated image data provided by the embodiments of the present disclosure.

[0050] Figure 1 is a flowchart of a method for generating annotated image data according to an exemplary embodiment of the present disclosure, as Figure 1 shown, the method includes:

[0051] S101. Input the first random matrix into the trained background image generation model to obtain the background image output by the background image generation model and the border coordinates, where the border coordinates are used to represent the position of the first sub-image of the foreground object to be generated in the background image.

[0052] It should be noted that the background image generation model can generate a corresponding background image according to the input random matrix. In the present disclosure, the first random matrix is input into the trained background image generation model to obtain the background image output by the background image generation model and the border coordinates. Among them, the border coordinates are used to represent the position of the first sub-image of the foreground object to be generated in the background image. The border coordinates can be randomly generated by a corresponding algorithm after the background image is generated, or on the premise that the background image generation model is trained to generate the background image and the border coordinates during the training process of the background image generation model, so that the background image generation model can generate the background image and the border coordinates according to the input random matrix.

[0053] S102. Split the background image according to the border coordinates to obtain the first sub-image and the second sub-image.

[0054] It can be understood that the border coordinates can be used to represent the position of the first sub-image of the foreground object to be generated in the background image. Therefore, the background image can be split according to the border coordinates to obtain the first sub-image and the second sub-image.

[0055] For example, if the border coordinates are (1, 2), (1, 4), (2, 2), and (2, 4), these four coordinates can be connected in sequence to form a border. Then, when one side of the background image is determined as the x-axis and the other side intersecting with this side is determined as the y-axis, the position of the border in the background image is determined, so that the background image can be split according to the border to obtain the first sub-image and the second sub-image.

[0056] S103. Input the second random matrix and the first sub-image into the trained foreground object generation model to obtain the target sub-image output by the foreground object generation model.

[0057] Among them, the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes the target foreground object.

[0058] It should be noted that during the training process of the foreground object generation model, the foreground object generation model can be trained to generate a target image with the same size as the input image according to the input random matrix and image, and the target image includes the target foreground object. On this basis, by inputting the second random matrix and the first sub-image into the trained foreground object generation model, the target sub-image output by the foreground object generation model can be obtained. The size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes the target foreground object.

[0059] Among them, if the annotated image data annotates an animal, the target foreground object can be any animal. If the annotated image data annotates a vehicle, the target foreground object can be any vehicle. The present disclosure does not make specific limitations on this.

[0060] S104. At least use the border coordinates as the annotation information of the target foreground object in the target sub-image, and fuse the annotated target sub-image with the second sub-image to obtain the annotated image data.

[0061] It can be understood that the border coordinates can be used to represent the position of the first sub-image of the foreground object to be generated in the background image, and the target sub-image is an image with the same size as the first sub-image obtained by inputting the first sub-image and the second random matrix into the trained foreground object generation model. Therefore, the border coordinates can be used as the annotation information of the target foreground object in the target sub-image.

[0062] In the present disclosure, at least the border coordinates are used as the annotation information of the target foreground object in the target sub-image, and the annotated target sub-image is fused with the second sub-image to obtain the annotated image data. Among them, fusing the annotated target sub-image with the second sub-image may refer to fusing the outer border of the target sub-image with the inner border of the second sub-image. It can be understood that the inner border of the second sub-image is generated after splitting the background image according to the border coordinates, and the inner border can fit with the outer border of the first sub-image.

[0063] Using the above method, by inputting the first random matrix into the trained background image generation model, the background image and the border coordinates output by the background image generation model are obtained. Since the border coordinates can be used to represent the position of the first sub-image of the foreground object to be generated in the background image, the background image can be segmented according to the border coordinates to obtain the first sub-image and the second sub-image. On this basis, by inputting the second random matrix and the first sub-image into the trained foreground object generation model, the target sub-image output by the foreground object generation model is obtained. Wherein, the size of the target sub-image is the same as the size of the first sub-image, and the target sub-image includes the target foreground object. At least the border coordinates are used as the annotation information of the target foreground object in the target sub-image, and the annotated target sub-image is fused with the second sub-image, so that the annotated image data can be obtained. Compared with the method of manually annotating data in the related art, the method of generating annotated image data by the model in the present disclosure is more efficient.

[0064] In order to make it easier for those of ordinary skill in the art to understand the technical solutions provided by the present disclosure, the above steps will be described in detail with examples below.

[0065] Optionally, the training process of the background image generation model may include:

[0066] Input the random matrix sample into the background image generation model to be trained, and obtain the synthetic background image output by the background image generation model to be trained;

[0067] Input the synthetic background image into the foreground recognizer to obtain the first discrimination result output by the foreground recognizer on whether there is a target foreground object in the synthetic background image;

[0068] Input the random matrix sample, the synthetic background image and the background image sample set into the discrimination model to be trained, so as to obtain the second discrimination result on whether the synthetic background image is discriminated as an image in the background image sample set by the discrimination model to be trained, and obtain the third discrimination result on whether each background image sample in the background image sample set is discriminated as an image generated by the background image generation model to be trained by the discrimination model to be trained;

[0069] Calculate the loss information according to the first discrimination result, the second discrimination result and the third discrimination result;

[0070] Adjust the training parameters of the discrimination model to be trained according to the loss information, obtain the updated discrimination model to be trained, and return to execute the step of inputting the random matrix sample, the synthetic background image and the background image sample set into the discrimination model to be trained to obtain the second discrimination result and the third discrimination result until the training parameters of the discrimination model to be trained are updated N times.

[0071] Optionally, the method provided by the embodiments of the present disclosure may further include:

[0072] After updating the training parameters of the discriminative model to be trained for the Nth time, update the training parameters of the background map generation model according to the loss information calculated for the (N + 1)th time.

[0073] It should be noted that the network architecture of the background map generation model in the training process may be similar to the network architecture of GAN. That is to say, in the training process of the background map generation model, the discriminative model and the background map generation model may be involved.

[0074] In the training process of the background map generation model, a random matrix sample may be input into the background map generation model to be trained, and a synthetic background image output by the background map generation model to be trained may be obtained. In order to ensure that the generated synthetic background image does not include the target foreground object and avoid missing the annotation of the target foreground object in the synthetic background image, a foreground recognizer may be added after generating the synthetic background image, and the discrimination result of the foreground recognizer may be incorporated into the calculation of the model loss information, so as to adjust the parameters. Specifically, the synthetic background image is input into the foreground recognizer to obtain a first discrimination result of whether there is a target foreground object in the synthetic background image output by the foreground recognizer. Among them, the foreground recognizer can be used to identify whether there is a target foreground object in the synthetic background image, and the algorithm used by the foreground recognizer can be the YOLO (You Only Look Once) algorithm or the like. The foreground recognizer can identify the target foreground object in the image and output the border coordinates of the target foreground object and the category to which the target foreground object belongs. Therefore, it is possible to determine whether there is a target foreground object in the synthetic background image according to the category to which the target foreground object output by the foreground recognizer belongs. For example, if the image data with annotations annotates a feline animal, then when the category to which the target foreground object output by the foreground recognizer belongs indicates that the target foreground object belongs to the feline animal, it is determined that there is a target foreground object in the synthetic background image.

[0075] Since, during the training process of the GAN, the parameters of the generation network are usually fixed first, and the parameters of the discrimination network are adjusted according to the judgment results of the discrimination network. After enabling the discrimination network to have a certain discrimination ability (i.e., the ability to correctly distinguish real data and the data generated by the generation network) through this training process, the parameters of the discrimination network are fixed, and the parameters of the generation network are adjusted according to the judgment results of the discrimination network, so that the generation network has a certain data generation ability (i.e., the generated data can enable the discrimination network to identify it as real data). Iterating in this way, a generation network with strong data generation ability can be trained. On this basis, with the training parameters of the background map generation model to be trained fixed, the random matrix sample, the synthetic background image, and the background image sample set (this background image sample set can correspond to real data) can be input into the discrimination model to be trained, so as to obtain the second discrimination result of whether the discrimination model to be trained identifies the synthetic background image as an image in the background image sample set, and the third discrimination result of whether the discrimination model to be trained identifies each background image sample in the background image sample set as an image generated by the background map generation model to be trained, and calculate the loss information according to the first discrimination result, the second discrimination result, and the third discrimination result, so that the training parameters of the discrimination model to be trained can be adjusted according to the loss information, and an updated discrimination model to be trained can be obtained.

[0076] It should be noted that, in order to enable the discrimination model to have better discrimination ability, the steps of inputting the random matrix sample into the background map generation model to be trained, obtaining the synthetic background image and the first discrimination result of the foreground recognizer for this synthetic background image, and then inputting the random matrix sample, the synthetic background image, and the background image sample set into the discrimination model to be trained to obtain the second discrimination result, the third discrimination result, and updating the training parameters of the discrimination model to be trained according to the first discrimination result, the second discrimination result, and the third discrimination result can be returned and executed until the training parameters of the discrimination model to be trained are updated N times. Among them, N can be a constant greater than or equal to 1.

[0077] It can be understood that after the training parameters of the discriminative model to be trained are updated for the Nth time, the discriminative model to be trained can be adjusted to a discriminative model with certain discriminative ability. On this basis, while fixing the training parameters of the discriminative model, the training parameters of the background image generation model to be trained can be updated according to the loss information calculated for the (N + 1)th time. And, the steps of inputting the random matrix sample into the background image generation model to be trained, obtaining the first discriminative result of the synthesized background image and the foreground recognizer for the synthesized background image, and then inputting the random matrix sample, the synthesized background image, and the background image sample set into the discriminative model to obtain the second discriminative result, the third discriminative result, and updating the training parameters of the background generation model to be trained according to the first discriminative result, the second discriminative result, and the third discriminative result can be returned until the training parameters of the background image generation model to be trained are updated N times. In this way, a background image generation model with certain data generation ability can be obtained. After that, by iteratively executing the above process of updating the training parameters of the discriminative model to be trained N times and updating the training parameters of the background image generation model to be trained N times for multiple times, a background image generation model with strong data generation ability can be trained.

[0078] In addition, it is worth noting that in the related art, usually only the synthesized data and the real data output by the generation network are input into the discriminative network, and the parameters are adjusted according to the discriminative results of the discriminative network. In the present disclosure, however, the random matrix sample, the synthesized background image, and the background image sample set are input into the discriminative model, and the parameters are adjusted according to the discriminative results of the discriminative model. This is because when the generation network outputs image data, the feature information of the image data is usually relatively redundant. Here, the redundancy can refer to spatial redundancy (i.e., the redundancy caused by the strong correlation between adjacent pixels within the image). The redundant feature information is not conducive to the convergence of the model. The random matrix sample is usually non-redundant in features. In this case, the random matrix sample, the synthesized background image, and the background image sample set can be jointly used as the input of the discriminative model, so as to reduce the redundancy of the input data and thus improve the training effect of the model.

[0079] Optionally, calculating the loss information according to the first discriminative result, the second discriminative result, and the third discriminative result includes:

[0080] Calculating the loss information through the following formula:

[0081]

[0082] where J represents the loss information, m represents the total number of sample samplings, x (i) represents the i-th background image sample, z (i) represents the i-th random matrix sample, G(z (i)) represents the i-th synthesized background image, D(x (i) ) represents the probability of identifying the i-th background image sample as an image generated by the background image generation model to be trained, T(G(z (i) )) represents the probability that the foreground object recognizer recognizes the target foreground object in the i-th synthesized background image, D(G(z (i) )) represents the probability of identifying the i-th synthesized background image as an image in the background image sample set.

[0083] Among them, the total number of sample samplings can be less than or equal to the number of random matrix samples input to the background image generation model to be trained.

[0084] It should be noted that since the loss information is usually very small, therefore, when adjusting the training parameters of the discriminative model to be trained while fixing the training parameters of the background image generation model to be trained or when adjusting the training parameters of the background image generation model to be trained while fixing the training parameters of the discriminative model, the loss information calculated can all use the above J. And in order to further improve the model accuracy of the background image generation model, when adjusting the training parameters of the background image generation model to be trained while fixing the training parameters of the discriminative model, the loss information calculated can be obtained by calculating the following formula obtained by simplifying the above formula:

[0085]

[0086] Among them, J1 represents the simplified loss information. It can be understood that D(x (i) ) represents the probability of identifying the i-th background image sample as an image generated by the background image generation model to be trained. This discrimination result has nothing to do with the model performance of the background image generation model, so it can be simplified.

[0087] Optionally, the foreground object generation model includes a network fusion module. On this basis, the above step S103 may include:

[0088] Input the second random matrix and the first sub-image into the network fusion module to obtain a fusion matrix after fusing the image matrix corresponding to the second random matrix and the first sub-image;

[0089] Based on the fusion matrix, generate an initial target sub-image, where the size of the initial target sub-image is the same as the size of the first sub-image;

[0090] Fuse the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image to obtain the initial target sub-image after edge fusion, and the target sub-image represents the initial target sub-image after edge fusion.

[0091] Exemplarily, the second random matrix and the first sub-image can be input into the network fusion module to obtain a fusion matrix after fusing the image matrix corresponding to the second random matrix and the first sub-image. Among them, fusing the image matrix corresponding to the second random matrix and the first sub-image includes methods such as matrix addition or matrix concatenation. For example, if the second random matrix is [32, 112, 112] and the image matrix corresponding to the first sub-image is [32, 112, 112], then when adding the second random matrix and the image matrix corresponding to the first sub-image, the obtained fusion matrix is [32, 112, 112], and when concatenating the second random matrix and the image matrix corresponding to the first sub-image, the obtained fusion matrix is [64, 112, 112].

[0092] On this basis, an initial target sub-image is generated according to the obtained fusion matrix. The size of the initial target sub-image is the same as that of the first sub-image. To make the fusion of the target sub-image and the second sub-image more natural, the first edge region sub-image in the initial target sub-image can be fused with the second edge region sub-image of the first sub-image to obtain an initial target sub-image after edge fusion. Among them, the target sub-image can represent the initial target sub-image after edge fusion. Among them, the edge region can refer to the region obtained by multiplying the length and width of the image by a preset range. The preset range should be set small to prevent the edge fusion operation from affecting the target foreground object generated in the target sub-image. For example, if the length of the image is 10 and the width is 5, the preset range can be set to 5%. In this case, the width of the edge region of any long side of the image is 0.25, and the width of the edge region of any wide side of the image is 0.5.

[0093] Optionally, fusing the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image includes:

[0094] For each first pixel point in the first edge region of the initial target image, the pixel value of the third pixel point at the same position as the first pixel point in the target sub-image is calculated according to the pixel value of the first pixel point and the pixel value of the second pixel point at the same position as the first pixel point in the second edge region of the first sub-image.

[0095] Among them, the method of fusing the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image can be that, when setting corresponding weight values for the first pixel point in the first edge region sub-image and the second pixel point in the second edge region, multiplying the pixel value of the first pixel point by the value of the first weight, adding the product of the pixel value of the second pixel point at the same position as the first pixel point and the value of the second weight, and determining the finally obtained sum as the pixel value of the third pixel point at the same position as the first pixel point in the target sub-image. Since the size of the initial target sub-image is the same as that of the first sub-image, the size of the first edge region can also be the same as that of the second edge region. The first weight value and the second weight value can be set according to the actual situation, and the present disclosure does not make specific limitations thereon.

[0096] For example, the pixel value of the first pixel point is 50, the pixel value of the second pixel point at the same position as the first pixel point is 60, the first weight value is 0.3, and the second weight value is 0.7. In this case, the pixel value of the third pixel point at the same position as the first pixel point in the target sub-image is 50 * 0.3 + 60 * 0.7 = 57.

[0097] Optionally, the method provided by the embodiments of the present disclosure may further include:

[0098] Obtaining the object type information of the target foreground object;

[0099] On this basis, including at least the border coordinates as the annotation information of the target foreground object in the target sub-image may include:

[0100] Taking the border coordinates and the object type information as the annotation information of the target foreground object in the target sub-image.

[0101] It should be noted that the annotation information may further include the object type information of the target foreground object. The object type information is related to the real image sample set used in the training process of the foreground object generation model, and the real image sample set includes real foreground objects. It can be understood that the network architecture in the training process of the foreground object generation model may be similar to the network architecture of GAN. That is to say, in the training process of the foreground object generation model, a foreground object discrimination model and a foreground object generation model may be involved. Among them, the real image sample set and the target sub-image sample generated by the foreground object generation model to be trained according to the input second random matrix sample and the first sub-image can be input into the foreground object discrimination model to be trained to obtain corresponding discrimination results, and the training parameters of the foreground object discrimination model to be trained and the foreground object generation model to be trained can be adjusted in turn according to a training process similar to that of the background image generation model, so as to obtain a foreground object generation model that can generate the target sub-image, and the target sub-image includes the target foreground object. In a possible implementation manner, the object type information of the target foreground object may be the same as the object type information of the real foreground object in the real image sample set. In this case, the object type information of the target foreground object can be determined according to the object type information of the real foreground object.

[0102] On this basis, the border coordinates and the object type information can be used as the annotation information of the target foreground object in the target sub-image.

[0103] In addition, it is worth noting that in the related art, a trained model can be used to pre-annotate data, or data with annotations can be manufactured by splicing. However, the pre-annotated data in the former still needs to be manually screened, with low efficiency, and the annotated data generated in the latter is relatively rigid, and the quality of the annotated data cannot be guaranteed. In contrast, the present disclosure obtains the first sub-image and the second sub-image by segmenting the background image, and generates a target sub-image based on the first sub-image, and the target sub-image includes the target foreground object. Thus, after fusing the target sub-image into the second sub-image, more natural annotated image data can be obtained. That is to say, the technical solution provided by the embodiments of the present disclosure can generate more natural annotated image data without manual annotation or manual screening of data.

[0104] Using the above method, by inputting the first random matrix into the trained background image generation model, the background image and the border coordinates output by the background image generation model are obtained. Since the border coordinates can be used to represent the position of the first sub-image of the foreground object to be generated in the background image, the background image can be segmented according to the border coordinates to obtain the first sub-image and the second sub-image. On this basis, by inputting the second random matrix and the first sub-image into the trained foreground object generation model, the target sub-image output by the foreground object generation model is obtained. Wherein, the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes the target foreground object. At least the border coordinates are used as the annotation information of the target foreground object in the target sub-image, and the annotated target sub-image is fused with the second sub-image, so that the annotated image data can be obtained. Compared with the method of manually annotating data in the related art, the method of generating annotated image data by the model in the present disclosure is more efficient.

[0105] Based on the same inventive concept, the present disclosure also provides an apparatus for generating annotated image data. Refer to Figure 2 , Figure 2 FIG. is a block diagram of an apparatus for generating annotated image data according to an exemplary embodiment of the present disclosure. As Figure 2 shown, the apparatus 100 for generating annotated image data includes:

[0106] A first input module 101, configured to input the first random matrix into the trained background image generation model, and obtain the background image and the border coordinates output by the background image generation model, where the border coordinates are used to represent the position of the first sub-image of the foreground object to be generated in the background image;

[0107] A segmentation module 102, configured to segment the background image according to the border coordinates to obtain the first sub-image and the second sub-image;

[0108] A second input module 103, configured to input the second random matrix and the first sub-image into the trained foreground object generation model, and obtain the target sub-image output by the foreground object generation model, where the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes the target foreground object;

[0109] A fusion module 104, configured to use at least the border coordinates as the annotation information of the target foreground object in the target sub-image, and fuse the annotated target sub-image with the second sub-image to obtain the annotated image data.

[0110] By using the above device, the first random matrix is input into the trained background image generation model to obtain the background image and the border coordinates output by the background image generation model. Since the border coordinates can be used to represent the position of the first sub-image of the foreground object to be generated in the background image, the background image can be segmented according to the border coordinates to obtain the first sub-image and the second sub-image. On this basis, by inputting the second random matrix and the first sub-image into the trained foreground object generation model, the target sub-image output by the foreground object generation model can be obtained. Wherein, the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes the target foreground object. At least the border coordinates are used as the annotation information of the target foreground object in the target sub-image, and the annotated target sub-image is fused with the second sub-image, so that the annotated image data can be obtained. Compared with the method of manually annotating data in the related art, the device for generating annotated image data by the model in the present disclosure has higher efficiency.

[0111] Optionally, the device 100 further includes a training module, and the training module is configured to:

[0112] Input the random matrix sample into the background image generation model to be trained to obtain the synthetic background image output by the background image generation model to be trained;

[0113] Input the synthetic background image into the foreground recognizer to obtain the first discrimination result output by the foreground recognizer on whether there is a target foreground object in the synthetic background image;

[0114] Input the random matrix sample, the synthetic background image, and the background image sample set into the discrimination model to be trained to obtain the second discrimination result on whether the discrimination model to be trained discriminates the synthetic background image as an image in the background image sample set, and to obtain the third discrimination result on whether the discrimination model to be trained discriminates each background image sample in the background image sample set as an image generated by the background image generation model to be trained;

[0115] Calculate the loss information according to the first discrimination result, the second discrimination result, and the third discrimination result;

[0116] Adjust the training parameters of the discrimination model to be trained according to the loss information to obtain the updated discrimination model to be trained, and return to execute the step of inputting the random matrix sample, the synthetic background image, and the background image sample set into the discrimination model to be trained to obtain the second discrimination result and the third discrimination result until the training parameters of the discrimination model to be trained are updated N times.

[0117] Optionally, the device 100 further includes:

[0118] An update module, configured to update the training parameters of the background image generation model according to the loss information calculated in the (N + 1)-th time after updating the training parameters of the discriminative model to be trained for the N-th time.

[0119] Optionally, the training module is further configured to:

[0120] Calculate the loss information through the following formula:

[0121]

[0122] where J represents the loss information, m represents the total number of sample samplings, x (i) represents the i-th background image sample, z (i) represents the i-th random matrix sample, G(z (i) ) represents the i-th synthesized background image, D(x (i) ) represents the probability of identifying the i-th background image sample as an image generated by the discriminative model to be trained for the background image, T(G(z (i) )) represents the probability that the foreground object recognizer recognizes a target foreground object in the i-th synthesized background image, and D(G(z (i) )) represents the probability of identifying the i-th synthesized background image as an image in the background image sample set.

[0123] Optionally, the foreground object generation model includes a network fusion module, and the second input module 103 is further configured to:

[0124] Input the second random matrix and the first sub-image into the network fusion module to obtain a fusion matrix after fusing the image matrices corresponding to the second random matrix and the first sub-image;

[0125] Generate an initial target sub-image based on the fusion matrix, where the size of the initial target sub-image is the same as the size of the first sub-image;

[0126] Fuse the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image to obtain the initial target sub-image after edge fusion, and the target sub-image represents the initial target sub-image after edge fusion.

[0127] Optionally, the apparatus 100 further includes:

[0128] A calculation module, configured to calculate, for each first pixel point in a first edge region of the initial target graph, a pixel value of a third pixel point at the same position as the first pixel point in the target sub-image according to a pixel value of the first pixel point and a pixel value of a second pixel point at the same position as the first pixel point in a second edge region of the first sub-image.

[0129] Optionally, the apparatus 100 further includes:

[0130] An acquisition module, configured to acquire object type information of the target foreground object;

[0131] The fusion module 104 is further configured to:

[0132] Use the border coordinates and the object type information as annotation information of the target foreground object in the target sub-image.

[0133] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0134] Based on the same inventive concept, an embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-described method for generating annotated image data are implemented.

[0135] Specifically, the computer-readable storage medium may be a flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, a server, a public cloud server, and so on.

[0136] Regarding the computer-readable storage medium in the above embodiments, the implementation of the steps of the method for generating annotated image data when the computer program stored thereon is executed has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0137] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, which includes:

[0138] A memory, on which a computer program is stored;

[0139] A processor, configured to execute the computer program in the memory to implement the steps of the above-described method for generating annotated image data.

[0140] Figure 3is a block diagram of an electronic device 200 shown according to an exemplary embodiment. As Figure 3 shown, the electronic device 200 may include: a processor 201, a memory 202. The electronic device 200 may also include one or more of a multimedia component 203, an input / output (I / O) interface 204, and a communication component 205.

[0141] Among them, the processor 201 is used to control the overall operation of the electronic device 200 to complete all or part of the steps in the above-mentioned method for generating annotated image data. The memory 202 is used to store various types of data to support the operation of the electronic device 200. These data may include, for example, instructions for any application or method operating on the electronic device 200, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 202 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 203 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 202 or sent through the communication component 205. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 204 provides an interface between the processor 201 and other interface modules, and the above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 205 is used for wired or wireless communication between the electronic device 200 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G or 5G, NB-IoT (Narrow Band Internet of Things), or a combination of one or more of them. Accordingly, the communication component 205 may include: a Wi-Fi module, a Bluetooth module, an NFC module.

[0142] In one exemplary embodiment, the electronic device 200 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the above-described method for generating annotated image data.

[0143] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0144] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.

[0145] Furthermore, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. A method for generating annotated image data, characterized in that, The method includes: Inputting a first random matrix into a trained background image generation model to obtain a background image and border coordinates output by the background image generation model, where the border coordinates are used to represent the position of a first sub-image of a foreground object to be generated in the background image; Segmenting the background image according to the border coordinates to obtain the first sub-image and a second sub-image; Inputting a second random matrix and the first sub-image into a trained foreground object generation model to obtain a target sub-image output by the foreground object generation model, where the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes a target foreground object; Using at least the border coordinates as annotation information of the target foreground object in the target sub-image, and fusing the annotated target sub-image with the second sub-image to obtain annotated image data; The training process of the background image generation model includes: Inputting a random matrix sample into a background image generation model to be trained to obtain a synthesized background image output by the background image generation model to be trained; Inputting the synthesized background image into a foreground recognizer to obtain a first discrimination result output by the foreground recognizer on whether there is a target foreground object in the synthesized background image; Inputting the random matrix sample, the synthesized background image, and a background image sample set into a discrimination model to be trained to obtain a second discrimination result on whether the discrimination model to be trained discriminates the synthesized background image as an image in the background image sample set, and a third discrimination result on whether the discrimination model to be trained discriminates each background image sample in the background image sample set as an image generated by the background image generation model to be trained; Calculating loss information according to the first discrimination result, the second discrimination result, and the third discrimination result; Adjusting the training parameters of the discrimination model to be trained according to the loss information to obtain an updated discrimination model to be trained, and returning to execute the step of inputting the random matrix sample, the synthesized background image, and the background image sample set into the discrimination model to be trained to obtain the second discrimination result and the third discrimination result until the training parameters of the discrimination model to be trained are updated N times; The calculating loss information according to the first discrimination result, the second discrimination result, and the third discrimination result includes: Calculating the loss information through the following formula: ; Where J represents the loss information, and m represents the total number of sample samplings, represents the i-th background image sample, represents the i-th random matrix sample, represents the i-th synthesized background image, represents the probability of identifying the i-th background image sample as an image generated by the to-be-trained background image generation model, represents the probability that the foreground object recognizer recognizes a target foreground object in the i-th synthesized background image, represents the probability of identifying the i-th synthesized background image as an image in the background image sample set.

2. The method according to claim 1, characterized in that, The method further includes: After the Nth update of the training parameters of the discrimination model to be trained, updating the training parameters of the background image generation model according to the loss information calculated in the (N + 1)th time.

3. The method according to claim 1, wherein The foreground object generation model includes a network fusion module. The inputting a second random matrix and the first sub-image into a trained foreground object generation model to obtain a target sub-image output by the foreground object generation model includes: Inputting the second random matrix and the first sub-image into the network fusion module to obtain a fusion matrix after fusing the image matrices corresponding to the second random matrix and the first sub-image; Generate an initial target sub-image based on the fusion matrix, where the size of the initial target sub-image is the same as that of the first sub-image; Fuse the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image to obtain the initial target sub-image after edge fusion, and the target sub-image represents the initial target sub-image after edge fusion.

4. The method according to claim 3, wherein The fusing the first edge region sub-image in the initial target sub-image with the second edge region sub-image of the first sub-image includes: For each first pixel point in the first edge region of the initial target sub-image, calculate the pixel value of the third pixel point at the same position as the first pixel point in the target sub-image according to the pixel value of the first pixel point and the pixel value of the second pixel point at the same position as the first pixel point in the second edge region of the first sub-image.

5. The method according to claim 1, wherein The method further includes: Obtain the object type information of the target foreground object; The at least using the border coordinates as the annotation information of the target foreground object in the target sub-image includes: Using the border coordinates and the object type information as the annotation information of the target foreground object in the target sub-image.

6. An image data generation device with annotations, characterized in that, The apparatus includes: A first input module, configured to input a first random matrix into a trained background image generation model to obtain a background image and border coordinates output by the background image generation model, where the border coordinates are used to represent the position of a first sub-image of a foreground object to be generated in the background image; A segmentation module, configured to segment the background image according to the border coordinates to obtain the first sub-image and the second sub-image; A second input module, configured to input a second random matrix and the first sub-image into a trained foreground object generation model to obtain a target sub-image output by the foreground object generation model, where the size of the target sub-image is the same as that of the first sub-image, and the target sub-image includes a target foreground object; A fusion module, configured to at least use the border coordinates as the annotation information of the target foreground object in the target sub-image, and fuse the annotated target sub-image with the second sub-image to obtain annotated image data; The apparatus further includes a training module, and the training module is configured to: Input a random matrix sample into a background image generation model to be trained to obtain a synthetic background image output by the background image generation model to be trained; Input the synthetic background image into a foreground recognizer to obtain a first discrimination result of whether there is a target foreground object in the synthetic background image output by the foreground recognizer; Input the random matrix sample, the synthetic background image, and a background image sample set into a discrimination model to be trained to obtain a second discrimination result of whether the discrimination model to be trained discriminates the synthetic background image as an image in the background image sample set, and a third discrimination result of whether the discrimination model to be trained discriminates each background image sample in the background image sample set as an image generated by the background image generation model to be trained; Calculate loss information according to the first discrimination result, the second discrimination result, and the third discrimination result; Adjust the training parameters of the discriminant model to be trained according to the loss information, obtain the updated discriminant model to be trained, and return to execute the steps of inputting the random matrix sample, the synthetic background image, and the background image sample set into the discriminant model to be trained to obtain the second discrimination result and the third discrimination result until the training parameters of the discriminant model to be trained are updated N times; The training module is further configured to: Calculate the loss information through the following formula: ; where J represents the loss information, m represents the total number of sample samplings, represents the i-th background image sample, represents the i-th random matrix sample, represents the i-th synthesized background image, represents the probability of identifying the i-th background image sample as an image generated by the to-be-trained background image generation model, represents the probability that the foreground object recognizer recognizes a target foreground object in the i-th synthesized background image, represents the probability of identifying the i-th synthesized background image as an image in the background image sample set.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-5.

8. An electronic device, characterized in that, Including: A memory storing a computer program thereon; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method, apparatus, system and storage medium for generating training data

    CN109146830A