Label data generation device and method

By generating features of line drawings and depth drawings, the labeled data is generated by combining prompt words, the problem of insufficient coverage of labeled data in the prior art is solved, and the efficiency and accuracy of model training are improved.

CN120298784APending Publication Date: 2025-07-11OMRON SHANGHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375859.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the existing manual or semi-automatic labeling methods, the coverage breadth of data is limited, resulting in a low coverage breadth of labeling data, which in turn leads to poor generalization of the final training model, and manual labeling is time-consuming and labor-intensive. Semi-automatic labeling has the problem of low data quality.

Method used

By acquiring the initial image data, the first labeling is used to generate line drawings and depth maps, the edge detection control network and depth map control network extract features, the text feature vector is generated in combination with the prompt words, and the diffusion generation model is used to generate label data.

Benefits of technology

It realizes the generation of labeled data with a larger breadth through a small amount of labeled data, which reduces the difficulty of training artificial intelligence models, and improves the accuracy of the model and the quality of labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298784A_ABST
    Figure CN120298784A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an annotation data generation device and method, and the device comprises an obtaining unit which obtains initial image data; the labeling unit is used for labeling the initial image data for the first time to obtain labeling information; the first generation unit is used for generating a line map and a depth map corresponding to the initial image data according to the annotation information; the second generation unit is used for generating a first condition according to the line draft map, generating a second condition according to the depth map and generating a third condition according to the prompt word; and the third generation unit is used for generating annotation data corresponding to the initial image data according to the first condition, the second condition and the third condition. Therefore, the annotation data with a larger scale can be obtained through a small amount of annotation data, the difficulty of training the annotation data of the artificial intelligence model is further reduced, the breadth of the generated annotation data is larger, and improvement of the precision of the model trained according to the generated annotation data is more facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to an apparatus and method for generating labeled data. Background Art

[0002] Data annotation is a key step in the development of machine learning and artificial intelligence (AI) systems. The core of data annotation is to attach semantic information to data so that machines can understand and use this data. That is to say, labeled data can help machine learning models understand complex data patterns. For example, the shape of an object in an image or the semantic structure of text, etc. Therefore, labeled data is the basis for training models. Most AI models rely on labeled data for supervised learning. A supervised learning model needs to learn the correspondence between inputs and outputs through a large amount of labeled data. High-quality labeled data is conducive to significantly improving the performance and accuracy of the model.

[0003] Existing data annotation methods include: manually annotating data and using semi-automatic annotation tools to assist in generating labeled data.

[0004] It should be noted that the above introduction to the technical background is only for the convenience of clearly and completely explaining the technical solutions of this application and facilitating the understanding of those skilled in the art. It cannot be considered that the above technical solutions are well-known to those skilled in the art just because these solutions are described in the background art part of this application. Summary of the Invention

[0005] The inventors found that in the existing manual annotation or semi-automatic annotation methods, the coverage breadth of data is limited by the data source. Therefore, the coverage breadth of the labeled data is relatively low, which in turn results in poor generalization of the final trained model.

[0006] During the training process of some AI models such as object detection models, a large amount of labeled data is required. Manual annotation depends on the understanding and judgment of annotators, which has strong subjectivity. Moreover, manually annotating large-scale data requires a large amount of time and economic costs, and long-term annotation may also lead to an increase in the error rate. Although a larger scale of labeled data can be obtained through semi-automatic annotation methods, there are still data samples that need to be corrected manually. In addition, if the initial data has class imbalance, this data deviation may be amplified in the semi-automatic process, resulting in low-quality labeled data.

[0007] To solve at least one of the above problems or other similar problems, the embodiments of this application provide an apparatus and method for generating labeled data that can obtain labeled data with a wider breadth through a small amount of labeled data.

[0008] According to the first aspect of the embodiments of this application, there is provided an apparatus for generating labeled data, the apparatus includes:

[0009] An acquisition unit that acquires initial image data;

[0010] A labeling unit that performs a first labeling on the initial image data to obtain labeling information;

[0011] A first generation unit that generates a line drawing and a depth map corresponding to the initial image data according to the labeling information;

[0012] A second generation unit that generates a first condition according to the line drawing, generates a second condition according to the depth map, and generates a third condition according to a prompt;

[0013] A third generation unit that generates labeling data corresponding to the initial image data according to the first condition, the second condition, and the third condition.

[0014] According to the second aspect of the embodiments of the present application, the labeling unit performs the first labeling on the initial image data by at least one of a manual labeling method, a semi-automatic labeling method, and an automatic labeling method.

[0015] According to the third aspect of the embodiments of the present application, the labeling unit performing a first labeling on the initial image data includes: performing at least one of rectangular box labeling, key point labeling, and region labeling on the initial image data;

[0016] The labeling information includes edge information of the initial image data.

[0017] According to the fourth aspect of the embodiments of the present application, the first generation unit generates the line drawing according to the edge information and obtains the depth map through a 2D depth model.

[0018] According to the fifth aspect of the embodiments of the present application, the first condition is an edge feature of the line drawing, and the second condition is a depth feature of the depth map.

[0019] According to the sixth aspect of the embodiments of the present application, the second generation unit uses an edge detection control network to generate the first condition and uses a depth map control network to generate the second condition;

[0020] Wherein, the second generation unit extracts the edge feature of the line drawing as the first condition through the edge detection control network, and the second generation unit extracts the depth feature of the depth map as the second condition through the depth map control network.

[0021] According to the seventh aspect of the embodiments of the present application, the third condition is a text feature vector generated according to the prompt;

[0022] The prompt words include positive prompt words and / or negative prompt words.

[0023] According to the eighth aspect of the embodiments of the present application, the third generation unit generates the annotation data corresponding to the initial image data according to the edge feature, the depth feature, and the text feature vector.

[0024] According to the ninth aspect of the embodiments of the present application, the third generation unit generates the annotation data corresponding to the initial image data through a diffusion generation model.

[0025] According to the tenth aspect of the embodiments of the present application, there is provided a method for generating annotation data, the method including:

[0026] Obtain initial image data;

[0027] Perform a first annotation on the initial image data to obtain annotation information;

[0028] Generate a line drawing and a depth map corresponding to the initial image data according to the annotation information;

[0029] Generate a first condition according to the line drawing, generate a second condition according to the depth map, and generate a third condition according to the prompt words;

[0030] Generate the annotation data corresponding to the initial image data according to the first condition, the second condition, and the third condition.

[0031] One of the beneficial effects of the embodiments of the present application is that: according to the annotation data generation device and method provided by the embodiments of the present application, it is possible to obtain a larger-scale annotation data through a small amount of annotation data, thereby reducing the difficulty of annotating data for training an artificial intelligence model, and the generated annotation data has a wider breadth, which is more conducive to improving the accuracy of the model trained according to the generated annotation data.

[0032] Referring to the following description and the drawings, specific embodiments of the present application are disclosed in detail, indicating the ways in which the principles of the present application can be adopted. It should be understood that the embodiments of the present application are not limited in scope thereby. Within the spirit and terms of the appended claims, the embodiments of the present application include many changes, modifications, and equivalents.

[0033] Features described and / or illustrated for one embodiment can be used in the same or similar way in one or more other embodiments, combined with the features in other embodiments, or replace the features in other embodiments.

[0034] It should be emphasized that the term "comprising / including", as used herein, refers to the presence of features, whole units, steps or components, but does not exclude the presence or addition of one or more other features, whole units, steps or components. Description of the Drawings

[0035] The accompanying drawings included are used to provide a further understanding of the embodiments of the present application, which form a part of the specification, illustrate the embodiments of the present application, and, together with the written description, explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0036] Figure 1 is a schematic diagram of a labeled data generation device according to an embodiment of the present application;

[0037] Figure 2 is a schematic diagram of the first labeling of the initial image data according to the present application;

[0038] Figure 3 is a schematic diagram of the initial image data according to an embodiment of the present application;

[0039] Figure 4 is as Figure 3 a schematic diagram of a line drawing corresponding to the initial image data shown;

[0040] Figure 5 is as Figure 3 a schematic diagram of a depth map corresponding to the initial image data shown;

[0041] Figure 6 is a schematic diagram of the workflow of the labeled data generation device according to an embodiment of the present application;

[0042] Figure 7 is another schematic diagram of the workflow of the labeled data generation device according to an embodiment of the present application;

[0043] Figure 8 is a schematic diagram of the labeled data generated by the device according to an embodiment of the present application;

[0044] Figure 9 is a schematic diagram of the labeled data generation method according to an embodiment of the present application;

[0045] Figure 10 is a schematic diagram of a computer device according to an embodiment of the present application. Detailed Embodiments

[0046] Referring to the accompanying drawings, the foregoing and other features of the present application will become apparent through the following description. In the description and drawings, specific embodiments of the present application are specifically disclosed, which show some embodiments in which the principles of the present application can be adopted. It should be understood that the present application is not limited to the described embodiments. On the contrary, the present application includes all modifications, variations, and equivalents falling within the scope of the appended claims.

[0047] In the embodiments of the present application, the term "and / or" includes any one and all combinations of one or more of the associated listed terms. Terms such as "comprising", "including", "having", etc. mean the presence of the stated features, elements, components, or assemblies, but do not exclude the presence or addition of one or more other features, elements, components, or assemblies.

[0048] In the embodiments of the present application, the singular forms "a", "the", etc. may include the plural forms and should be broadly understood as "a kind" or "a class" rather than being limited to the meaning of "one"; in addition, the term "the" should be understood to include both the singular form and the plural form unless the context clearly indicates otherwise. In addition, the term "according to" should be understood as "at least partially according to...", and the term "based on" should be understood as "at least partially based on...", unless the context clearly indicates otherwise.

[0049] The following describes the embodiments of the present application with reference to the accompanying drawings.

[0050] Embodiments of the first aspect

[0051] The embodiments of the present application provide a labeled data generation device. Figure 1 is a schematic diagram of the labeled data generation device 100 in the embodiments of the present application; as Figure 1 shown, the labeled data generation device 100 includes: an acquisition unit 101, a labeling unit 102, a first generation unit 103, a second generation unit 104, and a third generation unit 105.

[0052] Among them, the acquisition unit 101 is used to acquire initial image data, and the labeling unit 102 is used to perform the first labeling on the initial image data to obtain labeling information.

[0053] The first generation unit 103 generates a line drawing and a depth map corresponding to the initial image data according to the obtained labeling information. The second generation unit 104 generates a first condition according to the line drawing, a second condition according to the depth map, and a third condition according to the prompt word. The third generation unit 105 generates labeled data corresponding to the initial image data according to the generated first condition, second condition, and third condition.

[0054] The above-mentioned annotation data generation device 100 provided by the embodiments of the present application can obtain a larger-scale annotation data for training a target detection model through a small amount of annotation data, thereby reducing the difficulty of annotating data for training an artificial intelligence model, and the generated annotation data has a wider range, which is more conducive to improving the accuracy of the model trained according to the generated annotation data.

[0055] Figure 2 is a schematic diagram of the first annotation of the initial image data in the present application. As Figure 2 shown, in order to draw the line drawing of the initial image data, it is necessary to perform the first annotation on the initial image data.

[0056] In some embodiments, the annotation unit 102 can perform the first annotation on the initial image data by at least one of a manual annotation method, a semi-automatic annotation method, and an automatic annotation method.

[0057] Specifically, a fully manual annotation method can be adopted, or an automatic annotation tool such as SAM (SegmentAnything Model) can be used for automatic annotation, or semi-automatic annotation tools such as VIA (VGG ImageAnnotator) and Label studio can be used for semi-automatic annotation.

[0058] During the process of performing the first annotation on the initial image data, since the amount of initial image data to be processed is small, the amount of data to be processed by the above annotation methods is low, which avoids errors in manual annotation or automatic annotation when processing a large amount of data, reduces the time consumption, and improves the annotation efficiency.

[0059] As Figure 2 shown, in the above embodiment, as Figure 2 shown in (a), the initial image data contains multiple target objects, and each selected target object has its corresponding edge. In order to obtain the edge information that can represent the edge of the target object, it is necessary to annotate the target objects in the initial image data.

[0060] In the above embodiment, the annotation unit 102 performing the first annotation on the initial image data includes: performing at least one of rectangular box annotation, key point annotation, and region annotation on the initial image data to obtain the annotation information of the initial image data. For example, for rectangular box annotation, it can be to frame the area of the object with a rectangular box, and this rectangular box is the edge of the object; for key point annotation, it can be to select the key points of the object in the image, and connecting these key points can represent the edge of the object; for region annotation, it can be to select the area where the object is located, and the edge of the selected area is the edge of the object.

[0061] Through the above-mentioned first annotation of the initial image data by the annotation unit 102, the annotation information of the initial image data can be obtained, and the annotation information at least includes the edge information of the initial image data. This application is not limited to this, and the annotation information may also include other information, which can refer to related technologies.

[0062] In the above embodiment, the first generation unit 103 can generate a line drawing corresponding to the initial image data according to the above-mentioned edge information, and obtain a depth map corresponding to the initial image data through a 2D depth model.

[0063] Figure 3 is a schematic diagram of the initial image data in an embodiment of this application, Figure 4 and Figure 5 are respectively a schematic diagram of the line drawing and the depth map corresponding to the initial image data as shown in Figure 3 . As shown in Figure 3 , the initial image data is a photo, and there are four cardboard boxes with different heights in the direction perpendicular to the ground in the photo. For example, the four cardboard boxes are used as the target objects for object detection and are first annotated, and the edge information of the four cardboard boxes is obtained as the annotation information of the four cardboard boxes. Then, according to the Figure 3 edge information of the cardboard boxes in, the corresponding line drawing as shown in Figure 4 can be drawn. In addition, according to the initial image data, a depth map corresponding to the initial image data as shown in Figure 5 can be obtained through a 2D depth model. The depth map has the depth information of the initial image data and can reflect the spatial characteristics of the initial image data.

[0064] In the embodiment of this application, 2D depth models such as DepthNet, MonoDepth / MonoDepth2, and FastDepth can be used to obtain the depth map of the initial image data based on the annotation information of the initial image data. For details, reference can be made to the related technologies in this field, which will not be elaborated here.

[0065] In the above embodiment, according to the line drawing and the depth map, the second generation unit 104 can generate the first condition and the second condition respectively. In addition, the second generation unit 104 can also generate a third condition according to the prompt words.

[0066] In some embodiments, the first condition is the edge feature of the line drawing, and the second condition is the depth feature of the depth map. Among them, the edge feature of the line drawing can reflect the shape, position, and regional size of the target object in the initial image data, and the depth feature of the depth map further reflects the structural feature and spatial feature of the initial image data. Thus, it can be ensured that in the annotation data generated according to the first condition and the second condition, the shape and size of the target object and the spatial structure of the image are consistent with the initial image data, reducing the error rate in the process of generating annotation data.

[0067] In the above embodiment, the second generation unit 104 can use the edge detection control network to generate the first condition and use the depth map control network to generate the second condition. For example, the second generation unit 104 extracts the edge feature of the line drawing as the first condition through the edge detection control network and extracts the depth feature of the depth map as the second condition through the depth map control network.

[0068] In the above embodiment, the edge control network is, for example, a Canny edge detection network, etc., and the depth control network is, for example, PackNet-SfM, DPT (Dense Prediction Transformer), etc. The present application does not limit this.

[0069] In some embodiments, the third condition is the text feature vector generated according to the prompt; among them, the prompt can include, for example, a positive prompt, and / or, a negative prompt.

[0070] Specifically, the prompt can be a text description input by the user, directly guiding the model to generate image content that conforms to the semantics, which can determine the main body and style of the generated data. In addition, by inputting a negative prompt, some unnecessary elements can be excluded, reducing common errors in the generation process.

[0071] In the above embodiment, for example, the prompt can be encoded by CLIP (Contrastive Language-Image Pretraining) to convert the prompt into a text feature vector. CLIP (Contrastive Language-Image Pretraining) is a multimodal model that can map text and images to the same semantic space. In the diffusion model, CLIP encoding is the bridge connecting text prompts and image generation. By encoding the prompt to generate a text feature vector and using this text feature vector to guide the generation of annotation data can significantly improve the breadth of the generated annotation data.

[0072] In some embodiments, the third generation unit 105 can generate annotation data corresponding to the initial image data through a diffusion generation model according to the first condition, the second condition, and the third condition.

[0073] For example, the third generation unit 105 can generate annotation data corresponding to the initial image data through a diffusion generation model according to the above-mentioned edge features, the above-mentioned depth features, and the above-mentioned text feature vector.

[0074] This application does not limit the diffusion generation model. For example, it can be Stable Diffusion or other diffusion generation models. For details, reference can be made to related technologies and will not be elaborated here.

[0075] Figure 6 is a schematic diagram of the working process of the annotation data generation device 100 according to an embodiment of this application. As Figure 6 shown, the edge control network generates edge features corresponding to the initial image data based on the line drawing of the initial image data, the depth control network generates depth features corresponding to the initial image data based on the depth map of the initial image data, the encoder generates a text feature vector based on the prompt word. The above-mentioned edge features, depth features, and text feature vector are used as inputs to the diffusion generation model, and the diffusion generation model generates annotation data corresponding to the initial image data based on this.

[0076] In the above embodiment, as Figure 6 shown, the initial image data can also be used as an input to the diffusion generation model to control the noise of the generation result and make the generated annotation data maintain a unified style. Similarly, a random number (Random Seed) can also be input to control the noise value in the generation result. For details, reference can be made to related technologies in this field and will not be elaborated here.

[0077] Figure 7 is another schematic diagram of the working process of the annotation data generation device 100 according to an embodiment of this application; as Figure 7 shown, through the above first condition (taking edge features as an example), the second condition (taking depth features as an example), and the third condition (taking text feature vector as an example), annotation data with a structure feature consistent with the initial image data can be generated. In Figure 7 the example, the positive prompt words can be, for example, long shot, top view, clear background, corrugated cardboard box stack on the ground, many (stickers: 1.5) on box and many printed logo, QR code on box, etc. This application does not limit this.

[0078] Figure 8 is a schematic diagram of the labeled data generated by the device according to the embodiments of the present application. As Figure 8 shown, in the finally generated labeled data, there are different image features such as stickers and QR codes on the carton, but the edge features and depth features of the carton are consistent with the initial image data.

[0079] Through the labeled data generation device 100 provided by the above embodiments of the present application, it is possible to generate labeled data with a wider range through a small amount of labeled data, which not only reduces the difficulty of obtaining training data for the target detection model, but also improves the recognition accuracy of the target detection model. Moreover, through the diffusion generation model, any amount of labeled data can be extracted, with higher operation convenience and higher data quality obtained, which is more conducive to the subsequent training of the model.

[0080] Embodiments of the second aspect

[0081] The embodiments of the present application provide a labeled data generation method, which can be applied to the labeled data generation device 100 described in the embodiments of the first aspect. Figure 9 is a schematic diagram of the labeled data generation method of the embodiments of the present application. As Figure 9 shown, the method includes:

[0082] 901, obtaining initial image data;

[0083] 902, performing a first annotation on the initial image data to obtain annotation information;

[0084] 903, generating a line drawing and a depth map corresponding to the initial image data according to the annotation information;

[0085] 904, generating a first condition according to the line drawing, generating a second condition according to the depth map, and generating a third condition according to the prompt;

[0086] 905, generating labeled data corresponding to the initial image data according to the first condition, the second condition, and the third condition.

[0087] Since in the embodiments of the first aspect, the structure of the labeled data generation device 100 has been described in detail, the content is incorporated herein and will not be repeated here.

[0088] Through the labeled data generation method provided by the above embodiments of the present application, it is possible to generate labeled data with a wider range through a small amount of labeled data, which not only reduces the difficulty of obtaining training data for the target detection model, but also improves the recognition accuracy of the target detection model. Moreover, through the diffusion generation model, any amount of labeled data can be extracted, with higher operation convenience and higher data quality obtained, which is more conducive to the subsequent training of the model.

[0089] Embodiments of the third aspect

[0090] An embodiment of the present application provides a computer device, which includes the device 100 as described in the embodiment of the first aspect, the content of which is incorporated herein. The computer device may be, for example, a computer, a server, a workstation, a laptop computer, a smart phone, etc.; however, the embodiments of the present application are not limited thereto.

[0091] Figure 10 is a schematic diagram of the computer device according to the embodiment of the present application. As Figure 10 shown, the computer device 1000 according to the embodiment of the present application may include: a processor (such as a central processing unit CPU) 1010 and a memory 1020; the memory 1020 is coupled to the central processor 1010. The memory 1020 can store various data; in addition, it also stores a program for information processing and executes the program under the control of the processor 1010.

[0092] In some embodiments, the function of the annotation data generation device 100 is integrated into the processor 1010 for implementation. Among them, the processor 1010 is configured to implement the annotation data generation method as described in the embodiment of the second aspect.

[0093] In some embodiments, the annotation data generation device 100 is separately configured from the processor 1010. For example, the annotation data generation device 100 can be configured as a chip connected to the processor 1010, and the function of the annotation data generation device 100 is realized through the control of the processor 1010.

[0094] For example, the processor 1010 is configured to perform the following controls: obtain initial image data; perform a first annotation on the initial image data to obtain annotation information; generate a line drawing and a depth map corresponding to the initial image data according to the annotation information; generate a first condition according to the line drawing, generate a second condition according to the depth map, and generate a third condition according to the prompt word; generate annotation data corresponding to the initial image data according to the first condition, the second condition, and the third condition.

[0095] In addition, as Figure 10 shown, the computer device 1000 may further include: an input / output (I / O) device 1030, a display 1040, etc.; among them, the functions of the above components are similar to those in the prior art and will not be elaborated here. It should be noted that the computer device 1000 does not necessarily have to include Figure 10 all the components shown in; in addition, the computer device 1000 may further include Figure 10 components not shown in, and reference may be made to the related art.

[0096] An embodiment of the present application further provides a computer-readable program, which, when executed on a computer device, causes the computer to execute the annotation data generation method as described in the embodiments of the second aspect in the computer device.

[0097] An embodiment of the present application further provides a storage medium storing a computer-readable program, wherein the computer-readable program causes the computer to execute the annotation data generation method as described in the embodiments of the second aspect in the computer device.

[0098] The above devices and methods of the present application can be implemented by hardware or by a combination of hardware and software. The present application relates to such a computer-readable program that, when executed by a logic component, can cause the logic component to implement the above-described device or component, or cause the logic component to implement the above-described various methods or steps. The present application also relates to a storage medium for storing the above program, such as a hard disk, a magnetic disk, an optical disk, a DVD, a flash memory, etc.

[0099] The embodiments of the present application have been described above in conjunction with specific implementation manners, but those skilled in the art should understand that these descriptions are exemplary and do not limit the protection scope of the embodiments of the present application. Those skilled in the art can make various variations and modifications to the embodiments of the present application according to the spirit and principle of the embodiments of the present application, and these variations and modifications are also within the scope of the embodiments of the present application.

[0100] The method / device described in conjunction with the embodiments of the present application can be directly embodied as hardware, a software module executed by a processor, or a combination of the two. For example, one or more of the functional block diagrams shown in the figure and / or a combination of one or more of the functional block diagrams can correspond to each software module in the computer program flow, and can also correspond to each hardware module. These software modules can respectively correspond to the respective steps shown in the figure. These hardware modules can be implemented by solidifying these software modules using a field programmable gate array (FPGA).

[0101] A software module may be located in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. A storage medium may be coupled to the processor such that the processor can read information from the storage medium and write information to the storage medium; or the storage medium may be an integral part of the processor. The processor and the storage medium may be located in an ASIC. The software module may be stored in the memory of the mobile terminal or in a memory card that can be inserted into the mobile terminal. For example, if the device (such as a mobile terminal) uses a larger-capacity MEGA-SIM card or a large-capacity flash memory device, the software module may be stored in the MEGA-SIM card or the large-capacity flash memory device.

[0102] One or more of the functional blocks described in the drawings and / or one or more combinations of the functional blocks may be implemented as a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or any suitable combination thereof for performing the functions described in this application. One or more of the functional blocks described in the drawings and / or one or more combinations of the functional blocks may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in communication combination with a DSP, or any other such configuration.

[0103] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings. Many features and advantages of these embodiments are clear from this detailed description, and thus the appended claims are intended to cover all such features and advantages that fall within the true spirit and scope of these embodiments. In addition, since many modifications and changes are readily envisioned by those skilled in the art, the embodiments of the present application are not to be limited to the exact structures and operations illustrated and described, but may cover all suitable modifications and equivalents that fall within their scope.

Claims

1. A labeled data generation device, characterized in that, The device includes: An acquisition unit that acquires initial image data; An annotation unit that performs a first annotation on the initial image data to obtain annotation information; A first generation unit that generates a line drawing and a depth map corresponding to the initial image data according to the annotation information; A second generation unit that generates a first condition according to the line drawing, generates a second condition according to the depth map, and generates a third condition according to a prompt; A third generation unit that generates annotation data corresponding to the initial image data according to the first condition, the second condition, and the third condition.

2. The annotation data generation device according to claim 1, wherein The annotation unit performs the first annotation on the initial image data by at least one of a manual annotation method, a semi-automatic annotation method, and an automatic annotation method.

3. The annotation data generation device according to claim 2, wherein The annotation unit performing the first annotation on the initial image data includes: performing at least one of rectangular box annotation, key point annotation, and region annotation on the initial image data; The annotation information includes edge information of the initial image data.

4. The annotation data generation device according to claim 3, wherein The first generation unit generates the line drawing according to the edge information and obtains the depth map through a 2D depth model.

5. The annotation data generation device according to claim 1, wherein The first condition is the edge feature of the line drawing, and the second condition is the depth feature of the depth map.

6. The annotation data generation device according to claim 5, wherein The second generation unit uses an edge detection control network to generate the first condition and uses a depth map control network to generate the second condition; Wherein, the second generation unit extracts the edge feature of the line drawing as the first condition through the edge detection control network, and the second generation unit extracts the depth feature of the depth map as the second condition through the depth map control network.

7. The annotation data generation device according to claim 5, wherein The third condition is a text feature vector generated according to the prompt; The prompt includes a positive prompt and / or a negative prompt.

8. The annotation data generation device according to claim 7, wherein The third generation unit generates the annotation data corresponding to the initial image data according to the edge feature, the depth feature, and the text feature vector.

9. The annotation data generation device according to claim 1, wherein The third generation unit generates the annotation data corresponding to the initial image data through a diffusion generation model.

10. A method for generating labeled data, characterized in that, The method includes: Acquiring initial image data; Performing a first annotation on the initial image data to obtain annotation information; Generating a line drawing and a depth map corresponding to the initial image data according to the annotation information; Generating a first condition according to the line drawing, generating a second condition according to the depth map, and generating a third condition according to a prompt; Generate the annotation data corresponding to the initial image data according to the first condition, the second condition, and the third condition.