Image data set generation method and device in specific scene, equipment and storage medium
By using large language models and diffusion models to generate image data sets in specific scenarios, the problems of high cost of collecting real data and lack of diversity in images are solved, and efficient and diversified image data set generation is achieved.
Patent Information
- Application Number
- CN202510027513.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-27
AI Technical Summary
When constructing image data sets, the prior art has problems such as high cost and time to acquire real data, lack of realism and diversity in the generated image, and difficulty in obtaining real data in a specific scenario.
By inputting the text information corresponding to a specific scene into the large language model, a prompt word is obtained and the prompt word is input into the diffusion model to generate image information corresponding to the prompt word. Based on image information, multiple visual annotations are generated, and an image data set in a specific scene is constructed based on image information and multiple visual annotations.
It reduces the cost and time of collecting real data, and can effectively generate scarce image data in specific scenarios, thereby enriching the diversity of image data sets and improving the reality and quality of image data sets.
Smart Images

Figure CN120047767A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, apparatus, device, and storage medium for generating an image data set in a specific scenario. Background Art
[0002] With the wide application of machine learning in the field of computer vision, image data has become the basis for training and evaluating computer vision models. In recent years, with the continuous expansion of hardware computing power and the scale of deep learning models, the demand for the quantity of data by the models has also increased significantly. However, in the face of complex large-scale systems and specific tasks, the acquisition of real-world data remains a challenging task.
[0003] In the real world, collecting and establishing a large-scale data set through sensors requires specific channels or a large number of experiments. The accurate annotation of multi-modal sensors and different tasks also requires material and human resources. Moreover, real data sets usually have the long-tail problem, that is, the vast majority of the collected data comes from normal test scenarios, while the frequency of the crucial abnormal scenarios that researchers are concerned about is often relatively low. For some specific fields such as the autonomous driving of unmanned vehicles, it is almost impossible to obtain real data for system test and evaluation because it is difficult to achieve the extreme condition changes required in the real scenario and conduct variable control. At the same time, the test cost of attempting to collect and construct a test data set for the data corresponding to these extreme scenarios in the real physical world is also extremely high. In addition, although existing data augmentation techniques can alleviate the problem of data insufficiency to a certain extent, the generated images often lack realism and cannot fully capture the diversity and complexity of the target objects.
[0004] Therefore, in constructing an image data set, related technologies have problems such as high cost and time for collecting real data, lack of realism and diversity in the generated images, and difficulty in obtaining real data in a specific scenario. Summary of the Invention
[0005] In view of this, the present disclosure provides a method, apparatus, device, and storage medium for generating an image data set in a specific scenario to solve the problems in related technologies, such as high cost and time for collecting real data, lack of realism and diversity in the generated images, and difficulty in obtaining real data in a specific scenario when constructing an image data set.
[0006] In a first aspect, the present disclosure provides a method for generating an image data set in a specific scenario, the method including:
[0007] Inputting the text information corresponding to the specific scenario into a large language model to obtain a prompt;
[0008] Inputting the prompt into a diffusion model to generate image information corresponding to the prompt;
[0009] Generate multiple visual annotations based on the image information, where the visual annotations are used to add label identifications to the image information;
[0010] Construct an image dataset for a specific scenario according to the image information and multiple visual annotations.
[0011] In the embodiments of the present disclosure, by inputting the text information corresponding to a specific scenario into a large language model, a prompt is obtained; the prompt is input into a diffusion model to generate image information corresponding to the prompt; multiple visual annotations are generated based on the image information, where the visual annotations are used to add label identifications to the image information; an image dataset for a specific scenario is constructed according to the image information and multiple visual annotations. Since the embodiments of the present disclosure use a large language model and a diffusion model to generate image information with strong realism in a specific scenario, the cost and time of collecting real data are reduced, and scarce image data in a specific scenario can be effectively generated, thereby enriching the diversity of the image dataset.
[0012] In an alternative embodiment, inputting the prompt into the diffusion model to generate image information corresponding to the prompt includes:
[0013] Input the prompt into the text encoding module of the diffusion model to obtain a text vector;
[0014] Input the text vector into the image generation module of the diffusion model to obtain image information.
[0015] In the embodiments of the present disclosure, by inputting the prompt into the text encoding module of the diffusion model to obtain a text vector, the effective extraction and vectorized representation of the semantic features of the prompt are realized, and the association between text and image can be established. By inputting the text vector into the image generation module of the diffusion model to obtain image information, the conversion from text information to image information is realized, the cost and time of collecting real data can be reduced, and scarce image data in a specific scenario can be effectively generated, thereby enriching the diversity of the image dataset.
[0016] In an alternative embodiment, after inputting the prompt into the diffusion model to generate image information corresponding to the prompt, the method further includes:
[0017] Compare the image information with the prompt;
[0018] In the case of receiving the first feedback information, adjust the parameters of the diffusion model and / or the prompt, where the first feedback information is used to indicate that the image information is inconsistent with the prompt.
[0019] In the embodiments of the present disclosure, by comparing the image information with the prompt and adjusting the parameters of the diffusion model and / or the prompt when receiving the first feedback information, the realism of the image information can be enhanced and the quality of the generated image information can be improved.
[0020] In an alternative embodiment, adjusting the parameters of the diffusion model and / or the prompt includes:
[0021] Adjusting the relevant parameters of the diffusion model until the second feedback information is received, then stopping the adjustment of the relevant parameters to obtain an optimized diffusion model, where the relevant parameters are used to make the diffusion model converge, and the second feedback information is used to indicate that the image information is consistent with the prompt;
[0022] and / or,
[0023] Adjusting the prompt until the second feedback information is received, then stopping the adjustment of the prompt to obtain an optimized prompt.
[0024] In the embodiments of the present disclosure, by adjusting the relevant parameters of the diffusion model to obtain an optimized diffusion model, the quality of the generated image information can be improved. By adjusting the prompt to obtain an optimized prompt, the realism of the image information can be enhanced.
[0025] In an alternative embodiment, constructing an image dataset for a specific scenario according to the image information and multiple visual annotations includes:
[0026] Placing each piece of image information and the corresponding multiple visual annotations in the same folder, where the number of folders is at least one;
[0027] Constructing an image dataset according to the folders.
[0028] In the embodiments of the present disclosure, by placing each piece of image information and the corresponding multiple visual annotations in the same folder and constructing an image dataset according to the folders, an image dataset for a specific scenario is constructed based on the image information and multiple visual annotations.
[0029] In an alternative embodiment, constructing an image dataset for a specific scenario according to the image information and multiple visual annotations includes:
[0030] Placing the image information in an image folder;
[0031] According to the types of the multiple visual annotations, placing the multiple visual annotations in the corresponding annotation folders, where the names of the multiple visual annotations correspond to the name of the image information;
[0032] Constructing an image dataset according to the image folder and the annotation folders.
[0033] In the embodiments of the present disclosure, by placing image information in an image folder, placing multiple visual annotations in annotation folders corresponding to their respective types according to various types of visual annotations, and constructing an image dataset based on the image folder and the annotation folders, an image dataset in a specific scenario is constructed based on the image information and multiple visual annotations.
[0034] In an alternative embodiment, after constructing an image dataset in a specific scenario based on the image information and multiple visual annotations, the method further includes:
[0035] Obtaining request information of a user;
[0036] Matching the types of multiple visual annotations according to the request information to obtain a target visual annotation type;
[0037] Constructing a target image dataset based on the image dataset and the target visual annotation type.
[0038] In the embodiments of the present disclosure, by matching the types of multiple visual annotations according to the request information of the user to obtain a target visual annotation type, and constructing a target image dataset based on the image dataset and the target visual annotation type, a target image dataset can be constructed based on the user's needs, improving the pertinence and usability of the target image dataset.
[0039] In a second aspect, the present disclosure provides an apparatus for generating an image dataset in a specific scenario, the apparatus including:
[0040] A first obtaining module, configured to input text information corresponding to a specific scenario into a large language model to obtain a prompt;
[0041] A first generating module, configured to input the prompt into a diffusion model to generate image information corresponding to the prompt;
[0042] A second generating module, configured to generate multiple visual annotations based on the image information, where the visual annotations are used to add label identifications to the image information;
[0043] A first constructing module, configured to construct an image dataset in a specific scenario based on the image information and multiple visual annotations.
[0044] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method for generating an image dataset in a specific scenario according to the first aspect or any corresponding embodiment thereof.
[0045] Fourthly, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method for generating an image data set in a specific scenario according to the first aspect or any corresponding embodiment thereof.
[0046] Fifthly, the present disclosure provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the method for generating an image data set in a specific scenario according to the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0047] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 is a flowchart of the method for generating an image data set in a specific scenario according to an embodiment of the present disclosure;
[0049] Figure 2a is an image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0050] Figure 2b is a hard-edge annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0051] Figure 2c is a soft-edge annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0052] Figure 2d is a line drawing annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0053] Figure 2e is a normal map annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0054] Figure 2f is a semantic segmentation annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0055] Figure 2g is a depth map annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0056] Figure 2h is a straight line detection annotation of the image in the image data set in a specific scenario according to an embodiment of the present disclosure;
[0057] Figure 3 is a structural block diagram of an image dataset generation device in a specific scenario according to an embodiment of the present disclosure;
[0058] Figure 4 is a schematic hardware structure diagram of a computer device according to an embodiment of the present disclosure. Specific Embodiments
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0060] With the wide application of machine learning in the field of computer vision, image data has become the basis for training and evaluating computer vision models. In recent years, the continuous expansion of hardware computing power and the scale of deep learning models has led to a significant increase in the demand for the quantity of data by the models. However, in the face of complex large-scale systems and specific tasks, the acquisition of real-world data remains a challenging task.
[0061] In the real world, collecting and establishing a large-scale dataset through sensors requires specific channels or a large number of experiments. The accurate annotation of multi-modal sensors and different tasks also requires physical and human resources. Moreover, real datasets usually have the long-tail problem, that is, the vast majority of the collected data comes from normal test scenarios, while the frequency of the crucial abnormal scenarios that researchers are concerned about is often relatively low. For some specific fields such as the autonomous driving of unmanned vehicles, it is almost impossible to obtain real data for system test and evaluation because it is difficult to achieve the extreme condition changes required in the real scenario and conduct variable control. At the same time, the test cost of attempting to collect and construct a test dataset for the data corresponding to these extreme scenarios in the real physical world is also extremely high. In addition, although existing data augmentation technologies can alleviate the problem of insufficient data to a certain extent, the generated images often lack realism and cannot fully capture the diversity and complexity of the target objects.
[0062] Therefore, in constructing an image dataset, related technologies have problems such as high cost and time for collecting real data, lack of realism and diversity in the generated images, and difficulty in obtaining real data in specific scenarios.
[0063] To solve the above problems, according to an embodiment of the present disclosure, there is provided an embodiment of an image dataset generation method in a specific scenario. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0064] In this embodiment, there is provided an image dataset generation method in a specific scenario. As Figure 1 shown, Figure 1 is a flowchart of an image dataset generation method in a specific scenario according to an embodiment of the present disclosure. This process can be applied to a server and includes the following steps:
[0065] Step S101, input the text information corresponding to the specific scenario into the large language model to obtain a prompt.
[0066] Optionally, the embodiment of the present disclosure is applied to the field of autonomous driving. The specific scenario refers to the vehicle driving scenarios under various vehicles, various road scenarios, various weathers, and various perspectives.
[0067] A large language model (LLM) refers to a machine learning model with a large number of parameters and a complex computational structure. A large language model is usually constructed by a deep neural network, with billions or even hundreds of billions of parameters. Its design purpose is to improve the model's expressive ability and prediction performance and handle more complex tasks and data. The large language model learns complex patterns and features by training a large amount of data, can make accurate predictions on unseen data, and is currently widely used in various fields, including natural language processing, computer vision, speech recognition, and recommendation systems, etc.
[0068] A prompt is the text information input when interacting with the model, used to provide input to the model to guide it to generate a specific output. The prompt describes the information such as the answer that one wants to obtain from the model, so as to better control the generated output.
[0069] Specifically, the server obtains the vehicle driving scenarios under various vehicles, various road scenarios, various weathers, and various perspectives, inputs the text information corresponding to these driving scenarios into the large language model, uses a tokenizer to encode the semantic features of the input text information into a digital sequence convenient for the large language model to operate, then maps the digital sequence through an embedding layer to obtain high-dimensional vectors, and then performs complex inference operations on these vectors through a decoder to obtain an index sequence. Finally, the index sequence is restored using the tokenizer to output the prompt.
[0070] Step S102, input the prompt into the diffusion model to generate image information corresponding to the prompt.
[0071] Optionally, in the embodiments of the present disclosure, the prompt words output by the large language model are further expanded descriptions of the text information corresponding to a specific scenario, and are used to guide the diffusion model to generate images of various types of vehicle driving scenarios (such as Figure 2a shown).
[0072] The diffusion model is a deep learning method that, by simulating the physical diffusion process, gradually transforms data into noise, and then learns the reverse process to gradually recover the original data from the noise to achieve high-quality generation effects.
[0073] Specifically, the server selects a diffusion model that has been fine-tuned to generate high-quality and realistic images of specific scenarios, inputs the prompt words output by the large language model into the diffusion model, and through multiple denoising operations, generates image information corresponding to the prompt words.
[0074] Step S103: Generate multiple visual annotations based on the image information, where the visual annotations are used to add label identifications to the image information.
[0075] Optionally, in the embodiments of the present disclosure, the visual annotations include hard edge annotations (such as Figure 2b shown), soft edge annotations (such as Figure 2c shown), line drawing annotations (such as Figure 2d shown), normal map annotations (such as Figure 2e shown), semantic segmentation annotations (such as Figure 2f shown), depth map annotations (such as Figure 2g shown), and straight line detection annotations (such as Figure 2h shown).
[0076] Specifically, for each image, the server uses existing computer vision models in multiple fields to generate multiple visual annotations such as hard edges, soft edges, line drawings, normal maps, semantic segmentation, depth maps, and straight line detections for it.
[0077] In addition, for human images, the server can also choose to annotate human facial expressions, body postures, and hand postures.
[0078] Step S104: Construct an image dataset for a specific scenario according to the image information and multiple visual annotations.
[0079] Optionally, in the embodiments of the present disclosure, the image dataset includes image information and multiple visual annotations.
[0080] Specifically, the server places the image information and multiple visual annotations in the corresponding folders to construct an image dataset for a specific scenario.
[0081] In the embodiments of the present disclosure, by inputting the text information corresponding to a specific scenario into a large language model, a prompt is obtained; the prompt is input into a diffusion model to generate image information corresponding to the prompt; based on the image information, a variety of visual annotations are generated, where the visual annotations are used to add label identifiers to the image information; according to the image information and the variety of visual annotations, an image dataset under a specific scenario is constructed. Since the embodiments of the present disclosure use a large language model and a diffusion model to generate image information with strong realism under a specific scenario, the cost and time of collecting real data are reduced, and scarce image data under a specific scenario can be effectively generated, thereby enriching the diversity of the image dataset.
[0082] In some alternative embodiments, inputting the prompt into a diffusion model to generate image information corresponding to the prompt includes:
[0083] Inputting the prompt into the text encoding module of the diffusion model to obtain a text vector;
[0084] Inputting the text vector into the image generation module of the diffusion model to obtain image information.
[0085] Optionally, in the embodiments of the present disclosure, the text encoding module of the diffusion model employs a contrastive language-image pre-training model (CLIP) based on contrastive learning.
[0086] The CLIP model is a multi-modal pre-trained neural network. The core idea of this model is to use a large amount of paired data of images and texts for pre-training to learn the alignment relationship between images and texts. The architecture of the CLIP model includes two main parts: an image encoder and a text encoder. The image encoder usually uses a convolutional neural network (CNN) to extract the feature representation of the image, while the text encoder uses a recurrent neural network (RNN) or a Transformer to extract the vector representation of the text. The CLIP model jointly trains the image encoder and the text encoder, and the goal is to maximize the similarity between the image and text vectors. Through this training method, the CLIP model can effectively predict a variety of visual tasks without training for a specific task (zero-shot learning).
[0087] Specifically, the server first inputs the prompt output by the large language model into the text encoding module of the diffusion model, encodes the semantic features of the prompt text using the text encoder of the CLIP model to obtain a text vector, and then inputs the text vector into the image generation module of the diffusion model. After multiple denoising operations, image information is finally obtained.
[0088] In the embodiments of the present disclosure, by inputting a prompt into the text encoding module of a diffusion model, a text vector is obtained, which can effectively extract the semantic features of the prompt and represent them in a vectorized manner, and can establish the association between text and images. By inputting the text vector into the image generation module of the diffusion model, image information is obtained, realizing the conversion from text information to image information, which can reduce the cost and time of collecting real data and effectively generate scarce image data in a specific scenario, thus enriching the diversity of the image dataset.
[0089] In some optional embodiments, after inputting the prompt into the diffusion model to generate image information corresponding to the prompt, the method further includes:
[0090] Comparing the image information with the prompt;
[0091] In the case of receiving a first feedback message, adjusting the parameters of the diffusion model and / or the prompt, where the first feedback message is used to indicate that the image information is inconsistent with the prompt.
[0092] Optionally, in the embodiments of the present disclosure, the parameters of the diffusion model include a sampler, the number of sampling iteration steps, the relevance of the prompt, etc.
[0093] Specifically, the server compares the image information with the prompt, and in the case of receiving the first feedback message, that is, when the image information is inconsistent with the prompt, adjusts the parameters of the diffusion model (sampler, the number of sampling iteration steps, the relevance of the prompt, etc.) and / or the prompt.
[0094] In addition, since as the number of parameters increases, the upper limits of the understanding ability of the diffusion model for the prompt and the image generation ability are also higher, so when the computing power permits, the server can also choose to use a diffusion model with a larger number of parameters to generate images.
[0095] In the embodiments of the present disclosure, by comparing the image information with the prompt and adjusting the parameters of the diffusion model and / or the prompt in the case of receiving the first feedback message, the realism of the image information can be enhanced and the quality of the generated image information can be improved.
[0096] In some optional embodiments, adjusting the parameters of the diffusion model and / or the prompt includes:
[0097] Adjusting the relevant parameters of the diffusion model until a second feedback message is received, then stopping adjusting the relevant parameters to obtain an optimized diffusion model, where the relevant parameters are used to make the diffusion model converge, and the second feedback message is used to indicate that the image information is consistent with the prompt;
[0098] and / or,
[0099] Adjust the prompt until the second feedback message is received, then stop adjusting the prompt to obtain the optimized prompt.
[0100] Optionally, in the embodiments of the present disclosure, the server adjusts the relevant parameters of the diffusion model (sampler, number of sampling iteration steps, prompt relevance, etc.) until the second feedback message is received, that is, when the image information is consistent with the prompt, stop adjusting the relevant parameters to obtain the optimized diffusion model. The server can also adjust the prompt until the second feedback message is received, that is, when the image information is consistent with the prompt, stop adjusting the prompt to obtain the optimized prompt.
[0101] In the embodiments of the present disclosure, by adjusting the relevant parameters of the diffusion model to obtain the optimized diffusion model, the quality of the generated image information can be improved. By adjusting the prompt to obtain the optimized prompt, the realism of the image information can be enhanced.
[0102] In some alternative embodiments, an image dataset for a specific scenario is constructed based on the image information and multiple visual annotations, including:
[0103] Place each piece of image information and its corresponding multiple visual annotations in the same folder, where the number of folders is at least one;
[0104] Construct an image dataset based on the folders.
[0105] Optionally, in the embodiments of the present disclosure, the server can place each piece of image information and its corresponding multiple visual annotations in the same folder, and different image information and their corresponding multiple visual annotations in different folders, and construct an image dataset based on the folders.
[0106] In the embodiments of the present disclosure, by placing each piece of image information and its corresponding multiple visual annotations in the same folder and constructing an image dataset based on the folders, an image dataset for a specific scenario is constructed based on the image information and multiple visual annotations.
[0107] In some alternative embodiments, an image dataset for a specific scenario is constructed based on the image information and multiple visual annotations, including:
[0108] Place the image information in an image folder;
[0109] According to the types of multiple visual annotations, place the multiple visual annotations in the corresponding annotation folders, where the names of the multiple visual annotations correspond to the names of the image information;
[0110] Construct an image dataset based on the image folder and the annotation folders.
[0111] Optionally, in the embodiments of the present disclosure, the server may place the image information in an image folder (such as the images folder), and at the same time, according to the types of multiple visual annotations, place the multiple visual annotations in the corresponding annotation folders of the types (such as placing all semantic segmentation annotations in the segment folder), and construct an image dataset based on the image folder and the annotation folders.
[0112] It should be noted that in order to ensure that the corresponding annotation file can be indexed according to the image name, the names of the multiple visual annotations should correspond to the name of the image information.
[0113] In the embodiments of the present disclosure, by placing the image information in the image folder, placing the multiple visual annotations in the corresponding annotation folders according to the types of the multiple visual annotations, and constructing the image dataset based on the image folder and the annotation folders, an image dataset in a specific scenario is constructed based on the image information and the multiple visual annotations.
[0114] In some alternative embodiments, after constructing the image dataset in a specific scenario according to the image information and the multiple visual annotations, the method further includes:
[0115] Obtaining the request information of the user;
[0116] Matching the types of the multiple visual annotations according to the request information to obtain the target visual annotation type;
[0117] Constructing a target image dataset according to the image dataset and the target visual annotation type.
[0118] Optionally, in the embodiments of the present disclosure, the request information of the user includes the image dataset information required for the computer vision task to be performed by the user, the target visual annotation type refers to the visual annotation type required by the user, and the target image dataset refers to the image dataset required by the user.
[0119] Specifically, after the server constructs the image dataset in a specific scenario, it obtains the request information of the user, matches the types of the multiple visual annotations according to the image dataset information required for the computer vision task to be performed by the user in the request information to obtain the target visual annotation type, and finally constructs the image dataset required by the user according to the image dataset and the target visual annotation type to obtain the target image dataset.
[0120] In the embodiments of the present disclosure, by matching the types of the multiple visual annotations according to the request information of the user to obtain the target visual annotation type, and constructing the target image dataset according to the image dataset and the target visual annotation type, a target image dataset can be constructed based on the user's needs, and the pertinence and usability of the target image dataset are improved.
[0121] In this embodiment, an image dataset generation device in a specific scenario is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0122] This embodiment provides an image dataset generation device in a specific scenario. As Figure 3 shown, it includes:
[0123] A first obtaining module 301, configured to input the text information corresponding to the specific scenario into a large language model to obtain a prompt;
[0124] A first generating module 302, configured to input the prompt into a diffusion model to generate image information corresponding to the prompt;
[0125] A second generating module 303, configured to generate multiple visual annotations based on the image information, where the visual annotations are used to add label identifiers to the image information;
[0126] A first constructing module 304, configured to construct an image dataset in the specific scenario according to the image information and multiple visual annotations.
[0127] In the embodiments of the present disclosure, by inputting the text information corresponding to the specific scenario into a large language model to obtain a prompt, inputting the prompt into a diffusion model to generate image information corresponding to the prompt, generating multiple visual annotations based on the image information, where the visual annotations are used to add label identifiers to the image information, and constructing an image dataset in the specific scenario according to the image information and multiple visual annotations. Since the embodiments of the present disclosure use a large language model and a diffusion model to generate image information with strong realism in a specific scenario, the cost and time of collecting real data are reduced, and scarce image data in a specific scenario can be effectively generated, thereby enriching the diversity of the image dataset.
[0128] In some alternative implementation manners, the first generating module 302 includes:
[0129] A first obtaining sub-module, configured to input the prompt into the text encoding module of the diffusion model to obtain a text vector;
[0130] A second obtaining sub-module, configured to input the text vector into the image generation module of the diffusion model to obtain image information.
[0131] In some alternative implementation manners, the device further includes:
[0132] A comparison module, configured to compare the image information with the prompt;
[0133] An adjustment module, configured to adjust the parameters and / or prompt words of the diffusion model when receiving a first feedback message, where the first feedback message is used to indicate that the image information is inconsistent with the prompt words.
[0134] In some optional embodiments, the adjustment module includes:
[0135] A third obtaining sub-module, configured to adjust the relevant parameters of the diffusion model until receiving a second feedback message, stop adjusting the relevant parameters, and obtain an optimized diffusion model, where the relevant parameters are used to make the diffusion model converge, and the second feedback message is used to indicate that the image information is consistent with the prompt words;
[0136] A fourth obtaining sub-module, configured to adjust the prompt words until receiving a second feedback message, stop adjusting the prompt words, and obtain optimized prompt words.
[0137] In some optional embodiments, the first construction module 304 includes:
[0138] A first placing sub-module, configured to place each piece of image information and corresponding multiple visual annotations in the same folder, where the number of folders is at least one;
[0139] A first construction sub-module, configured to construct an image data set according to the folders.
[0140] In some optional embodiments, the first construction module 304 includes:
[0141] A second placing sub-module, configured to place the image information in an image folder;
[0142] A third placing sub-module, configured to place multiple visual annotations in corresponding annotation folders according to the types of multiple visual annotations, where the names of the multiple visual annotations correspond to the names of the image information;
[0143] A second construction sub-module, configured to construct an image data set according to the image folder and the annotation folders.
[0144] In some optional embodiments, the apparatus further includes:
[0145] An obtaining module, configured to obtain the request information of the user;
[0146] A second obtaining module, configured to match the types of multiple visual annotations according to the request information to obtain a target visual annotation type;
[0147] A second construction module, configured to construct a target image data set according to the image data set and the target visual annotation type.
[0148] The further function descriptions of the above modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0149] In this embodiment, the image dataset generation device in a specific scenario is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0150] The embodiments of the present disclosure also provide a computer device having the above Figure 3 shown image dataset generation device in a specific scenario.
[0151] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure. As Figure 4 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 4 In
[0152] FIG. 1, one processor 10 is taken as an example.
[0153] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0154] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device and the like. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely disposed relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0155] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.
[0156] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0157] The embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and to be stored in a local storage medium, so that the methods described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0158] A part of the present disclosure can be applied as a computer program product, for example, computer program instructions, which when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present disclosure through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for a computer to execute computer program instructions include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0159] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for generating an image dataset in a specific scenario, characterized in that: The method comprises: Input text information corresponding to a specific scene into a large language model to obtain prompt words; Inputting the prompt word into a diffusion model to generate image information corresponding to the prompt word; Based on the image information, generating a plurality of visual annotations, wherein the visual annotations are used to add label identifications to the image information; An image dataset for the specific scene is constructed according to the image information and the multiple visual annotations.
2. The method according to claim 1, characterized in that The step of inputting the prompt word into a diffusion model to generate image information corresponding to the prompt word includes: Inputting the prompt word into the text encoding module of the diffusion model to obtain a text vector; The text vector is input into the image generation module of the diffusion model to obtain the image information.
3. The method according to claim 2, characterized in that After inputting the prompt word into the diffusion model to generate image information corresponding to the prompt word, the method further includes: comparing the image information with the prompt word; In case of receiving first feedback information, adjusting the parameters of the diffusion model and / or the prompt word, wherein the first feedback information is used to indicate that the image information is inconsistent with the prompt word.
4. The method according to claim 3, characterized in that The adjusting the parameters of the diffusion model and / or the prompt word comprises: Adjusting relevant parameters of the diffusion model until receiving second feedback information, stopping adjusting the relevant parameters, and obtaining an optimized diffusion model, wherein the relevant parameters are used to make the diffusion model converge, and the second feedback information is used to indicate that the image information is consistent with the prompt word; and / or, The prompt word is adjusted until the second feedback information is received, and the adjustment of the prompt word is stopped to obtain an optimized prompt word.
5. The method according to claim 1, characterized in that The step of constructing the image dataset in the specific scene according to the image information and the multiple visual annotations includes: placing each piece of the image information and the corresponding multiple visual annotations in the same folder, wherein the number of the folders is at least one; The image dataset is constructed according to the folder.
6. The method according to claim 1, characterized in that The step of constructing the image dataset in the specific scene according to the image information and the multiple visual annotations includes: Placing the image information in an image folder; According to the types of the multiple visual annotations, the multiple visual annotations are placed in annotation folders of corresponding types, wherein the names of the multiple visual annotations correspond to the names of the image information; The image dataset is constructed according to the image folder and the annotation folder.
7. The method according to any one of claims 1 to 6, characterized in that: After constructing the image dataset in the specific scene according to the image information and the multiple visual annotations, the method further includes: Get the user's request information; Matching the multiple visual annotation types according to the request information to obtain a target visual annotation type; A target image dataset is constructed according to the image dataset and the target visual annotation type.
8. A device for generating an image data set in a specific scenario, characterized in that: The device comprises: The first obtaining module is used to input text information corresponding to a specific scene into a large language model to obtain a prompt word; A first generating module, used for inputting the prompt word into a diffusion model to generate image information corresponding to the prompt word; A second generating module is used to generate a plurality of visual annotations based on the image information, wherein the visual annotations are used to add label identifications to the image information; The first construction module is used to construct an image dataset in the specific scene according to the image information and the multiple visual annotations.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for generating an image data set in a specific scenario according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for generating an image data set in a specific scenario according to any one of claims 1 to 7.
11. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to enable a computer to execute the method for generating an image data set in a specific scenario according to any one of claims 1 to 7.