Training data generation method and device, storage medium and electronic equipment
By determining the target image and editing instructions in autonomous driving, and using the language model and the target model to generate image samples, the problem of high cost of training data generation in special driving scenarios is solved, high-quality and low-cost training data generation is achieved, and the generalization ability of the model is improved.
Patent Information
- Application Number
- CN202510315138.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
In the field of autonomous driving, especially for special driving scenarios (such as port driving scenarios), the cost of generating large-scale and high-quality training data is high and difficult to complete quickly.
By determining the target image and editing instructions, using the specified language model to generate editing instructions, combining the target model to output image samples, and generating training data required for deep learning tasks in autonomous driving.
This method can effectively reduce the cost of training data generation, improve the quality and coverage of data, and enhance the generalization ability of deep learning models.
Smart Images

Figure CN120147787A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular, to a method, apparatus, storage medium, and electronic device for generating training data. Background Art
[0002] Autopilot technology is a relatively popular cutting-edge artificial intelligence technology today. It uses a deep learning model as the key technology. In the research and development process of realizing autopilot using a deep learning model, large-scale and high-quality training data is required. In the field of autopilot, the so-called training data can be regarded as images of the road conditions ahead of a vehicle driving in a specified scenario taken in the observation direction of the vehicle driver, and is used for training deep learning tasks.
[0003] However, for training data in special driving scenarios (such as port driving scenarios), the cost of large-scale and high-quality training data is relatively high (for example, collecting images related to port driving scenarios on-site at the port requires a huge investment in labor costs and time costs, and it cannot be completed in a short time). How to produce large-scale and high-quality training data for special driving scenarios has become a technical problem that needs to be solved urgently at present. Summary of the Invention
[0004] The present application provides a method, apparatus, storage medium, and electronic device for generating training data, aiming to generate training data for deep learning tasks in autopilot.
[0005] To achieve the above object, the present application provides the following technical solutions:
[0006] A method for generating training data, comprising:
[0007] Determine a target image; the target image represents a target driving scenario corresponding to a target task in autopilot;
[0008] Use a specified language model to obtain an editing instruction for the target driving scenario; the editing instruction includes a description text of a restriction condition; the restriction condition includes a meteorological condition and / or an environmental condition;
[0009] Obtain training data for the target task through a target model corresponding to the target driving scenario; the target model is trained to output an image sample corresponding to the target driving scenario under the meteorological condition and / or environmental condition described in the editing instruction when the target image and the editing instruction are input; the training data includes a plurality of mutually different image samples.
[0010] Optionally, the training process of the target model includes:
[0011] Generate a corpus sample based on the specified language model; the corpus sample includes an input title, a corresponding editing instruction, and an output title; the input title includes an image description related to the target driving scenario; the output title includes a corrected image description obtained by correcting the image description based on the editing instruction;
[0012] Generate an image pair corresponding to the corpus sample through a specified text-to-image model; the image pair includes a first image matching the image description and a second image matching the corrected image description;
[0013] Train a preset diffusion model based on the image pair and the corpus sample as training samples to obtain the target model; wherein, the learning parameters of the diffusion model include meteorological conditions, environmental conditions, and the text complexity of the editing instruction.
[0014] Optionally, the target model includes:
[0015] An image encoding module for compressing an input target image to obtain compressed data and diffusing the compressed data to obtain intermediate features;
[0016] A text encoding module for encoding an input editing instruction to obtain an embedding vector;
[0017] An image information creation module for mapping the embedding vector into a residual network to determine a conditional diffusion network and using the conditional diffusion network to process the intermediate features to obtain an image information matrix;
[0018] The image encoding module is further configured to decode the image information matrix to obtain an image sample output by the target model.
[0019] Optionally, the method further includes:
[0020] Obtain a performance evaluation result obtained after the target task uses the training data for model training;
[0021] Optimize the target model according to the performance evaluation result to obtain an optimized target model;
[0022] Determine a target model corresponding to the target driving scenario based on the optimized target model.
[0023] Optionally, obtaining the training data of the target task through the target model corresponding to the target driving scenario includes:
[0024] If the target driving scenario is a port driving scenario, input the target image and the editing instruction into the target model corresponding to the port driving scenario to obtain an image sample set output by the target model; the image sample set includes multiple image samples; the image content of the image sample is related to the port.
[0025] From the image sample set, determine the image samples whose content similarity meets the first condition as the training data for the target task; the content similarity represents the similarity between the image content of the image sample and the image content of the target image.
[0026] Optionally, the method further includes:
[0027] From the image sample set, determine the image samples whose style similarity meets the second condition as the training data for the target task; the style similarity represents the similarity between the image style of the image sample and the image style of the target image.
[0028] Optionally, obtaining the editing instruction for the target driving scenario by using a specified language model includes:
[0029] Based on the specified language model, generate an editing instruction set; the editing instruction set includes multiple editing instructions; the editing instruction includes a description text whose character length meets the third condition.
[0030] From the editing instruction set, determine the editing instruction whose text semantics meet the fourth condition as the editing instruction for the target driving scenario.
[0031] A training data generation device includes:
[0032] An image determination unit for determining a target image; the target image represents a target driving scenario corresponding to a target task in autonomous driving.
[0033] An instruction generation unit for obtaining an editing instruction for the target driving scenario by using a specified language model; the editing instruction includes a description text of a limiting condition; the limiting condition includes a meteorological condition and / or an environmental condition.
[0034] A data generation unit for obtaining the training data for the target task through the target model corresponding to the target driving scenario; the target model is trained to output an image sample corresponding to the target driving scenario under the meteorological condition and / or environmental condition described by the editing instruction when the target image and the editing instruction are input; the training data includes multiple mutually different image samples.
[0035] Optionally, the process by which the data generation unit trains the target model includes:
[0036] Generate a corpus sample based on the specified language model; the corpus sample includes an input title, a corresponding editing instruction, and an output title; the input title includes an image description related to the target driving scenario; the output title includes a corrected image description obtained by correcting the image description based on the editing instruction.
[0037] Generate an image pair corresponding to the corpus sample through a specified text-to-image model; the image pair includes a first image matching the image description and a second image matching the corrected image description.
[0038] Train a preset diffusion model based on the image pair and the corpus sample as training samples to obtain the target model; wherein, the learning parameters of the diffusion model include meteorological conditions, environmental conditions, and the text complexity of the editing instruction.
[0039] Optionally, the target model includes:
[0040] An image encoding module for compressing an input target image to obtain compressed data and diffusing the compressed data to obtain intermediate features.
[0041] A text encoding module for encoding an input editing instruction to obtain an embedding vector.
[0042] An image information creation module for mapping the embedding vector into a residual network to determine a conditional diffusion network and using the conditional diffusion network to process the intermediate features to obtain an image information matrix.
[0043] The image encoding module is further configured to decode the image information matrix to obtain an image sample output by the target model.
[0044] Optionally, the device further includes:
[0045] A model optimization unit for: obtaining a performance evaluation result obtained after training the model using the training data for the target task; optimizing the target model according to the performance evaluation result to obtain an optimized target model; and determining the target model corresponding to the target driving scenario based on the optimized target model.
[0046] Optionally, the data generation unit is specifically configured to: if the target driving scenario is a port driving scenario, input the target image and the editing instruction into the target model corresponding to the port driving scenario to obtain an image sample set output by the target model; the image sample set includes a plurality of image samples; the image content of the image samples is related to the port; determine, from the image sample set, an image sample whose content similarity meets the first condition as the training data for the target task; the content similarity represents the similarity between the image content of the image sample and the image content of the target image.
[0047] Optionally, the data generation unit is further configured to: determine, from the image sample set, an image sample whose style similarity meets the second condition as the training data for the target task; the style similarity represents the similarity between the image style of the image sample and the image style of the target image.
[0048] Optionally, the instruction generation unit is specifically configured to:
[0049] generate an editing instruction set based on a specified language model; the editing instruction set includes a plurality of editing instructions; the editing instructions include descriptive texts whose character lengths meet the third condition;
[0050] determine, from the editing instruction set, an editing instruction whose text semantics meet the fourth condition as the editing instruction for the target driving scenario.
[0051] A storage medium, the storage medium includes a stored program, wherein the program, when run by a processor, executes the training data generation method described above.
[0052] An electronic device, comprising: a processor, a memory, and a bus; the processor is connected to the memory through the bus;
[0053] The memory is used to store a program, and the processor is used to run the program, wherein the program, when run by the processor, executes any of the training data generation methods described above.
[0054] The technical solution provided by this application determines a target image, obtains an editing instruction for a target driving scenario by using a specified language model, and obtains training data for a target task through a target model corresponding to the target driving scenario. This application uses the image samples output by the target model as training data, which can avoid the huge cost of making training data, reduce the model training cost of the target task in autonomous driving, and the image samples output by the target model can cover different meteorological conditions and environmental conditions, and can provide high-quality training data for the model training of the target task, thereby improving the model generalization ability of the target task. Description of the Drawings
[0055] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0056] Figure 1 A flowchart showing a method for generating training data provided by an embodiment of the present application;
[0057] Figure 2 A flowchart showing another method for generating training data provided by an embodiment of the present application;
[0058] Figure 3 A flowchart showing another method for generating training data provided by an embodiment of the present application;
[0059] Figure 4 A flowchart showing another method for generating training data provided by an embodiment of the present application;
[0060] Figure 5 A schematic diagram of the architecture of a training data generation device provided by an embodiment of the present application;
[0061] Figure 6 A schematic diagram of a target model processing flow provided by an embodiment of the present application;
[0062] Figure 7 A schematic diagram of a method for generating training data provided by an embodiment of the present application. Detailed implementation manners
[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0064] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0065] Embodiment 1
[0066] As Figure 1 shown, it is a schematic flowchart of a training data generation method provided by an embodiment of this application, including the following steps.
[0067] S101: Determine the target image.
[0068] Among them, the target image represents a target driving scenario corresponding to a target task in autonomous driving.
[0069] In some examples, the number of target images can be one or more. For the target image, the target image can be imported by the user according to the target task in autonomous driving.
[0070] In some examples, the types of target tasks include but are not limited to obstacle detection tasks, semantic segmentation tasks, target detection tasks, depth estimation tasks, image segmentation tasks, lane line detection tasks, etc.
[0071] In some examples, the types of target driving scenarios include but are not limited to port driving scenarios, mountain road driving scenarios, off-road driving scenarios, etc.
[0072] S102: Use a specified language model to obtain an editing instruction for the target driving scenario.
[0073] Among them, the editing instruction includes a description text of the restriction conditions, and the restriction conditions include meteorological conditions and / or environmental conditions.
[0074] In some examples, the specified language model includes but is not limited to models such as ChatGPT-3 and ChatGPT-4.
[0075] In a possible implementation, request information related to the limiting conditions of the target driving scenario can be preset or input in ChatGPT-4, so that ChatGPT-4 generates an editing instruction for the target driving scenario according to the request information. For example, after inputting "Create an instruction related to the meteorological conditions and environmental conditions of the target driving scenario" in ChatGPT-4, ChatGPT-4 generates multiple editing instructions related to meteorological conditions and / or environmental conditions.
[0076] In some examples, the so-called meteorological conditions can be used to define the meteorology of the target driving scenario. For example, the description text of the meteorological conditions can be "Remove raindrops", "Snowy day turns to sunny day", "From sunset to night".
[0077] In some examples, the so-called environmental conditions can be used to define the environment of the target driving scenario. For example, the description text of the environmental conditions can be "Urban road turns to highway", "Highway turns to rural road", "Night driving turns to day driving", "Congested traffic turns to smooth traffic", etc.
[0078] It should be noted that using a specified language model, multiple editing instructions for the target driving scenario can be generated. To ensure that relatively reliable training data can be obtained subsequently, the multiple editing instructions generated by the specified language model can be screened to determine the editing instructions that meet the corresponding requirements as the editing instructions for the target driving scenario.
[0079] Optionally, the implementation process of obtaining the editing instruction for the target driving scenario using a specified language model can be as follows: Based on the specified language model, an editing instruction set is generated. The editing instruction set includes multiple editing instructions, and the editing instruction includes a description text whose character length meets the third condition; from the editing instruction set, the editing instruction whose text semantics meet the fourth condition is determined as the editing instruction for the target driving scenario.
[0080] It can be understood that screening the editing instructions whose character length meets the third condition and text semantics meet the fourth condition from the editing instruction set as the editing instructions for the target driving scenario can effectively control the text complexity of the editing instructions, reduce the computational amount of the subsequent mentioned target model, and accelerate the generation efficiency of the training data.
[0081] In some examples, the so-called third condition can be set to that the character length is less than the specified length to avoid the character length of the editing instruction being too long and increasing the text complexity of the editing instruction.
[0082] In a possible implementation, by parsing the description text shown in the editing instruction, the character length of the editing instruction can be obtained. If the character length of the description text shown in the editing instruction is less than the specified length, it is determined that the editing instruction meets the third condition.
[0083] In some instances, the so-called fourth condition can be set such that the text semantics contains specified keywords to control that the description text shown in the editing instruction is all related to meteorological conditions or environmental conditions.
[0084] In possible implementation manners, the number of specified keywords can be multiple. If the description text shown in the editing instruction contains at least one specified keyword, it is determined that the editing instruction meets the fourth condition.
[0085] In some examples, the user can also select, according to the actual situation, an editing instruction that meets the user's needs from the editing instruction set generated by the specified language model.
[0086] S103: Obtain training data for the target task through the target model corresponding to the target driving scenario.
[0087] Among them, the target model is trained to output an image sample corresponding to the target driving scenario under the meteorological conditions and / or environmental conditions described in the editing instruction when the target image and the editing instruction are input. In addition, the training data includes a plurality of mutually different image samples.
[0088] It should be noted that obtaining the training data for the target task through the target model corresponding to the target driving scenario essentially means: inputting the target image and the editing instruction into the target model to obtain a plurality of image samples output by the target model. Using the target model to output the training data does not require investing a large amount of human and time costs in the production of the training data, effectively reducing the production cost of the training data, and can also provide corresponding image samples for the target driving scenario under different limiting conditions, greatly improving the scale and scenario coverage of the training data (the higher the scenario coverage, the higher the quality of the training data), thereby enhancing the generalization ability of the deep learning model in autonomous driving.
[0089] Optionally, the training process of the target model can refer to Figure 2 the steps shown and the explanatory notes of the steps.
[0090] Optionally, the target model includes an image encoding module, a text encoding module, and an image information creation module.
[0091] The image encoding module is used to compress the input target image to obtain compressed data, and diffuse the compressed data to obtain intermediate features.
[0092] In some examples, the image encoding module may adopt a VAE (Variational Auto-encoder). A VAE is a machine learning model used to compress high-dimensional data (i.e., the target image) into a lower-dimensional latent space, that is, to compress the target image to obtain compressed data. In the target model, the VAE is used to compress the target image from the pixel space to a latent space of a smaller dimension while capturing the semantic meaning of the target image. During the forward diffusion process of the target image, Gaussian noise is iteratively applied to the compressed latent representation to simulate the generation process of the target image, and this latent representation can be regarded as intermediate features.
[0093] A text encoding module for encoding the input editing instruction to obtain an embedding vector.
[0094] In some examples, the text encoding module may adopt a CLIP (Contrastive Language-Image Pre-Training) model. Through the CLIP model, the editing instruction can be converted into an embedding vector in the latent space, and this embedding vector represents the learned features of the descriptive text indicated by the editing instruction. In the image editing task, using the CLIP model to convert the editing instruction into an embedding vector can provide specified conditional information for the generation of image samples.
[0095] An image information creation module for mapping the embedding vector into a residual network to determine a conditional diffusion network, and using the conditional diffusion network to process the intermediate features to obtain an image information matrix.
[0096] In some examples, the image information creation module may adopt a U-Net. The so-called U-Net is a deep learning encoder-decoder architecture. In the target model, the main function of the U-Net is to perform a reverse denoising operation on the output of the forward diffusion to obtain the latent representation. The U-Net includes a residual network composed of multiple residual network structures, and the residual network generates the image information matrix in the latent space.
[0097] The image encoding module is also used to decode the image information matrix to obtain the image sample output by the target model.
[0098] In some examples, the image encoding module converts the image information matrix into the pixel space through the decoder in the VAE to obtain the image sample, thereby realizing the decoding process of the image information matrix.
[0099] In a possible implementation manner, for the processing process of the target model on the target image, reference can be made to Figure 6 as shown.
[0100] Optionally, after obtaining the training data of the target task through the target model corresponding to the target driving scenario, the target model can be further optimized to improve the image editing ability of the target model. For the implementation process of optimizing the target model, please refer to Figure 3 the steps shown and the explanatory notes of the steps.
[0101] It can be understood that the quality of the image samples is closely related to the model training of the target task. To improve the model performance of the target task, multiple image samples output by the target model can be further screened to determine higher-quality training data. Optionally, for the implementation process of obtaining the training data of the target task through the target model corresponding to the target driving scenario, please refer to Figure 4 the steps shown and the explanatory notes of the steps.
[0102] For the process shown in S101 - S103 above, combined with Figure 2 and Figure 3 the methods shown, for the process of providing large-scale and high-quality training data for the target task in autonomous driving, please refer to Figure 7 shown, which can be simply summarized as the following steps.
[0103] Step 1: Text triple generation. The text triple includes the input title, the corresponding editing instruction, and the output title, that is, based on the specified language model, generate corpus samples.
[0104] Step 2: Image pair generation, that is, through the specified text-to-image model, generate image pairs corresponding to the corpus samples.
[0105] Step 3: Model construction and parameter loading, that is, load the pre-trained parameters related to the image editing task for the diffusion model.
[0106] In some examples, the pre-trained parameters include the initial learning rate and the training batch size. The initial learning rate can be set to 1×10^(-4), and the training batch size is set to 32.
[0107] Step 4: Learnable parameter setting, that is, determine the learning parameters in the diffusion model.
[0108] Step 5: Model training, that is, based on the image pairs and corpus samples as training samples, train the preset diffusion model to obtain the target model.
[0109] Step 6: Model inference, that is, input the target image and the editing instruction into the target model to obtain the set of image samples output by the target model.
[0110] In some examples, the number of target images is 200, and the number of editing instructions is 300. Input them into the target model, and 1655 image samples output by the target model can be obtained.
[0111] Step 7, downstream task application and effect verification, that is, determining the performance evaluation result of the target task.
[0112] Step 8, performance analysis and optimization, that is, optimizing the target model according to the performance evaluation result of the target task.
[0113] The process shown in the above S101 - S103 uses the image samples output by the target model as training data, which can avoid the huge cost of making training data, reduce the model training cost of the target task in autonomous driving. The image samples output by the target model can cover different meteorological conditions and environmental conditions, and can provide high-quality training data for the model training of the target task, thereby improving the model generalization ability of the target task.
[0114] Embodiment 2
[0115] As Figure 2 shown, it is a schematic flowchart of another training data generation method provided by the embodiment of the present application, including the following steps.
[0116] S201: Generate corpus samples based on a specified language model.
[0117] Among them, the corpus samples include an input title, a corresponding editing instruction, and an output title. The input title includes an image description related to the target driving scenario, and the output title includes a corrected image description obtained by correcting the image description based on the editing instruction.
[0118] In a possible implementation, corresponding corpus request information can be input to the specified language model to enable the specified language model to generate corpus text.
[0119] In some examples, the image description related to the target driving scenario can include various description contents, such as urban roads, highways, night driving, rainy day driving, snow turning to sunny, sunset time, congested traffic, etc.
[0120] In some examples, the corrected image description obtained by correcting the image description based on the editing instruction can be a noun description containing adjectives, such as "clear traffic intersection", "city street illuminated at night".
[0121] In a possible implementation, based on a specified language model, the generated corpus sample can be: {"input_image": "0_clean.png", "edited_image": "0_rain.png", "edit_prompt": "Make the weather gloomy with rain"}. In the shown corpus sample, input_image represents the input caption, edited_image represents the output caption, and edit_prompt represents the editing instruction.
[0122] S202: Generate an image pair corresponding to the corpus sample through a specified text-to-image model.
[0123] Among them, the image pair includes a first image matching the image description and a second image matching the corrected image description.
[0124] In some examples, the specified text-to-image model can be a combination of the Stable Diffusion model and the P2P (Prompt-to-Prompt) model.
[0125] It should be noted that during the process of generating the image pair using the specified text-to-image model, the generated first image and second image can be filtered. According to the directional similarity of the first image and the second image in the CLIP (Clip Space), the consistency between the changes between the first image and the second image and the changes between the input caption and the output caption is measured, so as to obtain the first image matching the image description and the second image matching the corrected image description.
[0126] It can be understood that the changes between the first image and the second image are consistent with the changes between the input caption and the output caption. Therefore, the image content between the first image and the second image is not very different, and the difference between the first image and the second image can be regarded as the differences in various meteorological conditions and / or environmental conditions.
[0127] In a possible implementation, multiple image pairs corresponding to the corpus sample can be generated through the specified text-to-image model, and different corpus samples can also correspond to generate one or more image pairs, so as to obtain the images corresponding to the target driving scenarios under various meteorological conditions and environmental conditions.
[0128] S203: Train a preset diffusion model based on the image pair and the corpus sample as training samples to obtain a target model.
[0129] Among them, the learning parameters of the diffusion model include meteorological conditions, environmental conditions, and the text complexity of the editing instruction.
[0130] In some examples, the diffusion model can adopt a conditional diffusion model architecture suitable for complex image editing. Before training the diffusion model using training samples, pre-trained parameters related to the image editing task can be loaded into the diffusion model so that the diffusion model can effectively adapt to the image editing task.
[0131] In a possible implementation, the preset diffusion model can be the Instruct pix2pix model. Using image pairs and corpus samples as training samples to train the preset diffusion model to obtain the target model can be understood as: fine-tuning the Instruct pix2pix model using the training samples so that the Instruct pix2pix model adapts to complex image editing tasks and eliminates the limitations of the Instruct pix2pix model in the target tasks of autonomous driving.
[0132] The processes shown in S201 - S203 above can achieve the following beneficial effects:
[0133] 1. Performance optimization and task adaptability: By efficiently fine-tuning the diffusion model, the performance of the target model in the target task is significantly improved. This fine-tuning method specifically addresses the defects of the target model in the target task, such as detail loss or unnatural effects in image editing, thus achieving faster and more accurate results.
[0134] 2. Automation and diversity of data generation: The target model trained using the diffusion model automatically generates a large number of high-quality and diverse image samples, reducing the dependence on large-scale manually annotated data. This not only improves the efficiency of data preparation but also enriches the training data, helping the model learn richer image features and styles.
[0135] 3. Improving the generality and flexibility of the target model: By fine-tuning the diffusion model, it can better adapt to various different types of image editing tasks, thus improving the flexibility and generality of the target model in practical applications. This enables the target model to generate image samples matching the target driving scenario according to the specific instructions of the user.
[0136] 4. Accelerating the inference process of the target model: Training the diffusion model using the training samples can optimize the inference process of the target model, enabling the target model to process the target image in a shorter time, thereby improving the efficiency of image editing.
[0137] Embodiment III
[0138] As Figure 3 shown, it is a schematic flowchart of another training data generation method provided by the embodiment of the present application, including the following steps.
[0139] S301: Obtain the performance evaluation result obtained after training the model using the training data for the target task.
[0140] Among them, the target task is usually the training task of a deep learning model. After training the model using the training data of the target task, the obtained performance evaluation result belongs to conventional technical means and will not be elaborated here.
[0141] In some examples, the performance evaluation result of the target task can be determined based on the evaluation metrics of the deep learning model corresponding to the target task, such as the MAP (Mean Average Precision) metric, accuracy metric, recall metric, AP (Average Precision) metric, etc.
[0142] S302: Optimize the target model according to the performance evaluation result to obtain an optimized target model.
[0143] Among them, the process of optimizing the target model according to the performance evaluation result can be: optimizing the pre-training parameters and loss function of the target model according to the performance evaluation result to improve the performance and generalization ability of the target model.
[0144] It can be understood that compared with the target model before optimization, the training data output by the optimized target model better meets the training requirements of the target task, and the quality of the training data is effectively improved.
[0145] S303: Based on the optimized target model, determine the target model corresponding to the target driving scenario.
[0146] Among them, after determining the optimized target model, the optimized target model is determined as the target model corresponding to the target driving scenario.
[0147] The process shown in S301 - S303 above uses the performance evaluation result of the target task to optimize the target model, realizes further fine-tuning of the target model, and improves the quality of the training data output by the target model.
[0148] Example Four
[0149] As Figure 4 shown, it is a schematic flowchart of another training data generation method provided by the embodiments of the present application, including the following steps.
[0150] S401: If the target driving scenario is a port driving scenario, input the target image and the editing instruction into the target model corresponding to the port driving scenario to obtain an image sample set output by the target model.
[0151] Among them, the image sample set includes multiple image samples, and the image content of the image samples is related to ports.
[0152] It should be emphasized that the target model corresponding to the port driving scenario can be obtained by training the diffusion model based on the corpus samples related to the port driving scenario and the corresponding image pairs.
[0153] S402: Determine the image samples whose content similarity meets the first condition from the image sample set as the training data for the target task.
[0154] Among them, the content similarity represents the similarity between the image content of the image sample and the image content of the target image.
[0155] In some examples, the content similarity of each image sample can be determined by a preset similarity algorithm.
[0156] S403: Determine the image samples whose style similarity meets the second condition from the image sample set as the training data for the target task.
[0157] Among them, the style similarity represents the similarity between the image style of the image sample and the image style of the target image.
[0158] In some examples, the style similarity of each image sample can be determined by a preset similarity algorithm.
[0159] The processes shown in S401 - S403 above can use the content similarity and the style similarity as references to screen out higher - quality training data from the image sample set, providing effective help for the deep - learning tasks of autonomous driving.
[0160] Embodiment Five
[0161] Corresponding to the training data generation method provided in the above application, an embodiment of the present application also provides a training data generation device.
[0162] As Figure 5 shown, it is a schematic architecture diagram of a training data generation device provided by an embodiment of the present application, including the following units.
[0163] The image determination unit 100 is used to determine the target image; the target image represents the target driving scenario corresponding to the target task in autonomous driving.
[0164] The instruction generation unit 200 is used to obtain the editing instruction of the target driving scenario by using a specified language model; the editing instruction includes the description text of the limiting conditions; the limiting conditions include meteorological conditions and / or environmental conditions.
[0165] Optionally, the instruction generation unit 200 is specifically configured to: generate an editing instruction set based on a specified language model; the editing instruction set includes a plurality of editing instructions; the editing instructions include descriptive texts with a character length meeting a third condition; and determine, from the editing instruction set, an editing instruction whose text semantics meet a fourth condition as the editing instruction for the target driving scenario.
[0166] The data generation unit 300 is configured to obtain training data for a target task through a target model corresponding to the target driving scenario; the target model is trained to output an image sample corresponding to the target driving scenario under the meteorological conditions and / or environmental conditions described in the editing instruction when an input target image and the editing instruction are provided; the training data includes a plurality of mutually different image samples.
[0167] Optionally, the process by which the data generation unit 300 trains the target model includes: generating a corpus sample based on a specified language model; the corpus sample includes an input title, a corresponding editing instruction, and an output title; the input title includes an image description related to the target driving scenario; the output title includes a corrected image description obtained by correcting the image description based on the editing instruction; generating an image pair corresponding to the corpus sample through a specified text-to-image model; the image pair includes a first image matching the image description and a second image matching the corrected image description; and training a preset diffusion model based on the image pair and the corpus sample as training samples to train the target model; wherein the learning parameters of the diffusion model include meteorological conditions, environmental conditions, and the text complexity of the editing instruction.
[0168] Optionally, the target model includes: an image encoding module configured to compress an input target image to obtain compressed data and perform diffusion on the compressed data to obtain intermediate features; a text encoding module configured to encode an input editing instruction to obtain an embedding vector; an image information creation module configured to map the embedding vector into a residual network to determine a conditional diffusion network and use the conditional diffusion network to process the intermediate features to obtain an image information matrix; and the image encoding module is further configured to decode the image information matrix to obtain the image sample output by the target model.
[0169] Optionally, the data generation unit 300 is specifically configured to: if the target driving scenario is a port driving scenario, input the target image and the editing instruction into the target model corresponding to the port driving scenario to obtain a set of image samples output by the target model; the set of image samples includes a plurality of image samples; the image content of the image samples is related to the port; and determine, from the set of image samples, an image sample whose content similarity meets a first condition as the training data for the target task; the content similarity represents the similarity between the image content of the image sample and the image content of the target image.
[0170] Optionally, the data generation unit 300 is further configured to: determine, from the image sample set, an image sample whose style similarity meets the second condition as the training data for the target task; the style similarity represents the similarity between the image style of the image sample and the image style of the target image.
[0171] The model optimization unit 400 is configured to: obtain the performance evaluation result obtained after the target task is trained using the training data; optimize the target model according to the performance evaluation result to obtain an optimized target model; determine the target model corresponding to the target driving scenario based on the optimized target model.
[0172] Each of the above units uses the target model to output image samples as training data, which can avoid the huge cost of producing training data, reduce the model training cost of the target task in autonomous driving. The image samples output by the target model can cover different meteorological conditions and environmental conditions, and can provide high-quality training data for the model training of the target task, thereby improving the model generalization ability of the target task.
[0173] This application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the training data generation method provided by this application.
[0174] This application also provides an electronic device, including: a processor, a memory, and a bus. The processor is connected to the memory through the bus. The memory is used to store the program, and the processor is used to run the program, wherein the program executes the training data generation method provided by this application when running.
[0175] In addition, at least part of the functions described above in the embodiments of this application can be executed by one or more hardware logic components. For example, without limitation, the exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0176] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0177] Although several specific implementation details are included in the above description, these should not be construed as limiting the scope of the present application. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0178] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A method for generating training data, characterized in that: include: Determine the target image; The target image represents a target driving scene corresponding to a target task in autonomous driving; Using a specified language model, obtaining an editing instruction for the target driving scenario; the editing instruction includes a description text of a restriction; the restriction includes a meteorological condition and / or an environmental condition; The training data of the target task is obtained through the target model corresponding to the target driving scene; the target model is trained to output image samples corresponding to the target driving scene under the meteorological conditions and / or environmental conditions described in the editing instruction when the target image and the editing instruction are input; the training data includes a plurality of different image samples.
2. The method according to claim 1, characterized in that: The training process of the target model includes: Based on the specified language model, a corpus sample is generated; the corpus sample includes an input title, a corresponding editing instruction, and an output title; the input title includes an image description related to the target driving scene; the output title includes a modified image description obtained by modifying the image description based on the editing instruction; By specifying a text graph model, an image pair corresponding to the corpus sample is generated; the image pair includes a first image matching the image description and a second image matching the modified image description; Based on the image pair and the corpus sample as training samples, a preset diffusion model is trained to obtain the target model; wherein the learning parameters of the diffusion model include meteorological conditions, environmental conditions and the text complexity of the editing instruction.
3. The method according to claim 1, characterized in that The target model includes: An image encoding module, used for compressing an input target image to obtain compressed data, and diffusing the compressed data to obtain intermediate features; A text encoding module, used to encode the input editing instructions to obtain an embedding vector; An image information creation module, used for mapping the embedding vector into a residual network to determine a conditional diffusion network, and processing the intermediate features using the conditional diffusion network to obtain an image information matrix; The image encoding module is also used to decode the image information matrix to obtain image samples output by the target model.
4. The method according to claim 1, characterized in that: The method further comprises: Obtaining a performance evaluation result of the target task after model training using the training data; According to the performance evaluation result, the target model is optimized to obtain an optimized target model; Based on the optimized target model, a target model corresponding to the target driving scenario is determined.
5. The method according to claim 1, characterized in that Obtaining training data for the target task through a target model corresponding to the target driving scenario includes: If the target driving scene is a port driving scene, input the target image and the editing instruction into a target model corresponding to the port driving scene to obtain an image sample set output by the target model; the image sample set includes a plurality of image samples; and the image content of the image samples is related to the port; From the image sample set, an image sample whose content similarity meets a first condition is determined as training data for the target task; the content similarity represents the similarity between the image content of the image sample and the image content of the target image.
6. The method according to claim 5, characterized in that The method further comprises: From the image sample set, an image sample whose style similarity meets the second condition is determined as training data for the target task; the style similarity represents the similarity between the image style of the image sample and the image style of the target image.
7. The method according to claim 1, characterized in that Obtaining editing instructions for the target driving scenario using a specified language model, including: Based on the specified language model, an editing instruction set is generated; the editing instruction set includes a plurality of editing instructions; the editing instructions include a description text whose character length meets a third condition; From the editing instruction set, an editing instruction whose text semantics meets the fourth condition is determined as the editing instruction of the target driving scene.
8. A training data generating device, characterized in that: include: An image determination unit, used for determining a target image; The target image represents a target driving scene corresponding to a target task in autonomous driving; An instruction generating unit, configured to obtain an editing instruction of the target driving scenario by using a specified language model; the editing instruction includes a description text of a restriction condition; the restriction condition includes a meteorological condition and / or an environmental condition; A data generation unit is used to obtain training data for the target task through a target model corresponding to the target driving scene; the target model is trained to output image samples corresponding to the target driving scene under the meteorological conditions and / or environmental conditions described in the editing instruction when the target image and the editing instruction are input; and the training data includes a plurality of different image samples.
9. A storage medium, characterized in that: The storage medium includes a stored program, wherein the program, when executed by a processor, executes the training data generating method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: processor, memory, and bus; The processor is connected to the memory via the bus; The memory is used to store programs, and the processor is used to run programs, wherein the program, when run by the processor, executes the training data generating method according to any one of claims 1 to 7.
Citation Information
Cited By
Unstructured road scene data generation and end-to-end sensing method
CN121582894A
Multi-view video data generation method for unstructured mining area scene
CN121904503A