Data generation method, model training method, device, equipment, medium and chip
By using simulation data and mining data generation models to generate high-quality target mining area data in mining area autonomous driving, the problems of high cost and low quality of real environment data acquisition are solved, and safe and efficient data generation and model training are achieved.
Patent Information
- Application Number
- CN202510521223.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art relies on real environmental data collection in mining areas for autonomous driving, resulting in high costs and safety hazards, low data quality, and prone to errors and inconsistencies in the labeling results.
By obtaining simulated mining area data, extracting control condition data, and using the trained mining area data generation model to generate target mining area data, combined with random noise data, high-quality data for autonomous driving in mining area are generated.
It reduces data acquisition costs, improves data quality, reduces domain gaps, and ensures the stability and safety of the autonomous driving perception model in the mining area.
Smart Images

Figure CN120030626B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving and image processing in mining areas, and in particular to a mining area data generation method, a mining area data generation model training method, a mining area data generation device, a mining area data generation model training device, computer equipment, a readable storage medium and a chip. Background Art
[0002] The core of autonomous driving systems lies in the acquisition and labeling of high-quality perception data. Related technologies deploy multimodal sensors in real-world environments to collect scene information, and then manually or semi-automatically label task labels such as object detection and semantic segmentation. This approach has the following drawbacks:
[0003] (1) Data collection is completely dependent on the real environment, and the data collection cost is high. Especially in special scenarios such as mines, the data collection cost increases significantly. At the same time, there are many environmental risk factors in mining areas, and data on-site collection is limited, the collection process is complicated, and there are safety hazards.
[0004] (2) Real data requires a lot of manual work or auxiliary tools to label, which consumes a lot of resources. In addition, the labeling results are prone to errors and inconsistencies, which reduces data quality. Summary of the Invention
[0005] In view of this, the present application provides a mining area data generation method, a mining area data generation model training method, a mining area data generation device, a mining area data generation model training device, computer equipment, a readable storage medium and a chip, which solve the problems of high data collection cost and low data quality.
[0006] In a first aspect, an embodiment of the present application provides a method for generating mining area data, comprising:
[0007] Acquire simulated mining area data for a mining area scene, the simulated mining area data including a virtual scene image and label information of the virtual scene image, and extract first control condition data of the virtual scene image;
[0008] Random noise data is sampled, and the first control condition data and the random noise data are input into a trained mining area data generation model to obtain target mining area data; wherein, the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.
[0009] In a second aspect, an embodiment of the present application provides a mining area data generation model training method, comprising:
[0010] Acquire real mining area data in a mining area scenario, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image;
[0011] Performing noise processing on the real scene image to obtain noise data;
[0012] Based on the second control condition data and the noise data, a mining area data generation model is trained. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.
[0013] In a third aspect, an embodiment of the present application provides a mining area data generation device, comprising:
[0014] A first data acquisition module is configured to acquire simulated mining area data for a mining area scene, wherein the simulated mining area data includes a virtual scene image and label information of the virtual scene image;
[0015] A first control condition extraction module, configured to extract first control condition data of the virtual scene image;
[0016] Noise data sampling module, used to sample random noise data;
[0017] A data generation module is used to input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein, the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of the vehicle in the mining area scene.
[0018] In a fourth aspect, an embodiment of the present application provides a mining area data generation model training device, comprising:
[0019] A second data acquisition module is used to acquire real mining area data in a mining area scene, wherein the real mining area data includes a real scene image and label information of the real scene image;
[0020] A second control condition extraction module, configured to extract second control condition data of the real scene image;
[0021] A data noise adding module, configured to perform noise adding processing on the real scene image to obtain noise data;
[0022] A model training module is used to train a mining area data generation model based on the second control condition data and the noise data. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.
[0023] In a fifth aspect, an embodiment of the present application provides a computer device, which includes a first processor and a first memory, wherein the first memory stores programs or instructions that can be run on the first processor, and when the programs or instructions are executed by the first processor, the steps of the method of the first aspect or the second aspect are implemented.
[0024] In a sixth aspect, an embodiment of the present application provides a readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, the steps of the method of the first aspect or the second aspect are implemented.
[0025] In the seventh aspect, an embodiment of the present application provides a chip, which includes at least one second processor and a communication interface, the communication interface and the at least one second processor are coupled, and the at least one second processor is used to run programs or instructions to implement the steps of the method of the first aspect or the second aspect.
[0026] The beneficial effects of the present application are as follows: the mining area data generation method, mining area data generation model training method, mining area data generation device, mining area data generation model training device, computer equipment, readable storage medium and chip provided in the embodiments of the present application obtain simulated mining area data for mining area scenarios, extract control condition data from the simulated mining area data, thereby further optimizing the simulated mining area data, and then use the control condition data as a guide to generate target mining area data for controlling mining area autonomous driving through the trained mining area data generation model. The present application does not need to rely on real environment data collection and does not need to perform a large amount of data labeling work. It can generate a large amount of mining area autonomous driving data based on simulated mining area data and trained mining area data generation models. And through the guidance of control conditions, the data authenticity is effectively improved, the domain gap between the simulated data and the real mining area scenario is significantly reduced, the quality of the generated data is improved, and the stability of the training effect of the mining area autonomous driving perception model is ensured.
[0027] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0029] Figure 1 A schematic diagram showing a process of a method for generating mining area data according to an embodiment of the present application is shown;
[0030] Figure 2 A schematic diagram showing the construction of simulated mining area data in an embodiment of the present application is shown;
[0031] Figure 3 A logical diagram of a method for generating mining area data according to an embodiment of the present application is shown;
[0032] Figure 4 A schematic diagram showing a comparison of the effects of the embodiment of the present application and the related art;
[0033] Figure 5 A schematic diagram showing a process flow of a mining area data generation model training method according to an embodiment of the present application is shown;
[0034] Figure 6 A schematic diagram showing the mining area data generation model training of an embodiment of the present application is shown;
[0035] Figure 7 A schematic diagram showing the dimensionality reduction process of an embodiment of the present application is shown;
[0036] Figure 8 A schematic diagram showing a preset model of an embodiment of the present application is shown;
[0037] Figure 9 A screenshot of pseudo code for dynamic foreground weight calculation in an embodiment of the present application is shown;
[0038] Figure 10 A schematic diagram showing a dynamic weight scheduling curve according to an embodiment of the present application is shown;
[0039] Figure 11 A structural block diagram of a mining area data generating device according to an embodiment of the present application is shown;
[0040] Figure 12 The following is a structural block diagram of a mining area data generation model training device according to an embodiment of the present application;
[0041] Figure 13 A structural block diagram of a computer device according to an embodiment of the present application is shown;
[0042] Figure 14 A structural block diagram of a readable storage medium according to an embodiment of the present application is shown;
[0043] Figure 15The structure block diagram of the chip of the embodiment of the present application is shown. DETAILED DESCRIPTION
[0044] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0045] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0046] Below, in conjunction with the accompanying drawings, the mining area data generation method, mining area data generation model training method, mining area data generation device, mining area data generation model training device, computer equipment, readable storage medium and chip provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios. Unless there is a conflict, the following embodiments and features in the embodiments can be combined with each other.
[0047] The present application embodiment provides a method for generating mining area data, such as Figure 1 As shown, the method includes:
[0048] Step 101: Acquire simulated mining area data for a mining area scene, the simulated mining area data including a virtual scene image and label information of the virtual scene image, and extract first control condition data of the virtual scene image.
[0049] In this step, simulated mining area data for the mining area scene is obtained. The simulated mining area data includes a virtual scene image and label information of the virtual scene image. The label information includes a category label and detection frame geometry information. The simulated mining area data is used to subsequently generate target mining area data for controlling the automatic driving of vehicles in the mining area scene.
[0050] In one embodiment of the present application, obtaining simulated mining area data for a mining area scenario includes:
[0051] Construct a virtual environment of the mining area based on the collected data and digital environment model of the mining area;
[0052] Carry out vehicle dynamics simulation and sensor simulation in the virtual environment of the mining area to obtain a virtual vehicle moving in the virtual environment of the mining area and sensors running on the virtual vehicle;
[0053] Based on the virtual environment of the mining area, a variety of mining scenarios are constructed, and simulated mining data collected by sensors of virtual vehicles in the mining area scenarios are obtained. The mining area scenarios must include at least one of the following: long-tail scenarios, complex working condition scenarios, and extreme climate scenarios.
[0054] In this embodiment, Figure 2 As shown in the figure, the simulation engine (Unity) is used to construct the simulated mining area data. The specific implementation steps include:
[0055] (1) Through mining area data collection and digital environment model, a virtual environment of the mining area scene is built, and the collected data is post-processed and visually rendered to construct a realistic mining area virtual environment.
[0056] (2) Carry out vehicle dynamic simulation and sensor simulation. Sensors include lidar, visual cameras, etc. to realize the movement of virtual vehicles in mining scenes and the operation of sensors on virtual vehicles.
[0057] (3) Obtain scenario parameter configurations, construct various mining scenarios such as long-tail scenarios, complex working condition scenarios, and extreme climate scenarios, and perform mining operation scheduling and management, including vehicle loading, transportation, and unloading. Long-tail scenarios include special obstacles or abnormal events, and extreme climates include sandstorms and heavy rain. Furthermore, obtain virtual scene images output by sensors in real time, and automatically annotate the virtual scene images with the help of automated annotation tools to obtain label information for the virtual scene images.
[0058] Through the above methods, accurate simulation of the dynamic work of the virtual mining area is achieved, providing a data source for the dynamic environment of complex mining areas, and ensuring the authenticity and diversity of the generated data.
[0059] The real data collection in related technologies is costly and time-consuming, especially in dangerous areas of mines, where data acquisition is difficult and presents safety risks. The present application embodiment uses simulation-driven data generation technology, with the help of virtual sensors and automated annotation tools, to efficiently generate high-precision simulation data and precise annotations in batches. This serves as input for generating target mining data used to control autonomous driving of vehicles in mining scenarios, thereby avoiding the high cost of data collection and manual annotation in real scenarios and significantly improving the efficiency and safety of data generation.
[0060] It is worth noting that through flexible scenario parameter configuration, this application can effectively simulate the complex terrain, extreme weather and long-tail scenarios unique to mining areas in a simulation environment, so that the autonomous driving dataset covers more diverse boundary cases and improves the robustness and generalization ability of the autonomous driving perception model in complex scenarios.
[0061] Furthermore, after constructing the simulated mining area data, in order to improve the authenticity and generalization ability of the simulated mining area data, control condition data is extracted from the simulated mining area data, that is, the first control condition data of the virtual scene image is extracted as input for the subsequent mining area data generation model.
[0062] In one embodiment of the present application, extracting first control condition data of a virtual scene image includes:
[0063] Converting the label information of the virtual scene image into a first text description feature, and performing feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature;
[0064] Performing position embedding processing on the spatial information of the virtual scene image to obtain a first spatial vector, where the spatial information includes resolution information and position information;
[0065] Performing encoding processing and feature fusion processing on the segmentation mask data of the virtual scene image to obtain a first fusion feature;
[0066] First control condition data is obtained according to the first global semantic feature, the first spatial vector and the first fusion feature, and the first control condition data is mapped to a latent space corresponding to the mining area data generation model.
[0067] In this embodiment, Figure 3 As shown, the label information of the virtual scene image is converted into a natural language description, namely the first text description feature, through prompt engineering technology. Feature extraction and further pooling are performed on the first text description feature to obtain the first global semantic feature. At the same time, the resolution and position information of the virtual scene image are introduced to obtain the first spatial vector of the virtual scene image to supplement the spatial information of the image, ensuring that the subsequent mining area data generation model can accurately locate the target object in the virtual scene image. In addition, the segmentation mask data of the virtual scene image undergoes encoding and feature fusion processing to provide structured prior information to assist the mining area data generation model in image generation.
[0068] The first control condition data includes the first global semantic feature, the first spatial vector and the first fusion feature. Finally, the first control condition data is mapped to the latent space corresponding to the mining area data generation model as the input of the mining area data generation model for the mining area data generation model to generate data.
[0069] Through the above method, the control conditions are extracted to achieve data enhancement.
[0070] In one embodiment of the present application, the label information includes a category label and detection frame geometry information; converting the label information of the virtual scene image into a first text description feature includes:
[0071] The category label and detection frame geometry information of the virtual scene image are converted into a first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label.
[0072] In this embodiment, the prompt engineering technology converts the category labels and detection box geometry information of the virtual scene image into natural language descriptions through the prompt engineering method, which can be used by the mining area data generation model to utilize the multimodal information of the data.
[0073] In target detection tasks in related technologies, the annotations of target objects are usually stored in the following format: {"ID":"truck","bbox":[x, y, w, h]}. Among them, ID represents the category of the detection box of the target object, for example, "truck" represents a mining truck, and "bbox" represents the detection box (bounding_box), which is a list of four values, representing the pixel coordinates (x, y) of the center point of the detection box and the width (w) and height (h) of the detection box. However, this format only provides geometric information of the target object and it is difficult to express higher-level semantic relationships, such as the color, state, background environment, etc. of the target object. In this application, if Figure 3 As shown in the figure, the hint engineering converts the target object's annotation information into a natural language description, that is, the first text description feature. For example: In the image, a large yellow dump truck is positioned at coordinates (x, y), with its body partially obscured by airborne dust. The vehicle spans w pixels in width and h pixels in height. This conversion method enhances the expressiveness of the target object and enables the mining area data generation model to understand more complex visual relationships.
[0074] In one embodiment of the present application, feature extraction and pooling processing are performed on the first text description feature to obtain a first global semantic feature, including:
[0075] The first text description feature is segmented by a first CLIP segmenter to obtain a first target text description feature, and the first target text description feature is encoded by a first CLIP encoder to obtain a first deep feature;
[0076] The first text description feature is segmented by a second CLIP segmenter to obtain a second target text description feature, and the second target text description feature is encoded by a second CLIP encoder to obtain a second deep feature;
[0077] Performing feature splicing processing on the first deep feature and the second deep feature to obtain a first spliced feature;
[0078] Performing pooling processing on the second target text description feature to obtain the first global feature;
[0079] Obtaining a first global semantic feature according to the first concatenated feature and the first global feature;
[0080] The word segmentation process includes: splitting the first text description feature into subword units, adding text tags to the subword units, and performing length unification processing on the subword units after adding text tags to obtain target text description features with unified length. Specifically, when the first text description feature is segmented using the first CLIP word segmenter, a first target text description feature with unified length is obtained; when the first text description feature is segmented using the second CLIP word segmenter, a second target text description feature with unified length is obtained.
[0081] In this embodiment, in order to enhance the ability of the mining area data generation model to understand text control signals, CLIP (Contrastive Language-Image Pre-training, based on the contrastive language-image pre-training model) is used to construct a text encoder to map the natural language description to the latent space of the mining area data generation model, where the mining area data generation model is constructed based on the principle of the diffusion model, thereby improving the text's ability to guide the generated content.
[0082] The CLIP model consists of a CLIP tokenizer and a CLIP encoder, which uses a Transformer architecture. In the generation task, the first text description feature is input into the CLIP processing pipeline in natural language form. The input text is normalized to ensure format consistency. Subsequently, the WordPiece tokenization method is used to break the text into smaller sub-word units, adapting them to the CLIP pre-trained vocabulary, which consists of (49,408) tokens (features). Special markers, such as [SOS], [EOS], and [PAD], are added to indicate the start and end of the text and padding content. Finally, a padding or truncation strategy is used to unify the length of the sub-word units after adding text markers, bringing the text length to the maximum number of features, which can be 77 tokens. This results in the target text description feature, ensuring that it meets the input specifications of the CLIP encoding process. This concludes the tokenization phase. The encoding phase then proceeds, using the Transformer architecture to extract deep features from the text, which serve as the text conditional input for the mining area data generation model.
[0083] In one embodiment, Figure 3 As shown in the figure, the first CLIP word segmenter and the first CLIP encoder act on the first text description feature to obtain the first deep feature s1; the second CLIP word segmenter and the second CLIP encoder act on the first text description feature to obtain the second deep feature s2. Furthermore, the first deep feature s1 and the second deep feature s2 are concatenated to obtain the first concatenated feature s3.
[0084] In addition, in order to further improve the mining area data generation model’s ability to perceive the overall semantics of the text, the hidden state of the first token is extracted and pooled into a global vector to represent the overall semantic features of the input text. Figure 3 As shown, the second target text description feature is pooled to obtain the first global feature s4. The first concatenated feature s3 and the first global feature s4 are then fused into the first global semantic feature s5.
[0085] In this embodiment, two different CLIP models are used to extract text features and then pool them to obtain global semantic features. This dual encoding strategy can enhance the model's understanding of text semantics and improve the expression of complex scenes. It can also combine multimodal alignment information between text and image, thereby improving the quality and consistency of the synthesized image.
[0086] Step 102: sample random noise data, and input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein, the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.
[0087] In this embodiment, random noise data of the same size as the scene image is sampled from the normally distributed data. The first control condition data and the random noise data are then input into a trained mining area data generation model. This outputs target mining area data with a realistic mining area style. This target mining area data is used to expand the mining area autonomous driving dataset and train the perception model. In some embodiments, the quality of the simulated mining area data generation can be continuously improved based on the effectiveness of the perception model.
[0088] In one embodiment, the mining area data generation model is constructed based on the principle of the stable diffusion model.
[0089] In an embodiment of the present application, simulated mining area data for a mining area scenario is obtained, and control condition data is extracted from the simulated mining area data, thereby further optimizing the simulated mining area data. The control condition data is then used as a guide to generate target mining area data for controlling autonomous driving in the mining area through a trained mining area data generation model. This application does not need to rely on real-world environment data collection and does not require a large amount of data labeling work. It can generate a large amount of mining area autonomous driving data based on simulated mining area data and trained mining area data generation models. In addition, the guidance of control conditions effectively improves the authenticity of the data, significantly reduces the domain gap between the simulated data and the real mining area scenario, improves the quality of the generated data, and ensures the stability of the training effect of the mining area autonomous driving perception model.
[0090] In one embodiment of the present application, Figure 3 As shown, the mining area data generation model includes an original neural network, a conditional control network and a connection network. The connection network is a zero convolution layer. The original neural network and the conditional control network are connected through the connection network. The original neural network includes multiple sub-networks connected in sequence.
[0091] In one embodiment of the present application, the first control condition data includes: a first global semantic feature, a first spatial vector, and a first fusion feature; the first control condition data and random noise data are input into a trained mining area data generation model to obtain target mining area data, including:
[0092] The first control condition data and random noise data are input into the mining area data generation model, and the target mining area data is output; wherein, the first control condition data and random noise data are input into the mining area data generation model, including: inputting the random noise data into the first sub-network of the original neural network, and fusing the random noise data and the first fusion feature and inputting them into the conditional control network, and inputting the first global semantic feature and the first space vector into each sub-network of the conditional control network and the original neural network, respectively.
[0093] In this embodiment, random noise data is input into the first subnetwork of the original neural network, and the random noise data and the first fusion feature are fused and input into the conditional control network, and the first global semantic feature and the first spatial vector are respectively input into the conditional control network and each subnetwork of the original neural network. Finally, the mining area data generation model outputs target mining area data with higher authenticity. In this embodiment of the application, multiple conditional information such as text description, segmentation mask, resolution information, and location information of the simulated mining area data is used to guide the mining area data generation model to generate high-quality images of the mining area in a realistic style.
[0094] The following combination Figure 4 The following is a comparison of the effects of the embodiments of the present application and related technologies: Figure 4 (a) in the figure is the original mining scene data generated by the simulation platform (Unity). Although it can simulate the typical scene and layout of the mining area, it lacks visual details and there is an obvious visual gap between it and the real scene. Figure 4 (b) in the figure is the semantic segmentation label data of the corresponding scene. These labels are automatically generated and can effectively support the training of subsequent perception models. Figure 4 (c) and (d) are the optimized mining scene data proposed in this application. (c) and (d) are the effects of two different parameter quantities. It can be observed that the optimized and enhanced data are more realistic in texture details, color and light and shadow performance, the visual effect is significantly improved, the domain gap is significantly reduced, and it is closer to the real mining area image, reflecting the significant advantages of this application scheme. In summary, this application proposes an autonomous driving data generation method that is adapted to the mining environment based on the special terrain, extreme weather conditions, long-tail scenes and other characteristics of mining scenes. Compared with related technologies, the generated data quality is higher and can support the perception model training and practical application of mining autonomous driving systems more effectively, safely and accurately.
[0095] In one embodiment of the present application, the method further comprises:
[0096] Acquire real mining area data in a mining area scenario, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image;
[0097] Perform noise processing on the real scene image to obtain noise data;
[0098] Based on the second control condition data and noise data, a mining area data generation model is trained.
[0099] In this embodiment, real mining area data for a mining area scenario is constructed based on real mining area data for the mining area scenario. Specifically, real mining area data for the mining area scenario is acquired, second control condition data for the real scene image is extracted, and noise data is generated by adding noise to the real scene image. Based on the second control condition data and the noise data, a mining area data generation model is trained to generate the mining area data generation model, which can subsequently be used to generate autonomous driving data for the mining area. The construction of the mining area data generation model is described in detail in a mining area data generation model training method provided in an embodiment of this application.
[0100] The present application embodiment provides a mining area data generation model training method, such as Figure 5 As shown, the method includes:
[0101] Step 501: Acquire real mining area data in a mining area scenario, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image;
[0102] Step 502: performing noise processing on the real scene image to obtain noise data;
[0103] Step 503: Based on the second control condition data and the noise data, a mining area data generation model is trained. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.
[0104] In this embodiment, real mining data from a mining scenario is obtained. To improve data generalization, control condition data is extracted from the real mining data, specifically, second control condition data for the real scene image. The real scene image is then subjected to noise processing to generate noise data. The noise data and second control condition data serve as the basis for training a mining data generation model, which is then used to generate autonomous driving mining data.
[0105] In one embodiment, a mining area data generation model is constructed based on the principle of a diffusion model. The principle of the diffusion model is to gradually add a small amount of noise to the real data through a forward diffusion process, so that the data distribution gradually approaches a smooth Gaussian noise distribution. Subsequently, when generating samples, the denoising process is reversed. That is, through a series of designed steps, clear samples that conform to the real data distribution are gradually recovered from completely random noise. In this reverse denoising process, each denoising step requires estimating the score function of the corresponding data distribution. This function is actually the gradient field of the probability density function, which guides the generation process to gradually move closer to areas with high probability density and good data quality.
[0106] Specifically, data generation follows a two-step process: (1) extracting a random vector from the prior distribution; (2) using an inverse Markov chain to gradually denoise and reconstruct high-quality new data points. This approach ensures that the generated data accurately fits the target data distribution while maintaining the stability and controllability of the generation process.
[0107] The embodiment of the present application can obtain real mining area data, and perform conditional optimization and enhancement on the real mining area data, and then train to obtain an accurate mining area data generation model. Subsequently, it can generate authentic data based on the mining area data generation model, ensuring the stability of the training effect of the mining area autonomous driving perception model.
[0108] In one embodiment of the present application, extracting second control condition data from a real scene image includes:
[0109] Convert the label information of the real scene image into a second text description feature, and perform feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature;
[0110] Performing position embedding processing on the spatial information of the real scene image to obtain a second spatial vector, where the spatial information includes resolution information and position information;
[0111] Performing encoding and feature fusion processing on the segmentation mask data of the real scene image to obtain a second fusion feature;
[0112] Second control condition data is obtained according to the second global semantic feature, the second spatial vector and the second fusion feature, and the second control condition data is mapped to a latent space corresponding to the mining area data generation model.
[0113] In this embodiment, Figure 6As shown, the label information of the real scene image is converted into a natural language description, namely the second text description feature, through hint engineering technology. The second text description feature is then feature extracted and further pooled to obtain the second global semantic feature. At the same time, the resolution information and position information of the real scene image are introduced to obtain the second spatial vector of the real scene image to supplement the spatial information of the image and ensure that the model can accurately locate the target object in the real scene image. In addition, the segmentation mask data of the real scene image is encoded and feature fused to provide structured prior information to assist the model in image generation.
[0114] The second control condition data includes a second global semantic feature, a second space vector and a second fusion feature. Finally, the second control condition data is mapped to the latent space corresponding to the model as input for model training.
[0115] Through the above method, the control conditions are extracted to achieve data enhancement.
[0116] In one embodiment of the present application, the label information includes a category label and detection box geometry information; converting the label information of the real scene image into a second text description feature includes:
[0117] The category label and detection frame geometry information of the real scene image are converted into a second text description feature, which includes the category label, detection frame geometry information, and the color, state, and background environment of the category label.
[0118] In this embodiment, the category labels and detection box geometry information of real scene images are converted into natural language descriptions through prompt engineering technology, which enables the model to utilize the multimodal information of the data.
[0119] like Figure 6 As shown in the figure, the target object's annotation information is converted into a natural language description through the hint engineering, which is also the second text description feature. This conversion method enhances the expressive ability of the target object and enables the model to understand more complex visual relationships.
[0120] In one embodiment of the present application, feature extraction and pooling processing are performed on the second text description feature to obtain a second global semantic feature, including:
[0121] The second text description feature is segmented by the first CLIP segmenter to obtain a third target text description feature, and the second target text description feature is encoded by the first CLIP encoder to obtain a third deep feature;
[0122] The second text description feature is segmented by a second CLIP segmenter to obtain a fourth target text description feature, and the fourth target text description feature is encoded by a second CLIP encoder to obtain a fourth deep feature;
[0123] Perform feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature;
[0124] Performing pooling processing on the fourth target text description feature to obtain the second global feature;
[0125] Obtain a second global semantic feature according to the second concatenated feature and the second global feature;
[0126] The word segmentation process includes: splitting the second text description feature into subword units, adding text tags to the subword units, and performing length unification processing on the subword units after adding text tags to obtain target text description features with unified length. Specifically, when the first CLIP word segmenter is used to perform word segmentation processing on the second text description feature, a third target text description feature with unified length is obtained; when the second CLIP word segmenter is used to perform word segmentation processing on the second text description feature, a fourth target text description feature with unified length is obtained.
[0127] In this embodiment, in order to enhance the model's ability to understand text control signals, CLIP is used to construct a text encoder to map natural language descriptions to the model's latent space.
[0128] The CLIP model consists of a CLIP tokenizer and a CLIP encoder, which uses a Transformer architecture. In the generation task, the second text description feature is input into the CLIP processing pipeline in natural language format, and the input text is normalized to ensure format consistency. Subsequently, the WordPiece tokenization method is used to break the text into smaller sub-word units, adapting them to the CLIP pre-trained vocabulary, which consists of (49,408) tokens (features). Special markers, such as [SOS], [EOS], and [PAD], are added to indicate the start and end of the text and padding content. Finally, a padding or truncation strategy is used to unify the length of the sub-word units after adding text markers, bringing the text length to the maximum number of features, which can be 77 tokens. This results in the target text description feature, ensuring that it meets the input specifications of the CLIP encoding process. This concludes the tokenization phase. The encoding phase then proceeds, using the Transformer architecture to extract deep features from the text, which serve as the text conditional input for the mining area data generation model.
[0129] In one embodiment, Figure 6As shown, the first CLIP segmenter and the first CLIP encoder act on the second text description feature to obtain the third deep feature s1'; the second CLIP segmenter and the second CLIP encoder act on the second text description feature to obtain the fourth deep feature s2'. Furthermore, the third deep feature s1' and the fourth deep feature s2' are concatenated to obtain the second concatenated feature s3'.
[0130] In addition, in order to further improve the model's ability to perceive the overall semantics of the text, the hidden state of the first token is extracted and pooled into a global vector to represent the overall semantic features of the input text. Figure 6 As shown, the fourth target text description feature is pooled to obtain the second global feature s4'. The second concatenated feature s3' and the second global feature s4' are then fused into the second global semantic feature s5'.
[0131] In this embodiment, two different CLIP models are used to extract text features and then pool them to obtain global semantic features. This dual encoding strategy can enhance the model's understanding of text semantics and improve the expression of complex scenes. It can also combine multimodal alignment information between text and image, thereby improving the quality and consistency of the synthesized image.
[0132] In one embodiment of the present application, performing noise processing on a real scene image to obtain noise data includes:
[0133] Perform dimensionality reduction processing on real scene images to obtain dimensionality reduction feature data;
[0134] The dimension-reduced feature data is subjected to noise processing to obtain noise data.
[0135] In this embodiment, during the model training process, directly modeling high-dimensional data often leads to a sharp increase in computational complexity and increases the difficulty of optimization. Therefore, the real scene image is first subjected to dimensionality reduction processing to obtain reduced-dimensionality feature data. After that, noise is added based on the reduced-dimensionality feature data, and then the model training continues.
[0136] In one embodiment of the present application, dimensionality reduction processing is performed on a real scene image to obtain dimensionality reduction feature data, including:
[0137] The probability distribution parameters of the real scene image in the latent space corresponding to the mining area data generation model are obtained through the VAE encoder. The probability distribution parameters include mean and variance.
[0138] According to the probability distribution parameters and standard Gaussian noise, the latent variables are calculated;
[0139] The latent variables are adjusted based on the adjustment factors, and the adjusted latent variables are decoded to obtain the reduced-dimensional feature data.
[0140] In this embodiment, VAE (Variational Auto Encoder) is used to encode real scene images and project them into the model’s latent space, thereby reducing data dimensions while retaining key semantic information, improving computational efficiency, and enhancing model stability. Specifically, the real scene images are passed through the VAE encoder to learn the probability distribution parameters, i.e., the mean. and variance , and then use the reparameterization technique to introduce standard Gaussian noise , I is the identity matrix, through the formula Calculating the latent variable z retains the gradient information, making the optimization process more stable. It also allows the model to adaptively adjust the representation of the latent space during the learning process. The latent variable z is then adjusted by the adjustment factor s and sent to the decoder to obtain the reduced dimension feature data. The overall process is as follows: Figure 7 shown.
[0141] Figure 7 The right side of the figure shows a channel visualization of the latent variable, where the four channels correspond to different latent feature components generated by the VAE. As can be seen from the visualization, the latent representation not only preserves the main structural information of the original image, but also captures texture details and local features in certain channels. This demonstrates that the VAE can compress the data while still preserving key information, enabling subsequent models to perform more stable reasoning based on this low-dimensional feature.
[0142] In one embodiment of the present application, the second control condition data includes: a second global semantic feature, a second spatial vector, and a second fusion feature; based on the second control condition data and the noise data, a mining area data generation model is trained, including:
[0143] Inputting the second control condition data and the noise data into a preset model, and outputting predicted noise data; wherein the preset model includes an original neural network, a conditional control network, and a connection network, the original neural network and the conditional control network are connected through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; inputting the second control condition data and the noise data into the preset model, including: inputting the noise data into the first sub-network of the original neural network, and fusing the noise data and the second fusion feature and inputting them into the conditional control network, and inputting the second global semantic feature and the second spatial vector into each sub-network of the conditional control network and the original neural network respectively;
[0144] Calculate the loss function based on the predicted noise data and the noise data;
[0145] The parameters of the preset model are iteratively optimized based on the loss function to obtain the mining area data generation model.
[0146] In this embodiment, the pre-set model uses a control network (ControlNet) to inject additional control condition data into the original neural network. This strategy allows the diffusion model to rely on external structured input during the generation process, thereby enhancing the model's perception and control of input constraints, allowing for more precise guidance of image generation results. ControlNet employs a freeze-copy strategy to ensure that the additional control condition data does not disrupt the generation capabilities of the original neural network.
[0147] Specifically, if Figure 8 As shown, the original neural network is frozen while a trainable copy, the conditional control network, is created, and the two are connected using a zero-convolution layer. A zero-convolution layer is a 1×1 convolution layer whose weights and biases are initialized to zero. That is, in the untrained state, this layer has no effect on the input data. In this way, in the early stages of training, ControlNet will not interfere with the generation process of the original neural network, thereby ensuring the stability of the model during the initialization phase and avoiding unnecessary deviations during inference. In addition, because the 1×1 zero-convolution layer only performs linear transformations between channels, its computational cost is extremely low, but it can gradually adjust the weights during training, thereby learning how to incorporate control condition data into the model's generation process.
[0148] In one embodiment of the present application, Figure 6 As shown in the figure, the noise data is input into the first sub-network of the original neural network, and the noise data and the second fusion feature are fused and input into the conditional control network. The second global semantic feature and the second spatial vector are respectively input into the conditional control network and each sub-network of the original neural network, and the predicted noise data is output. Furthermore, based on the predicted noise data and the noise data, the loss function is calculated. In ControlNet, the loss function is optimized based on the original neural network and is as follows:
[0149]
[0150] Where E is the expected value of the loss function, is the noise data, To predict noisy data, is the initial dimension reduction feature data, is the dimensionality reduction feature data of the real scene image, represents the second global semantic feature and the second space vector in the second control condition data, represents the second fusion feature in the second control condition data, t is the time step of the denoising process, are model parameters.
[0151] The parameters of the preset model are iteratively optimized based on the loss function until the iteration termination condition is met, and the mining area data generation model is obtained.
[0152] It should be noted that this application implements a data closed-loop system, which can use real mining data to verify the generalization performance of the mining data generation model, and provide timely feedback to the optimization process of the diffusion model and conditional control network, continuously improving the quality of the mining data generation model, so that data generation always remains consistent with the actual mining environment, ensuring the stability and continuous improvement of the training effect of the autonomous driving perception model.
[0153] In one embodiment of the present application, a cross-attention mechanism is introduced into each sub-network of the conditional control network and the original neural network.
[0154] In this embodiment, in order to ensure the effective transmission of control condition data, a cross-attention mechanism is introduced in each sub-network of the conditional control network and the original neural network, so that text embedding and spatial position information can fully influence the diffusion process.
[0155] In autonomous driving perception systems, vehicle detection and tracking is one of the core tasks, and its accuracy directly affects the reliability of high-level decisions such as path planning, collision avoidance, and behavior prediction. However, related technologies often adopt a unified generation strategy for foreground and background areas, which may result in missing details of foreground vehicles, blurred boundaries, and morphological distortion, thereby affecting the generalization ability of the perception model and its adaptability to real driving environments. To this end, in one embodiment of the present application, the method further includes:
[0156] For real scene images, define the detection frame set of the foreground area of the real scene image at the time step, and define the loss weight matrix of the detection frame set. The initial value of the loss weight matrix is set to an all-1 matrix.
[0157] The loss weight matrix is mapped to the latent space corresponding to the mining area data generation model using a bilinear interpolation method;
[0158] During the model training process, the loss weight matrix is dynamically adjusted. In the initial stage of model training, the loss weight matrix is controlled to increase so as to preferentially capture the key features of the foreground area. In the later stage of model training, the loss weight matrix is controlled to decrease so as to balance the features of the foreground area with the features of the background area of the real scene image.
[0159] In this example, a dynamic foreground weight loss mechanism is proposed. By adaptively adjusting the loss weight of the foreground area, the model is guided to gradually focus on the generation of foreground objects during training, thereby improving the quality and authenticity of the synthetic data. The core idea is to dynamically adjust the loss contribution of the foreground and background areas according to the different stages of the training process to achieve focused learning of foreground information while maintaining the integrity of the scene. The pseudo code is as follows: Figure 9 shown.
[0160] For each input image x , there are different numbers of detection boxes, using a set To represent the time step t The set of detection boxes under . Define a loss weight matrix , its initial value is set to a matrix of all 1s, that is, all pixels are given equal weight by default. During training, the optimization strength of the foreground area is dynamically adjusted through a two-stage cosine scheduling function.
[0161] Specifically, the initial stage of model training, i.e. When To train the threshold, the model hasn't fully learned the features of foreground objects. Therefore, a rapid growth strategy is used to rapidly increase the foreground weight, prompting the model to prioritize the structure, edges, and lighting characteristics of foreground vehicles. A cosine growth function is used for smooth control to ensure a stable training process and prevent sudden changes in gradient updates from impacting model convergence.
[0162] As the training progresses, the model gradually grasps the key features of the foreground area. If the foreground weight is still kept high in the subsequent training process, it may cause the model to overfit the foreground area, thereby affecting the learning of background information. Therefore, in the later stage of model training, that is, When , the foreground weight is gradually reduced to achieve a balance between foreground feature learning and background modeling.
[0163] The dynamic weight scheduling curve of this application is as follows Figure 10 As shown, the maximum loss weight value is set to , it reaches its highest value at 30% of the training cycle and then gradually decays.
[0164] This strategy ensures that the model effectively learns the scene's global information while generating foreground objects, thereby improving adaptability in diverse environments. Furthermore, given that the model's computations primarily occur in the latent space, a bilinear interpolation method is used to map the foreground loss weight matrix to the latent space to ensure consistency across scales. This method constructs a smooth foreground optimization matrix, thereby enhancing the influence of foreground objects in the diffusion model.
[0165] In one embodiment of the present application, the final loss function of the model is:
[0166]
[0167] Among them, E is the expected value of the loss function, is the loss weight matrix, is the set of detection boxes, is the noise data, To predict noisy data, is the dimensionality reduction feature data of the real scene image, c is the second control condition data, t is the time step, are model parameters.
[0168] To further improve the fidelity of foreground objects, this embodiment of the application uses foreground annotation boxes to calculate loss weights. A dynamic foreground weight loss mechanism is employed to adaptively adjust the optimization effort in the foreground region, enabling the model to properly balance the quality of foreground and background generation at different stages. Ultimately, by calculating the loss function and optimizing the model parameters, the model maintains global structural consistency while enhancing the generation of foreground details, ensuring that the generated results are both consistent with the text description and possess high visual fidelity and structural rationality.
[0169] As a specific implementation of the above-mentioned mining area data generation method, the embodiment of the present application provides a mining area data generation device. Figure 11 As shown, the mining area data generating device 1100 includes: a first data acquisition module 1101 , a first control condition extraction module 1102 , a noise data sampling module 1103 and a data generating module 1104 .
[0170] The first data acquisition module 1101 is used to acquire simulated mining area data for a mining area scene, where the simulated mining area data includes a virtual scene image and label information of the virtual scene image;
[0171] A first control condition extraction module 1102 is used to extract first control condition data of a virtual scene image;
[0172] Noise data sampling module 1103, used for sampling random noise data;
[0173] The data generation module 1104 is used to input the first control condition data and random noise data into the trained mining area data generation model to obtain target mining area data; wherein, the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.
[0174] Furthermore, the first data acquisition module 1101 is specifically configured to:
[0175] Construct a virtual environment of the mining area based on the collected data and digital environment model of the mining area;
[0176] Carry out vehicle dynamics simulation and sensor simulation in the virtual environment of the mining area to obtain a virtual vehicle moving in the virtual environment of the mining area and sensors running on the virtual vehicle;
[0177] Based on the virtual environment of the mining area, a variety of mining scenarios are constructed, and simulated mining data collected by sensors of virtual vehicles in the mining area scenarios are obtained. The mining area scenarios must include at least one of the following: long-tail scenarios, complex working condition scenarios, and extreme climate scenarios.
[0178] Furthermore, the first control condition extraction module 1102 is specifically configured to:
[0179] Converting the label information of the virtual scene image into a first text description feature, and performing feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature;
[0180] Performing position embedding processing on the spatial information of the virtual scene image to obtain a first spatial vector, where the spatial information includes resolution information and position information;
[0181] Performing encoding processing and feature fusion processing on the segmentation mask data of the virtual scene image to obtain a first fusion feature;
[0182] First control condition data is obtained according to the first global semantic feature, the first spatial vector and the first fusion feature, and the first control condition data is mapped to a latent space corresponding to the mining area data generation model.
[0183] Furthermore, the label information includes the category label and the detection box geometry information; the first control condition extraction module 1102 is specifically used to:
[0184] Converting the category label and detection frame geometry information of the virtual scene image into a first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label;
[0185] The first text description feature is segmented by a first CLIP segmenter to obtain a first target text description feature, and the first target text description feature is encoded by a first CLIP encoder to obtain a first deep feature;
[0186] The first text description feature is segmented by a second CLIP segmenter to obtain a second target text description feature, and the second target text description feature is encoded by a second CLIP encoder to obtain a second deep feature;
[0187] Performing feature splicing processing on the first deep feature and the second deep feature to obtain a first spliced feature;
[0188] Performing pooling processing on the second target text description feature to obtain the first global feature;
[0189] Obtaining a first global semantic feature according to the first concatenated feature and the first global feature;
[0190] The word segmentation process includes: splitting the first text description feature into subword units, adding text tags to the subword units, and performing length unification processing on the subword units after adding the text tags to obtain target text description features with unified length.
[0191] Furthermore, the mining area data generation model includes an original neural network, a conditional control network, and a connection network. The original neural network and the conditional control network are connected through the connection network. The original neural network includes a plurality of sub-networks connected in sequence. The first control condition data includes: a first global semantic feature, a first spatial vector, and a first fusion feature.
[0192] The data generation module 1104 is specifically configured to:
[0193] The first control condition data and random noise data are input into the mining area data generation model, and the target mining area data is output; wherein, the first control condition data and random noise data are input into the mining area data generation model, including: inputting the random noise data into the first sub-network of the original neural network, and fusing the random noise data and the first fusion feature and inputting them into the conditional control network, and inputting the first global semantic feature and the first space vector into each sub-network of the conditional control network and the original neural network, respectively.
[0194] Furthermore, the device also includes:
[0195] A mining area data generation model training module is used to obtain real mining area data in a mining area scenario, the real mining area data including real scene images and label information of the real scene images, and extract the second control condition data of the real scene images;
[0196] Perform noise processing on the real scene image to obtain noise data;
[0197] Based on the second control condition data and noise data, a mining area data generation model is trained.
[0198] The mining area data generating device 1100 in the embodiment of the present application can be a computer device, or a component in a computer device, such as an integrated circuit or a chip. The mining area data generating device 1100 provided in the embodiment of the present application can realize Figure 1To avoid repetition, the various processes implemented in the embodiment of the mining area data generation method are not described here.
[0199] As a specific implementation of the above-mentioned mining area data generation model training method, the embodiment of the present application provides a mining area data generation model training device. Figure 12 As shown, the mining area data generation model training device 1200 includes: a second data acquisition module 1201, a second control condition extraction module 1202, a data noise addition module 1203 and a model training module 1204.
[0200] The second data acquisition module 1201 is used to acquire real mining area data in a mining area scene, where the real mining area data includes real scene images and label information of the real scene images;
[0201] A second control condition extraction module 1202 is used to extract second control condition data of the real scene image;
[0202] The data noise adding module 1203 is used to perform noise adding processing on the real scene image to obtain noise data;
[0203] The model training module 1204 is used to train a mining area data generation model based on the second control condition data and the noise data. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.
[0204] Furthermore, the second control condition extraction module 1202 is specifically configured to:
[0205] Convert the label information of the real scene image into a second text description feature, and perform feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature;
[0206] Performing position embedding processing on the spatial information of the real scene image to obtain a second spatial vector, where the spatial information includes resolution information and position information;
[0207] Performing encoding and feature fusion processing on the segmentation mask data of the real scene image to obtain a second fusion feature;
[0208] Second control condition data is obtained according to the second global semantic feature, the second spatial vector and the second fusion feature, and the second control condition data is mapped to a latent space corresponding to the mining area data generation model.
[0209] Furthermore, the label information includes the category label and the detection box geometry information; the second control condition extraction module 1202 is specifically used to:
[0210] Convert the category label and detection frame geometry information of the real scene image into a second text description feature, where the second text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label;
[0211] The second text description feature is segmented by the first CLIP segmenter to obtain a third target text description feature, and the second target text description feature is encoded by the first CLIP encoder to obtain a third deep feature;
[0212] The second text description feature is segmented by a second CLIP segmenter to obtain a fourth target text description feature, and the fourth target text description feature is encoded by a second CLIP encoder to obtain a fourth deep feature;
[0213] Perform feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature;
[0214] Performing pooling processing on the fourth target text description feature to obtain the second global feature;
[0215] Obtain a second global semantic feature according to the second concatenated feature and the second global feature;
[0216] The word segmentation processing includes: splitting the second text description feature into subword units, adding text tags to the subword units, and performing length unification processing on the subword units after adding the text tags to obtain target text description features with unified length.
[0217] Furthermore, the data noise adding module 1203 is specifically configured to:
[0218] Perform dimensionality reduction processing on real scene images to obtain dimensionality reduction feature data;
[0219] The dimension-reduced feature data is subjected to noise processing to obtain noise data.
[0220] Furthermore, the data noise adding module 1203 is specifically configured to:
[0221] The probability distribution parameters of the real scene image in the latent space corresponding to the mining area data generation model are obtained through the VAE encoder. The probability distribution parameters include mean and variance.
[0222] According to the probability distribution parameters and standard Gaussian noise, the latent variables are calculated;
[0223] The latent variables are adjusted based on the adjustment factors, and the adjusted latent variables are decoded to obtain the reduced-dimensional feature data.
[0224] Furthermore, the second control condition data includes: a second global semantic feature, a second spatial vector, and a second fusion feature; the model training module 1204 is specifically used to:
[0225] Inputting the second control condition data and the noise data into a preset model, and outputting predicted noise data; wherein the preset model includes an original neural network, a conditional control network, and a connection network, the original neural network and the conditional control network are connected through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; inputting the second control condition data and the noise data into the preset model, including: inputting the noise data into the first sub-network of the original neural network, and fusing the noise data and the second fusion feature and inputting them into the conditional control network, and inputting the second global semantic feature and the second spatial vector into each sub-network of the conditional control network and the original neural network respectively;
[0226] Calculate the loss function based on the predicted noise data and the noise data;
[0227] The parameters of the preset model are iteratively optimized based on the loss function to obtain the mining area data generation model.
[0228] Furthermore, a cross-attention mechanism is introduced in the conditional control network and each sub-network of the original neural network.
[0229] Furthermore, the model training module 1204 is further configured to:
[0230] For real scene images, define the detection frame set of the foreground area of the real scene image at the time step, and define the loss weight matrix of the detection frame set. The initial value of the loss weight matrix is set to an all-1 matrix.
[0231] The loss weight matrix is mapped to the latent space corresponding to the mining area data generation model using a bilinear interpolation method;
[0232] During the model training process, the loss weight matrix is dynamically adjusted. In the initial stage of model training, the loss weight matrix is controlled to increase to prioritize capturing the key features of the foreground area. In the later stage of model training, the loss weight matrix is controlled to decrease to balance the features of the foreground area with the features of the background area of the real scene image.
[0233] The loss function is:
[0234]
[0235] Among them, E is the expected value of the loss function, is the loss weight matrix, is the set of detection boxes, is the noise data, To predict noisy data, is the dimensionality reduction feature data of the real scene image, c is the second control condition data, t is the time step, are model parameters.
[0236] The mining area data generation model training device 1200 in the embodiment of the present application can be a computer device, or a component in a computer device, such as an integrated circuit or a chip. The mining area data generation model training device 1200 provided in the embodiment of the present application can achieve Figure 5 To avoid repetition, the various processes implemented in the mining area data generation model training method embodiment will not be described here.
[0237] The present application also provides a computer device, such as Figure 13 As shown, the computer device 1300 includes a first processor 1301 and a first memory 1302. The first memory 1302 stores programs or instructions that can be run on the first processor 1301. When the program or instruction is executed by the first processor 1301, the various steps of the above-mentioned mining area data generation method or mining area data generation model training method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0238] The first memory 1302 can be used to store software programs and various data. The first memory 1302 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store an operating system, applications or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the first memory 1302 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM RAM (DRRAM). The first memory 1302 in the embodiment of the present application includes, but is not limited to, these and any other suitable types of memory.
[0239] The first processor 1301 may include one or more processing units. Optionally, the first processor 1301 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into the first processor 1301.
[0240] The present application also provides a readable storage medium. Figure 14 As shown, a program or instruction 1401 is stored on the readable storage medium 1400. When the program or instruction 1401 is executed by the processor, each process of the above-mentioned mining area data generation method or mining area data generation model training method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0241] The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. Computer-readable storage medium 1400 may include computer storage media and communication media, and may also include any medium that can transfer a computer program from one location to another. The storage medium may be any target medium that can be accessed by a computer.
[0242] As one possible design, computer-readable storage medium 1400 may include a Compact Disc Read-Only Memory (CD-ROM), RAM, ROM, EEPROM, or other optical disk storage; computer-readable storage media may include magnetic disk storage or other magnetic disk storage devices. Furthermore, any connection line may also be appropriately referred to as a computer-readable storage medium. For example, if software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL (Digital Subscriber Line), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers.
[0243] The embodiment of the present application also provides a chip, such as Figure 15 As shown, the chip 1500 includes at least one second processor 1501 and a communication interface 1502. The communication interface 1502 is coupled to the second processor 1501. The second processor 1501 is used to run programs or instructions to implement the various processes of the above-mentioned mining area data generation method or mining area data generation model training method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0244] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0245] Preferably, the chip 1500 further includes a memory, such as a second memory 1503 , which stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.
[0246] In the embodiment of the present application, the second memory 1503 may include a read-only memory and a random access memory, and provide instructions and data to the second processor 1501. A portion of the second memory 1503 may also include a non-volatile random access memory (NVRAM).
[0247] In the embodiment of the present application, the second processor 1501, the communication interface 1502 and the second memory 1503 are coupled together via a bus system 1504. In addition to the data bus, the bus system 1504 may also include a power bus, a control bus and a status signal bus. Figure 15 Various buses are labeled as bus system 1504.
[0248] The mining area map updating method described in the above embodiments of the present application can be applied to or implemented by the second processor 1501. The second processor 1501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the second processor 1501. The above-mentioned second processor 1501 can be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The second processor 1501 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.
[0249] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0250] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A method for generating mining area data, characterized in that: include: Acquire simulated mining area data for a mining area scene, the simulated mining area data including a virtual scene image and label information of the virtual scene image, and extract first control condition data of the virtual scene image; Sampling random noise data, and inputting the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of the vehicle in the mining area scenario; The first control condition data includes a first global semantic feature, a first spatial vector, and a first fusion feature. The first global semantic feature is obtained by converting the label information of the virtual scene image into a first text description feature and performing feature extraction and pooling processing on the first text description feature. The first spatial vector is obtained by performing position embedding processing on spatial information of the virtual scene image, and the spatial information includes resolution information and position information. The label information includes a category label and detection frame geometry information; converting the label information of the virtual scene image into a first text description feature, and performing feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature, including: Converting the category label and the detection frame geometry information of the virtual scene image into the first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; Performing word segmentation processing on the first text description feature through a first CLIP word segmenter to obtain a first target text description feature, and performing encoding processing on the first target text description feature through a first CLIP encoder to obtain a first deep feature; Performing word segmentation processing on the first text description feature through a second CLIP word segmenter to obtain a second target text description feature, and encoding the second target text description feature through a second CLIP encoder to obtain a second deep feature; Performing feature splicing processing on the first deep feature and the second deep feature to obtain a first spliced feature; Performing pooling processing on the second target text description feature to obtain a first global feature; The first global semantic feature is obtained according to the first splicing feature and the first global feature.
2. The method according to claim 1, characterized in that The obtaining of simulated mining area data for a mining area scenario includes: Construct a virtual environment of the mining area based on the collected data and digital environment model of the mining area; Performing vehicle dynamics simulation processing and sensor simulation processing in the virtual mining environment to obtain a virtual vehicle moving in the virtual mining environment and sensors operating on the virtual vehicle; A variety of mining scenarios are constructed based on the mining virtual environment, and the simulated mining data collected by the sensors of the virtual vehicle in the mining scenarios are obtained. The mining scenarios are at least one of the following: long-tail scenarios, complex working condition scenarios, and extreme climate scenarios.
3. The method according to claim 1, characterized in that The extracting the first control condition data of the virtual scene image includes: Converting the label information of the virtual scene image into a first text description feature, and performing feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature; Performing position embedding processing on the spatial information of the virtual scene image to obtain a first spatial vector; Performing encoding processing and feature fusion processing on the segmentation mask data of the virtual scene image to obtain a first fusion feature; The first control condition data is obtained according to the first global semantic feature, the first spatial vector and the first fusion feature, and the first control condition data is mapped to a latent space corresponding to the mining area data generation model.
4. The method according to claim 1, wherein The word segmentation process includes: splitting the first text description feature into subword units, adding text tags to the subword units, and performing length unification processing on the subword units after adding text tags to obtain target text description features with unified length.
5. The method according to claim 1, wherein The mining area data generation model includes an original neural network, a conditional control network and a connection network, wherein the original neural network and the conditional control network are connected via the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; The first control condition data includes: a first global semantic feature, a first space vector and a first fusion feature; The step of inputting the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data includes: The first control condition data and the random noise data are input into the mining area data generation model, and the target mining area data is output; wherein, the inputting of the first control condition data and the random noise data into the mining area data generation model includes: inputting the random noise data into the first sub-network of the original neural network, and fusing the random noise data and the first fusion feature and inputting them into the conditional control network, and inputting the first global semantic feature and the first spatial vector into each of the sub-networks of the conditional control network and the original neural network, respectively.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Acquire real mining area data in the mining area scene, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image; Performing noise processing on the real scene image to obtain noise data; The mining area data generation model is trained based on the second control condition data and the noise data.
7. A mining area data generation model training method, characterized in that: include: Acquire real mining area data in a mining area scenario, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image; Performing noise processing on the real scene image to obtain noise data; Based on the second control condition data and the noise data, a mining area data generation model is trained to generate target mining area data based on the simulated mining area data in the mining area scenario, and the target mining area data is used to control the automatic driving of the vehicle in the mining area scenario; The second control condition data includes a second global semantic feature, a second spatial vector, and a second fusion feature. The second global semantic feature is obtained by converting the label information of the real scene image into a second text description feature and performing feature extraction and pooling processing on the second text description feature. The second spatial vector is obtained by performing position embedding processing on spatial information of the real scene image, and the spatial information includes resolution information and position information. The label information includes category labels and detection frame geometry information; The method of converting the label information of the real scene image into a second text description feature, and performing feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature includes: Converting the category label and the detection frame geometry information of the real scene image into a second text description feature, where the second text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; Performing word segmentation processing on the second text description feature through a first CLIP word segmenter to obtain a third target text description feature, and performing encoding processing on the third target text description feature through a first CLIP encoder to obtain a third deep feature; Performing word segmentation processing on the second text description feature through a second CLIP word segmenter to obtain a fourth target text description feature, and performing encoding processing on the fourth target text description feature through a second CLIP encoder to obtain a fourth deep feature; Performing feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature; Performing pooling processing on the fourth target text description feature to obtain a second global feature; The second global semantic feature is obtained according to the second splicing feature and the second global feature.
8. The method according to claim 7, characterized in that The extracting the second control condition data of the real scene image includes: Converting the label information of the real scene image into a second text description feature, and performing feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature; Performing position embedding processing on the spatial information of the real scene image to obtain a second spatial vector; Performing encoding processing and feature fusion processing on the segmentation mask data of the real scene image to obtain a second fusion feature; The second control condition data is obtained according to the second global semantic feature, the second spatial vector and the second fusion feature, and the second control condition data is mapped to a latent space corresponding to the mining area data generation model.
9. The method according to claim 7, characterized in that The word segmentation process includes: splitting the second text description feature into subword units, adding text tags to the subword units, and performing length unification process on the subword units after adding the text tags to obtain a target text description feature with unified length.
10. The method according to claim 7, characterized in that The performing noise processing on the real scene image to obtain noise data includes: Performing dimensionality reduction processing on the real scene image to obtain dimensionality reduction feature data; Noise processing is performed on the dimension-reduced feature data to obtain the noise data.
11. The method according to claim 10, characterized in that The performing dimensionality reduction processing on the real scene image to obtain dimensionality reduction feature data includes: Obtaining probability distribution parameters of the real scene image in the latent space corresponding to the mining area data generation model through a VAE encoder, wherein the probability distribution parameters include mean and variance; Calculating latent variables based on the probability distribution parameters and standard Gaussian noise; The latent variables are adjusted based on the adjustment factors, and the adjusted latent variables are decoded to obtain the dimension-reduced feature data.
12. The method according to claim 7, characterized in that The second control condition data includes: a second global semantic feature, a second space vector and a second fusion feature; The training of a mining area data generation model based on the second control condition data and the noise data includes: Inputting the second control condition data and the noise data into a preset model, and outputting predicted noise data; wherein the preset model includes an original neural network, a conditional control network, and a connection network, the original neural network and the conditional control network are connected through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; inputting the second control condition data and the noise data into the preset model, including: inputting the noise data into a first sub-network of the original neural network, and fusing the noise data with the second fusion feature and inputting it into the conditional control network, and inputting the second global semantic feature and the second spatial vector into each of the sub-networks of the conditional control network and the original neural network respectively; Calculating a loss function based on the predicted noise data and the noise data; The parameters of the preset model are iteratively optimized based on the loss function to obtain the mining area data generation model.
13. The method according to claim 12, characterized in that A cross-attention mechanism is introduced into the conditional control network and each sub-network of the original neural network.
14. The method according to claim 12, characterized in that The method further comprises: For the real scene image, defining a detection frame set of the foreground area of the real scene image at a time step, and defining a loss weight matrix of the detection frame set, wherein an initial value of the loss weight matrix is set to an all-one matrix; Using a bilinear interpolation method, the loss weight matrix is mapped to a latent space corresponding to the mining area data generation model; During the model training process, the loss weight matrix is dynamically adjusted; wherein, in the initial stage of the model training, the loss weight matrix is controlled to increase so as to preferentially capture the key features of the foreground area; and in the later stage of the model training, the loss weight matrix is controlled to decrease so as to balance the features of the foreground area with the features of the background area of the real scene image; The loss function is: Where E is the expected value of the loss function, is the loss weight matrix, is the detection box set, is the noise data, is the predicted noise data, is the dimension reduction feature data of the real scene image, c is the second control condition data, t is the time step, are model parameters.
15. A mining area data generating device, characterized in that: include: A first data acquisition module is configured to acquire simulated mining area data for a mining area scene, wherein the simulated mining area data includes a virtual scene image and label information of the virtual scene image; A first control condition extraction module, configured to extract first control condition data of the virtual scene image; Noise data sampling module, used to sample random noise data; a data generation module, configured to input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of a vehicle in the mining area scenario; The first control condition data includes a first global semantic feature, a first spatial vector, and a first fusion feature. The first global semantic feature is obtained by converting the label information of the virtual scene image into a first text description feature and performing feature extraction and pooling processing on the first text description feature. The first spatial vector is obtained by performing position embedding processing on spatial information of the virtual scene image, and the spatial information includes resolution information and position information. The label information includes a category label and detection frame geometry information; converting the label information of the virtual scene image into a first text description feature, and performing feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature, including: Converting the category label and the detection frame geometry information of the virtual scene image into the first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; Performing word segmentation processing on the first text description feature through a first CLIP word segmenter to obtain a first target text description feature, and performing encoding processing on the first target text description feature through a first CLIP encoder to obtain a first deep feature; Performing word segmentation processing on the first text description feature through a second CLIP word segmenter to obtain a second target text description feature, and encoding the second target text description feature through a second CLIP encoder to obtain a second deep feature; Performing feature splicing processing on the first deep feature and the second deep feature to obtain a first spliced feature; Performing pooling processing on the second target text description feature to obtain a first global feature; The first global semantic feature is obtained according to the first splicing feature and the first global feature.
16. A mining area data generation model training device, characterized in that: include: A second data acquisition module is used to acquire real mining area data in a mining area scene, wherein the real mining area data includes a real scene image and label information of the real scene image; A second control condition extraction module, configured to extract second control condition data of the real scene image; A data noise adding module, configured to perform noise adding processing on the real scene image to obtain noise data; a model training module, configured to train a mining area data generation model based on the second control condition data and the noise data, wherein the mining area data generation model is configured to generate target mining area data based on simulated mining area data in the mining area scenario, and the target mining area data is configured to control automatic driving of a vehicle in the mining area scenario; The second control condition data includes a second global semantic feature, a second spatial vector, and a second fusion feature. The second global semantic feature is obtained by converting the label information of the real scene image into a second text description feature and performing feature extraction and pooling processing on the second text description feature. The second spatial vector is obtained by performing position embedding processing on spatial information of the real scene image, and the spatial information includes resolution information and position information. The label information includes category labels and detection frame geometry information; The method of converting the label information of the real scene image into a second text description feature, and performing feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature includes: Converting the category label and the detection frame geometry information of the real scene image into a second text description feature, where the second text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; Performing word segmentation processing on the second text description feature through a first CLIP word segmenter to obtain a third target text description feature, and performing encoding processing on the third target text description feature through a first CLIP encoder to obtain a third deep feature; Performing word segmentation processing on the second text description feature through a second CLIP word segmenter to obtain a fourth target text description feature, and performing encoding processing on the fourth target text description feature through a second CLIP encoder to obtain a fourth deep feature; Performing feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature; Performing pooling processing on the fourth target text description feature to obtain a second global feature; The second global semantic feature is obtained according to the second splicing feature and the second global feature.
17. A computer device, characterized in that: It includes a first processor and a first memory, the first memory stores a program or instruction running on the first processor, and when the program or instruction is executed by the first processor, it implements the steps of the mining area data generation method as described in any one of claims 1 to 6, or implements the steps of the mining area data generation model training method as described in any one of claims 7 to 14.
18. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by the processor, the steps of the mining area data generation method according to any one of claims 1 to 6 are implemented, or the steps of the mining area data generation model training method according to any one of claims 7 to 14 are implemented.
19. A chip, characterized in that: The chip includes at least one second processor and a communication interface, the communication interface is coupled to the at least one second processor, and the at least one second processor is used to run programs or instructions to implement the steps of the mining area data generation method as described in any one of claims 1 to 6, or to implement the steps of the mining area data generation model training method as described in any one of claims 7 to 14.