Data generation method and device, model training method and device, equipment, medium and chip

By using simulation data and data generation models in the mining area autonomous driving system, the problems of high data acquisition cost and low quality are solved, and efficient and safe data generation and stable training of autonomous driving perception models are achieved.

CN120030626AActive Publication Date: 2025-05-23UNIV OF CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510521223.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art has high data acquisition costs and low data quality in mining autonomous driving systems. Especially in special mine scenarios, data acquisition has safety hazards and errors.

Method used

By obtaining simulated mining area data, extracting control condition data, and using trained mining area data to generate the target mining area data to control autonomous driving.

Benefits of technology

There is no need to rely on real-world data collection and manual annotation, which significantly reduces data generation costs, improves data quality and safety, and enhances the training effect of autonomous driving perception models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030626A_ABST
    Figure CN120030626A_ABST
Patent Text Reader

Abstract

The invention provides a data generation method and device, a model training method and device, equipment, a medium and a chip, and relates to the field of mining area automatic driving and image processing. The method comprises the steps that simulation mining area data in a mining area scene are acquired, the simulation mining area data comprise a virtual scene image and label information of the virtual scene image, and first control condition data of the virtual scene image are extracted; sampling random noise data, and inputting the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is obtained based on training of real mining area data, and the target mining area data is used for controlling automatic driving of vehicles in the mining area scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving and image processing in mining areas, and in particular to a mining area data generation method, a mining area data generation model training method, a mining area data generation device, a mining area data generation model training device, computer equipment, readable storage media and chips. Background Art

[0002] The core of the autonomous driving system lies in the acquisition and labeling of high-quality perception data. The methods in related technologies deploy multimodal sensors in real environments, collect scene information, and then manually or semi-automatically label task labels such as target detection and semantic segmentation. This method has the following disadvantages: (1) Data collection is completely dependent on the real environment, and the data collection cost is high. Especially in special scenarios such as mines, the data collection cost increases significantly. At the same time, there are many environmental risk factors in mining areas, and data collection on site is limited. The collection process is complicated and there are safety hazards.

[0003] (2) Real data requires a lot of manual or auxiliary tool annotation, which consumes a lot of resources. In addition, the annotation results are prone to errors and inconsistencies, which reduces data quality. Summary of the invention

[0004] In view of this, the present application provides a mining area data generation method, a mining area data generation model training method, a mining area data generation device, a mining area data generation model training device, computer equipment, readable storage medium and chip, which solve the problems of high data collection cost and low data quality.

[0005] In a first aspect, an embodiment of the present application provides a method for generating mining area data, comprising: Acquire simulated mining area data for a mining area scene, the simulated mining area data including a virtual scene image and label information of the virtual scene image, and extract first control condition data of the virtual scene image; Sample random noise data, and input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.

[0006] In a second aspect, an embodiment of the present application provides a mining area data generation model training method, comprising: Acquire real mining area data in a mining area scene, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image; Performing noise processing on the real scene image to obtain noise data; Based on the second control condition data and the noise data, a mining area data generation model is trained, and the mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario, and the target mining area data is used to control the automatic driving of vehicles in the mining area scenario.

[0007] In a third aspect, an embodiment of the present application provides a mining area data generating device, comprising: A first data acquisition module, used to acquire simulated mining area data for a mining area scene, wherein the simulated mining area data includes a virtual scene image and label information of the virtual scene image; A first control condition extraction module, used to extract first control condition data of the virtual scene image; Noise data sampling module, used to sample random noise data; A data generation module is used to input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.

[0008] In a fourth aspect, an embodiment of the present application provides a mining area data generation model training device, comprising: A second data acquisition module is used to acquire real mining area data in a mining area scene, wherein the real mining area data includes a real scene image and label information of the real scene image; A second control condition extraction module, used to extract second control condition data of the real scene image; A data noise adding module, used for performing noise adding processing on the real scene image to obtain noise data; A model training module is used to train a mining area data generation model based on the second control condition data and the noise data. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.

[0009] In a fifth aspect, an embodiment of the present application provides a computer device, which includes a first processor and a first memory, wherein the first memory stores programs or instructions that can be executed on the first processor, and when the programs or instructions are executed by the first processor, the steps of the method of the first aspect or the second aspect are implemented.

[0010] In a sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method of the first aspect or the second aspect are implemented.

[0011] In the seventh aspect, an embodiment of the present application provides a chip, which includes at least one second processor and a communication interface, wherein the communication interface is coupled to the at least one second processor, and the at least one second processor is used to run programs or instructions to implement the steps of the method of the first aspect or the second aspect.

[0012] The beneficial effects of the present application are: the mining area data generation method, mining area data generation model training method, mining area data generation device, mining area data generation model training device, computer equipment, readable storage medium and chip provided in the embodiments of the present application obtain simulated mining area data for mining area scenes, extract control condition data from the simulated mining area data, thereby further optimizing the simulated mining area data, and then use the control condition data as a guide to generate target mining area data for controlling mining area autonomous driving through the trained mining area data generation model. The present application does not need to rely on real environment data collection and does not need to perform a large amount of data labeling work, and can generate a large amount of mining area autonomous driving data based on simulated mining area data and trained mining area data generation models. And through the guidance of control conditions, the data authenticity is effectively improved, the domain gap between simulation data and real mining area scenes is significantly reduced, the quality of generated data is improved, and the stability of the training effect of the mining area autonomous driving perception model is ensured.

[0013] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram showing a process flow of a method for generating mining area data according to an embodiment of the present application is shown; Figure 2 A schematic diagram showing the construction of simulated mining area data in an embodiment of the present application is shown; Figure 3 A logical schematic diagram of a method for generating mining area data according to an embodiment of the present application is shown; Figure 4 A schematic diagram showing a comparison of the effects of the embodiments of the present application and the related art; Figure 5 A schematic diagram showing a process flow of a mining area data generation model training method according to an embodiment of the present application is shown; Figure 6A schematic diagram showing the mining area data generation model training of an embodiment of the present application is shown; Figure 7 A schematic diagram showing the dimensionality reduction process of an embodiment of the present application is shown; Figure 8 A schematic diagram showing a preset model of an embodiment of the present application is shown; Fig. 9 A pseudo code screenshot of the dynamic foreground weight calculation of an embodiment of the present application is shown; Fig.10 A schematic diagram showing a dynamic weight scheduling curve according to an embodiment of the present application is shown; Fig.11 A structural block diagram of a mining area data generating device according to an embodiment of the present application is shown; Fig.12 A structural block diagram of a mining area data generation model training device according to an embodiment of the present application is shown; Fig.13 A structural block diagram of a computer device according to an embodiment of the present application is shown; Fig.14 A structural block diagram of a readable storage medium according to an embodiment of the present application is shown; Fig.15 The structure block diagram of the chip of the embodiment of the present application is shown. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0016] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0017] In conjunction with the accompanying drawings, the mining area data generation method, mining area data generation model training method, mining area data generation device, mining area data generation model training device, computer equipment, readable storage medium and chip provided in the embodiments of the present application are described in detail below through specific embodiments and their application scenarios. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0018] The present application embodiment provides a method for generating mining area data, such as Figure 1 As shown, the method includes: Step 101, obtaining simulated mining area data for a mining area scene, the simulated mining area data including a virtual scene image and label information of the virtual scene image, and extracting first control condition data of the virtual scene image.

[0019] In this step, simulated mining area data for the mining area scene is obtained, and the simulated mining area data includes a virtual scene image and label information of the virtual scene image. The label information includes a category label and detection frame geometry information. The simulated mining area data is used to subsequently generate target mining area data for controlling the automatic driving of vehicles in the mining area scene.

[0020] In one embodiment of the present application, obtaining simulated mining area data for a mining area scenario includes: Construct a virtual environment of the mining area based on the mining area collection data and digital environment model; Carry out vehicle dynamic simulation and sensor simulation in the virtual environment of the mining area to obtain a virtual vehicle moving in the virtual environment of the mining area and sensors running on the virtual vehicle; Based on the virtual environment of the mining area, various mining scenes are constructed, and simulated mining data collected by sensors of virtual vehicles in the mining scenes are obtained. The mining scenes include at least one of the following: long-tail scenes, complex working condition scenes, and extreme climate scenes.

[0021] In this embodiment, if Figure 2 As shown in the figure, the simulation engine (Unity) is used to construct the simulated mining area data. The specific implementation steps include: (1) Through mining area data collection and digital environment modeling, a virtual environment for generating mining area scenes is built, and the collected data is post-processed and visually rendered to construct a realistic mining area virtual environment.

[0022] (2) Carry out vehicle dynamics simulation and sensor simulation. Sensors include lidar, visual cameras, etc., to realize the movement of virtual vehicles in mining scenes and the operation of sensors on virtual vehicles.

[0023] (3) Obtain scene parameter configuration, construct various mining scenarios such as long-tail scenarios, complex working condition scenarios, and extreme climate scenarios, and perform mining operation scheduling and management, including vehicle loading, transportation, and unloading. Among them, long-tail scenarios include special obstacles or abnormal events, and extreme climates include sandstorms and rainstorms. In addition, obtain the virtual scene image output by the sensor in real time, and use the automatic annotation tool to automatically annotate the virtual scene image to obtain the label information of the virtual scene image.

[0024] Through the above-mentioned method, accurate simulation of the dynamic work of the virtual mine area is achieved, providing a data source for the dynamic environment of the complex mine area, and ensuring the authenticity and diversity of the generated data.

[0025] The real data collection in related technologies is costly and time-consuming, especially in dangerous areas of mines where data acquisition is difficult and presents safety risks. The embodiment of the present application adopts simulation-driven data generation technology, with the help of virtual sensors and automated annotation tools, to efficiently generate high-precision simulation data and precise annotations in batches as input for generating target mine data for controlling the automatic driving of vehicles in mine scenes, thereby avoiding the high cost of data collection and manual annotation in real scenes, and significantly improving the efficiency and safety of data generation.

[0026] It is worth noting that through flexible scenario parameter configuration, this application can effectively simulate the complex terrain, extreme weather and long-tail scenarios unique to mining areas in a simulation environment, so that the autonomous driving dataset covers more diverse boundary cases and improves the robustness and generalization ability of the autonomous driving perception model in complex scenarios.

[0027] Furthermore, after constructing the simulated mining area data, in order to improve the authenticity and generalization ability of the simulated mining area data, control condition data is extracted from the simulated mining area data, that is, the first control condition data of the virtual scene image is extracted as input for the subsequent mining area data generation model.

[0028] In one embodiment of the present application, extracting first control condition data of a virtual scene image includes: Convert the label information of the virtual scene image into a first text description feature, and perform feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature; Performing position embedding processing on the spatial information of the virtual scene image to obtain a first spatial vector, the spatial information including resolution information and position information; Performing encoding processing and feature fusion processing on the segmentation mask data of the virtual scene image to obtain a first fusion feature; According to the first global semantic feature, the first space vector and the first fusion feature, first control condition data are obtained, and the first control condition data are mapped to a latent space corresponding to the mining area data generation model.

[0029] In this embodiment, if Figure 3 As shown, the label information of the virtual scene image is converted into a natural language description, that is, the first text description feature, through the prompt engineering technology, and the first text description feature is feature extracted and further pooled to obtain the first global semantic feature. At the same time, the resolution information and position information of the virtual scene image are introduced to obtain the first spatial vector of the virtual scene image to supplement the spatial information of the image and ensure that the subsequent mining area data generation model can accurately locate the target object of the virtual scene image. In addition, the segmentation mask data of the virtual scene image is encoded and feature fused to provide structured prior information for assisting the mining area data generation model in image generation.

[0030] The first control condition data includes the first global semantic feature, the first space vector and the first fusion feature. Finally, the first control condition data is mapped to the latent space corresponding to the mining area data generation model as the input of the mining area data generation model for data generation by the mining area data generation model.

[0031] Through the above method, the control conditions are extracted to achieve data enhancement.

[0032] In one embodiment of the present application, the label information includes a category label and detection frame geometry information; converting the label information of the virtual scene image into a first text description feature includes: The category label and the detection frame geometry information of the virtual scene image are converted into a first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label.

[0033] In this embodiment, the prompt engineering technology converts the category labels and detection box geometry information of the virtual scene image into natural language descriptions through the prompt engineering method, so that the mining area data generation model can utilize the multimodal information of the data.

[0034] In target detection tasks in related technologies, the annotations of target objects are usually stored in the following format: {"ID":"truck","bbox":[x, y, w, h]}. Among them, ID represents the category of the detection box of the target object. For example, "truck" represents a mining truck, and "bbox" represents the detection box (bounding_box), which is a list of four values, representing the pixel coordinates (x, y) of the center point of the detection box and the width (w) and height (h) of the detection box. However, this format only provides geometric information of the target object and it is difficult to express higher-level semantic relationships, such as the color, state, background environment, etc. of the target object. In this application, if Figure 3 As shown in the figure, the annotation information of the target object is converted into a natural language description, that is, the first text description feature, through the prompt engineering. For example: In the image, a large yellow dump truck is positioned at coordinates (x, y), with its body partially obscured by airborne dust. The vehicle spans w pixels in width and h pixels in height. This conversion method enhances the expressiveness of the target object and enables the mining area data generation model to understand more complex visual relationships.

[0035] In one embodiment of the present application, feature extraction and pooling processing are performed on the first text description feature to obtain a first global semantic feature, including: The first text description feature is segmented by a first CLIP segmenter to obtain a first target text description feature, and the first target text description feature is encoded by a first CLIP encoder to obtain a first deep feature; The first text description feature is segmented by a second CLIP segmenter to obtain a second target text description feature, and the second target text description feature is encoded by a second CLIP encoder to obtain a second deep feature; Perform feature splicing processing on the first deep feature and the second deep feature to obtain a first splicing feature; Performing pooling processing on the second target text description feature to obtain the first global feature; Obtaining a first global semantic feature according to the first concatenated feature and the first global feature; The word segmentation process includes: splitting the first text description feature into subword units, adding text tags to the subword units, and performing length unification process on the subword units after adding the text tags to obtain a target text description feature with a unified length. Specifically, when the first text description feature is subjected to word segmentation process using the first CLIP word segmenter, a first target text description feature with a unified length is obtained; when the first text description feature is subjected to word segmentation process using the second CLIP word segmenter, a second target text description feature with a unified length is obtained.

[0036] In this embodiment, in order to enhance the ability of the mining area data generation model to understand text control signals, CLIP (Contrastive Language-Image Pre-training) is used to build a text encoder to map the natural language description to the latent space of the mining area data generation model, where the mining area data generation model is built based on the principle of the diffusion model, thereby improving the text's ability to guide the generated content.

[0037] The CLIP model includes a CLIP word segmenter and a CLIP encoder, and the encoder adopts a Transformer structure. In the generation task, the first text description feature is input into the CLIP processing flow in the form of natural language, and the input text is normalized to ensure format consistency. Subsequently, the WordPiece word segmentation method is used to decompose the text into smaller sub-word units to adapt it to the CLIP pre-trained vocabulary, which is (49,408) tokens (features). Next, special tags are added, such as [SOS], [EOS], [PAD], etc., to indicate the start, end and filling content of the text. Finally, a padding or truncation strategy is used to unify the length of the sub-word units after adding text tags, unify the text length to the maximum number of features, which can be 77 tokens, and obtain the target text description feature to ensure that it meets the input specifications of the CLIP encoding processing flow. At this time, the word segmentation part ends. Then the encoding part is carried out, that is, the deep features of the text are extracted using the Transformer structure as the text conditional input of the mining area data generation model.

[0038] In one embodiment, Figure 3 As shown, the first CLIP word segmenter and the first CLIP encoder act on the first text description feature to obtain the first deep feature s1; the second CLIP word segmenter and the second CLIP encoder act on the first text description feature to obtain the second deep feature s2. Further, the first deep feature s1 and the second deep feature s2 are feature concatenated to obtain the first concatenated feature s3.

[0039] In addition, in order to further improve the mining data generation model’s ability to perceive the overall semantics of the text, the hidden state of the first token is extracted and pooled into a global vector to represent the overall semantic features of the input text. Figure 3 As shown, the second target text description feature is pooled to obtain the first global feature s4. Then the first concatenated feature s3 and the first global feature s4 are fused into the first global semantic feature s5.

[0040] In the embodiment of the present application, two different CLIP models are used to extract features from text, and further pooling is performed to obtain global semantic features. This dual encoding strategy can enhance the model's ability to understand text semantics and improve the expression effect of complex scenes. It can also combine the multimodal alignment information of text and image, thereby improving the quality and consistency of the synthesized image.

[0041] Step 102, sampling random noise data, and inputting the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.

[0042] In this embodiment, random noise data of the same size as the scene image is sampled from the normal distribution data, and the first control condition data and the random noise data are input into the trained mining area data generation model, and the target mining area data with the real mining area style is output, and the target mining area data is used to expand the mining area automatic driving data set to train the perception model. In some embodiments, the quality of the simulated mining area data generation can also be continuously improved according to the effect of the perception model.

[0043] In one embodiment, the mining area data generation model is constructed based on the principle of the stable diffusion model.

[0044] In the embodiment of the present application, simulated mining area data for mining area scenarios are obtained, and control condition data is extracted from the simulated mining area data, thereby further optimizing the simulated mining area data, and then the control condition data is used as a guide to generate target mining area data for controlling the autonomous driving of the mining area through the trained mining area data generation model. The present application does not need to rely on real environment data collection and does not need to perform a large amount of data labeling work. It can generate a large amount of mining area autonomous driving data based on simulated mining area data and trained mining area data generation models. And through the guidance of control conditions, the data authenticity is effectively improved, the domain gap between the simulated data and the real mining area scenario is significantly reduced, the quality of the generated data is improved, and the stability of the training effect of the mining area autonomous driving perception model is ensured.

[0045] In one embodiment of the present application, as Figure 3 shown, the mining area data generation model includes an original neural network, a conditional control network, and a connection network. The connection network is a zero convolutional layer. The original neural network and the conditional control network are connected through the connection network. The original neural network includes a plurality of sub-networks connected in sequence.

[0046] In one embodiment of the present application, the first control condition data includes: a first global semantic feature, a first spatial vector, and a first fusion feature; inputting the first control condition data and random noise data into the trained mining area data generation model to obtain target mining area data, including: Inputting the first control condition data and random noise data into the mining area data generation model, and outputting target mining area data; wherein, inputting the first control condition data and random noise data into the mining area data generation model includes: inputting the random noise data into the first sub-network of the original neural network, and inputting the fused random noise data and the first fusion feature into the conditional control network, and inputting the first global semantic feature and the first spatial vector into the conditional control network and each sub-network of the original neural network respectively.

[0047] In this embodiment, inputting the random noise data into the first sub-network of the original neural network, inputting the fused random noise data and the first fusion feature into the conditional control network, and inputting the first global semantic feature and the first spatial vector into the conditional control network and each sub-network of the original neural network respectively, finally the mining area data generation model outputs target mining area data with higher authenticity. In the embodiment of the present application, using various conditional information such as text descriptions, segmentation masks, resolution information, and location information of the simulated mining area data to guide the mining area data generation model to generate high-quality images with the real style of the mining area.

[0048] The following combines Figure 4 to illustrate the effect comparison between the embodiment of the present application and the related technology: Figure 4 (a) in is the original mine scene data generated by the simulation platform (Unity). Although it can simulate the typical scenes and layouts of the mining area, the visual details are insufficient, and there is an obvious visual gap with the real scene. Figure 4 (b) in is the semantic segmentation label data of the corresponding scene. These labels are automatically generated and can effectively support the training of subsequent perception models. Figure 4(c) and (d) are the optimized mining scene data proposed in this application. (c) and (d) are the effects of two different parameter quantities. It can be observed that the optimized and enhanced data are more realistic in texture details, color and light and shadow performance, the visual effect is significantly improved, the domain gap is significantly reduced, and it is closer to the real mining area image, reflecting the significant advantages of the present application scheme. In summary, this application proposes an autonomous driving data generation method adapted to the mining environment based on the special terrain, extreme weather conditions, long-tail scenes and other characteristics of mining scenes. Compared with related technologies, the generated data is of higher quality and can more effectively, safely and accurately support the perception model training and practical application of mining autonomous driving systems.

[0049] In one embodiment of the present application, the method further includes: Acquire real mining area data in a mining area scene, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image; Perform noise processing on the real scene image to obtain noise data; Based on the second control condition data and the noise data, a mining area data generation model is trained.

[0050] In this embodiment, real mining area data in a mining area scenario is constructed based on real mining area data in a mining area scenario. Specifically, real mining area data in a mining area scenario is obtained, second control condition data of the real scene image is extracted, and noise data is obtained by adding noise to the real scene image. Based on the second control condition data and the noise data, a mining area data generation model is trained, and the mining area data generation model can be subsequently used to generate mining area autonomous driving data. The construction of the mining area data generation model is described in detail in a mining area data generation model training method provided in an embodiment of the present application.

[0051] The present application embodiment provides a mining area data generation model training method, such as Figure 5 As shown, the method includes: Step 501, obtaining real mining area data in a mining area scene, the real mining area data including a real scene image and label information of the real scene image, and extracting second control condition data of the real scene image; Step 502, performing noise processing on the real scene image to obtain noise data; Step 503, based on the second control condition data and the noise data, a mining area data generation model is trained, the mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario, and the target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.

[0052] In this embodiment, real mining area data in a mining area scenario is obtained, and control condition data is extracted from the real mining area data to improve the generalization ability of the data, that is, the second control condition data of the real scene image is extracted. The real scene image is subjected to noise processing to obtain noise data, and the noise data and the second control condition data are used as the basis for training the mining area data generation model, and the mining area data generation model is used to subsequently generate mining area autonomous driving data.

[0053] In one embodiment, a mining area data generation model is constructed based on the principle of a diffusion model. The principle of the diffusion model is to gradually add a small amount of noise to the real data through a forward diffusion process, so that the data distribution gradually tends to a smooth Gaussian noise distribution. Subsequently, the denoising process is reversed when generating samples, that is, through a series of design steps, clear samples that conform to the real data distribution are gradually restored from completely random noise. In this reverse denoising process, each denoising step needs to estimate the score function of the corresponding data distribution. This function is actually the gradient field of the probability density function, which guides the generation process to gradually move closer to the area with high probability density and good data quality.

[0054] Specifically, data generation follows a two-step process: (1) extracting a random vector from the prior distribution; (2) using an inverse Markov chain to gradually denoise and reconstruct high-quality new data points. This approach ensures that the generated data can accurately fit the target data distribution while maintaining the stability and controllability of the generation process.

[0055] The embodiment of the present application can obtain real mining area data, and conditionally optimize and enhance the real mining area data, and then train to obtain an accurate mining area data generation model. Subsequently, it can generate authentic data based on the mining area data generation model, thereby ensuring the stability of the training effect of the mining area autonomous driving perception model.

[0056] In one embodiment of the present application, extracting the second control condition data of the real scene image includes: Convert the label information of the real scene image into a second text description feature, and perform feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature; Performing position embedding processing on the spatial information of the real scene image to obtain a second spatial vector, wherein the spatial information includes resolution information and position information; Performing encoding processing and feature fusion processing on the segmentation mask data of the real scene image to obtain a second fusion feature; According to the second global semantic feature, the second space vector and the second fusion feature, the second control condition data is obtained, and the second control condition data is mapped to the latent space corresponding to the mining area data generation model.

[0057] In this embodiment, if Figure 6 As shown, the label information of the real scene image is converted into a natural language description, that is, the second text description feature, through the hint engineering technology, and the second text description feature is feature extracted and further pooled to obtain the second global semantic feature. At the same time, the resolution information and position information of the real scene image are introduced to obtain the second spatial vector of the real scene image to supplement the spatial information of the image and ensure that the model can accurately locate the target object of the real scene image. In addition, the segmentation mask data of the real scene image is encoded and feature fused to provide structured prior information to assist the model in image generation.

[0058] The second control condition data includes a second global semantic feature, a second space vector and a second fusion feature. Finally, the second control condition data is mapped to the latent space corresponding to the model as an input for model training.

[0059] Through the above method, the control conditions are extracted to achieve data enhancement.

[0060] In one embodiment of the present application, the label information includes a category label and detection frame geometry information; converting the label information of the real scene image into a second text description feature includes: The category label and detection frame geometry information of the real scene image are converted into a second text description feature, where the second text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label.

[0061] In this embodiment, the category labels and detection box geometry information of real scene images are converted into natural language descriptions through prompt engineering technology, which enables the model to utilize the multimodal information of the data.

[0062] like Figure 6 As shown in the figure, the annotation information of the target object is converted into a natural language description, that is, the second text description feature, through the prompt engineering. This conversion method enhances the expressive ability of the target object and enables the model to understand more complex visual relationships.

[0063] In one embodiment of the present application, feature extraction and pooling processing are performed on the second text description feature to obtain a second global semantic feature, including: The second text description feature is segmented by the first CLIP segmenter to obtain a third target text description feature, and the second target text description feature is encoded by the first CLIP encoder to obtain a third deep feature; The second text description feature is segmented by a second CLIP segmenter to obtain a fourth target text description feature, and the fourth target text description feature is encoded by a second CLIP encoder to obtain a fourth deep feature; Perform feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature; Performing pooling processing on the fourth target text description feature to obtain a second global feature; Obtain a second global semantic feature according to the second concatenated feature and the second global feature; The word segmentation process includes: splitting the second text description feature into subword units, adding text tags to the subword units, and performing length unification process on the subword units after adding the text tags to obtain a target text description feature with a unified length. Specifically, when the first CLIP word segmenter is used to perform word segmentation process on the second text description feature, a third target text description feature with a unified length is obtained; when the second CLIP word segmenter is used to perform word segmentation process on the second text description feature, a fourth target text description feature with a unified length is obtained.

[0064] In this embodiment, in order to enhance the model's ability to understand text control signals, CLIP is used to construct a text encoder to map natural language descriptions to the model's latent space.

[0065] The CLIP model includes a CLIP word segmenter and a CLIP encoder, and the encoder adopts a Transformer structure. In the generation task, the second text description feature is input into the CLIP processing flow in the form of natural language, and the input text is normalized to ensure format consistency. Subsequently, the WordPiece word segmentation method is used to decompose the text into smaller sub-word units to adapt it to the CLIP pre-trained vocabulary, which is (49,408) tokens (features). Next, special tags are added, such as [SOS], [EOS], [PAD], etc., to indicate the start, end and filling content of the text. Finally, a padding or truncation strategy is used to unify the length of the sub-word units after adding text tags, unify the text length to the maximum number of features, which can be 77 tokens, and obtain the target text description feature to ensure that it meets the input specifications of the CLIP encoding processing flow. At this time, the word segmentation part ends. Then the encoding part is carried out, that is, the deep features of the text are extracted using the Transformer structure as the text conditional input of the mining area data generation model.

[0066] In one embodiment, Figure 6As shown, the first CLIP segmenter and the first CLIP encoder act on the second text description feature to obtain the third deep feature s1'; the second CLIP segmenter and the second CLIP encoder act on the second text description feature to obtain the fourth deep feature s2'. Further, the third deep feature s1' and the fourth deep feature s2' are processed by feature splicing to obtain the second splicing feature s3'.

[0067] In addition, in order to further improve the model's ability to perceive the overall semantics of the text, the hidden state of the first token is extracted and pooled into a global vector to represent the overall semantic features of the input text. Figure 6 As shown, the fourth target text description feature is pooled to obtain the second global feature s4'. The second concatenated feature s3' and the second global feature s4' are then fused into the second global semantic feature s5'.

[0068] In the embodiment of the present application, two different CLIP models are used to extract features from text, and further pooling is performed to obtain global semantic features. This dual encoding strategy can enhance the model's ability to understand text semantics and improve the expression effect of complex scenes. It can also combine the multimodal alignment information of text and image, thereby improving the quality and consistency of the synthesized image.

[0069] In one embodiment of the present application, a real scene image is subjected to noise processing to obtain noise data, including: Perform dimensionality reduction processing on real scene images to obtain dimensionality reduction feature data; The dimension-reduced feature data is subjected to noise processing to obtain noise data.

[0070] In this embodiment, during the training process of the model, directly modeling high-dimensional data often leads to a sharp increase in computational complexity and increases the difficulty of optimization. Therefore, the real scene image is first subjected to dimensionality reduction processing, and after obtaining dimensionality reduction feature data, noise is added based on the dimensionality reduction feature data, and then the model training is continued.

[0071] In one embodiment of the present application, a dimensionality reduction process is performed on a real scene image to obtain dimensionality reduction feature data, including: The probability distribution parameters of the real scene image in the latent space corresponding to the mining area data generation model are obtained through the VAE encoder. The probability distribution parameters include mean and variance. According to the probability distribution parameters and standard Gaussian noise, the latent variables are calculated; The latent variables are adjusted based on the adjustment factors, and the adjusted latent variables are decoded to obtain the reduced dimension feature data.

[0072] In this embodiment, VAE (Variational Auto Encoder) is used to encode the real scene image and project it into the latent space of the model, thereby reducing the data dimension while retaining key semantic information, improving computational efficiency and enhancing the stability of the model. Specifically, the real scene image passes through the VAE encoder and learns the probability distribution parameters, i.e., the mean and variance , and then introduce standard Gaussian noise using reparameterization technique , I is the identity matrix, through the formula The calculation of the latent variable z retains the gradient information, making the optimization process more stable, and also allows the model to adaptively adjust the representation of the latent space during the learning process. The latent variable z is then adjusted by the adjustment factor s and sent to the decoder to obtain the reduced dimension feature data. The overall process is as follows: Figure 7 shown.

[0073] Figure 7 The right side of the figure shows the channel visualization of the latent variables, where the four channels correspond to different latent feature components generated by VAE. From the visualization, we can observe that the latent representation not only retains the main structural information of the original image, but also captures texture details and local features in some channels. This shows that VAE can still maintain key information while compressing data, allowing subsequent models to perform more stable reasoning based on this low-dimensional feature.

[0074] In one embodiment of the present application, the second control condition data includes: a second global semantic feature, a second space vector and a second fusion feature; based on the second control condition data and the noise data, a mining area data generation model is trained, including: Input the second control condition data and the noise data into a preset model, and output the predicted noise data; wherein the preset model includes an original neural network, a conditional control network, and a connection network, the original neural network and the conditional control network are connected through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; input the second control condition data and the noise data into the preset model, including: inputting the noise data into the first sub-network of the original neural network, and fusing the noise data and the second fusion feature and inputting them into the conditional control network, and inputting the second global semantic feature and the second space vector into each sub-network of the conditional control network and the original neural network respectively; Calculate the loss function based on the predicted noise data and the noise data; The parameters of the preset model are iteratively optimized based on the loss function to obtain the mining area data generation model.

[0075] In the embodiment of the present application, the preset model uses a control network (ControlNet) to inject additional control condition data into the original neural network. This strategy allows the diffusion model to rely on external structured inputs during the generation process, thereby enhancing the model's perception and control of input constraints, so that the image generation results can be guided more accurately. ControlNet adopts a freeze-copy strategy to ensure that the additional control condition data does not destroy the generation ability of the original neural network.

[0076] Specifically, if Figure 8 As shown in the figure, the original neural network is frozen, and a trainable copy, namely the conditional control network, is created, and the two are connected using a zero convolution layer. A zero convolution layer refers to a 1×1 convolution layer whose weights and biases are initialized to zero, that is, in an untrained state, this layer will not have any effect on the input data. In this way, in the early stages of training, ControlNet will not interfere with the generation process of the original neural network, thereby ensuring the stability of the model in the initialization stage and avoiding unnecessary deviations during the reasoning process. In addition, since the 1×1 zero convolution layer only performs linear transformations between channels, its computational cost is extremely low, but it can gradually adjust the weights during training, thereby learning how to incorporate control condition data into the generation process of the model.

[0077] In one embodiment of the present application, Figure 6 As shown, the noise data is input into the first sub-network of the original neural network, and the noise data and the second fusion feature are fused and input into the conditional control network, and the second global semantic feature and the second space vector are respectively input into each sub-network of the conditional control network and the original neural network, and the predicted noise data is output. Furthermore, based on the predicted noise data and the noise data, the loss function is calculated. In ControlNet, the loss function is optimized based on the original neural network, and the form is as follows:

[0078] Where E is the expected value of the loss function, is the noise data, To predict noisy data, is the initial dimension reduction feature data, is the dimension reduction feature data of the real scene image, represents the second global semantic feature and the second space vector in the second control condition data, represents the second fusion feature in the second control condition data, t is the time step of the denoising process, is the model parameter.

[0079] The parameters of the preset model are iteratively optimized based on the loss function until the iteration termination condition is met, and the mining area data generation model is obtained.

[0080] It should be noted that the data closed-loop system of this application can use real mining data to verify the generalization performance of the mining data generation model, and promptly feedback to the optimization process of the diffusion model and the conditional control network, continuously improve the quality of the mining data generation model, and ensure that data generation always remains consistent with the actual mining environment, ensuring the stability and continuous improvement of the training effect of the autonomous driving perception model.

[0081] In one embodiment of the present application, a cross-attention mechanism is introduced into each sub-network of the conditional control network and the original neural network.

[0082] In this embodiment, in order to ensure the effective transmission of control condition data, a cross-attention mechanism is introduced in each sub-network of the conditional control network and the original neural network, so that text embedding and spatial position information can comprehensively affect the diffusion process.

[0083] In the autonomous driving perception system, vehicle detection and tracking is one of the core tasks, and its accuracy directly affects the reliability of high-level decisions such as path planning, collision avoidance, and behavior prediction. However, related technologies often use a unified generation strategy for foreground and background areas, which may result in missing foreground vehicle details, blurred boundaries, and morphological distortion, thereby affecting the generalization ability of the perception model and its adaptability in real driving environments. To this end, in one embodiment of the present application, the method also includes: For a real scene image, a set of detection frames of the foreground area of ​​the real scene image at a time step is defined, as well as a loss weight matrix of the detection frame set. The initial value of the loss weight matrix is ​​set to an all-1 matrix. The loss weight matrix is ​​mapped to the latent space corresponding to the mining area data generation model using a bilinear interpolation method; During the model training process, the loss weight matrix is ​​dynamically adjusted; in the initial stage of model training, the loss weight matrix is ​​controlled to increase so as to preferentially capture the key features of the foreground area; in the later stage of model training, the loss weight matrix is ​​controlled to decrease so as to balance the features of the foreground area with the features of the background area of ​​the real scene image.

[0084] In this embodiment, a dynamic foreground weight loss mechanism is proposed. By adaptively adjusting the loss weight of the foreground area, the model is guided to gradually focus on the generation of foreground targets during the training process, thereby improving the quality and authenticity of the synthetic data. The core idea is to dynamically adjust the loss contribution of the foreground area and the background area according to the different stages of the training process to achieve focused learning of the foreground information while maintaining the integrity of the scene. The pseudo code is as follows: Fig. 9 shown.

[0085] For each input image x , there are different numbers of detection boxes, using a set To represent the time step t Define a loss weight matrix , its initial value is set to a matrix of all 1s, that is, all pixels have the same weight by default. During training, the optimization strength of the foreground area is dynamically adjusted through a two-stage cosine scheduling function.

[0086] Specifically, the initial stage of model training, i.e. When To train the threshold, the model has not yet fully learned the features of the foreground target, so a rapid growth strategy is adopted to rapidly increase the foreground weight, so as to prompt the model to preferentially capture the structure, edge and lighting features of the foreground vehicle. The cosine growth function is used for smooth regulation to ensure the stability of the training process and avoid the impact of sudden changes in gradient updates on model convergence.

[0087] As training progresses, the model gradually grasps the key features of the foreground area. If the foreground weight is still kept high in the subsequent training process, it may cause the model to overfit the foreground area, thus affecting the learning of background information. When , the foreground weight is gradually reduced to achieve a balance between foreground feature learning and background modeling.

[0088] The dynamic weight scheduling curve of this application is as follows Fig.10 As shown, the maximum loss weight value is set to , reaching its maximum value at 30% of the training cycle and then gradually decaying.

[0089] This strategy ensures that the model can effectively learn the global information of the scene while generating foreground targets, thereby improving adaptability in different environments. In addition, considering that the model's calculations are mainly performed in the latent space, in order to ensure a smooth transition of the loss weights, a bilinear interpolation method is used to map the foreground loss weight matrix to the latent space so that it remains consistent at different scales. This method can construct a smooth foreground optimization matrix, thereby enhancing the influence of foreground targets in the diffusion model.

[0090] In one embodiment of the present application, the final loss function of the model is:

[0091] Where E is the expected value of the loss function, is the loss weight matrix, is the set of detection boxes, is the noise data, To predict noisy data, is the dimensionality-reduced feature data of the real-scene image, c is the second control condition data, and t is the time step, which are model parameters.

[0092] In the embodiments of the present application, in order to further improve the fidelity of the foreground target, the loss weight is calculated using the foreground annotation box, and a dynamic foreground weight loss mechanism is adopted to adaptively adjust the optimization intensity of the foreground area, so that the model can reasonably balance the generation quality of the foreground and the background at different stages. Finally, the model parameters are optimized by calculating the loss function, so that the model can strengthen the generation ability of the foreground details while maintaining the global structural consistency, thereby ensuring that the generated result not only conforms to the text description, but also has high visual fidelity and structural rationality.

[0093] As a specific implementation of the above-mentioned mining area data generation method, the embodiments of the present application provide a mining area data generation device. As Fig.11 shown, the mining area data generation device 1100 includes: a first data acquisition module 1101, a first control condition extraction module 1102, a noise data sampling module 1103, and a data generation module 1104.

[0094] Among them, the first data acquisition module 1101 is used to acquire the simulated mining area data for the mining area scene, and the simulated mining area data includes the virtual scene image and the label information of the virtual scene image; The first control condition extraction module 1102 is used to extract the first control condition data of the virtual scene image; The noise data sampling module 1103 is used to sample random noise data; The data generation module 1104 is used to input the first control condition data and the random noise data into the trained mining area data generation model to obtain the target mining area data; wherein, the mining area data generation model is trained based on the real mining area data, and the target mining area data is used to control the autonomous driving of the vehicle in the mining area scene.

[0095] Further, the first data acquisition module 1101 is specifically used for: Construct a mining area virtual environment based on the mining area acquisition data and the digital environment model; Perform vehicle power simulation processing and sensor simulation processing in the mining area virtual environment to obtain a virtual vehicle moving in the mining area virtual environment and sensors operating on the virtual vehicle; Construct diverse mining area scenes based on the mining area virtual environment, and obtain the simulated mining area data collected by the sensors of the virtual vehicle in the mining area scene. The mining area scene includes at least one of the following: long-tail scene, complex working condition scene, and extreme climate scene.

[0096] Furthermore, the first control condition extraction module 1102 is specifically used to: Convert the label information of the virtual scene image into a first text description feature, and perform feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature; Performing position embedding processing on the spatial information of the virtual scene image to obtain a first spatial vector, the spatial information including resolution information and position information; Performing encoding processing and feature fusion processing on the segmentation mask data of the virtual scene image to obtain a first fusion feature; According to the first global semantic feature, the first space vector and the first fusion feature, first control condition data are obtained, and the first control condition data are mapped to a latent space corresponding to the mining area data generation model.

[0097] Furthermore, the label information includes a category label and detection box geometry information; the first control condition extraction module 1102 is specifically used to: Converting the category label and the detection frame geometry information of the virtual scene image into a first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; The first text description feature is segmented by a first CLIP segmenter to obtain a first target text description feature, and the first target text description feature is encoded by a first CLIP encoder to obtain a first deep feature; The first text description feature is segmented by a second CLIP segmenter to obtain a second target text description feature, and the second target text description feature is encoded by a second CLIP encoder to obtain a second deep feature; Perform feature splicing processing on the first deep feature and the second deep feature to obtain a first splicing feature; Performing pooling processing on the second target text description feature to obtain the first global feature; Obtaining a first global semantic feature according to the first concatenated feature and the first global feature; The word segmentation process includes: splitting the first text description feature into sub-word units, adding text tags to the sub-word units, and performing length unification process on the sub-word units after adding the text tags to obtain a target text description feature with a unified length.

[0098] Furthermore, the mining area data generation model includes an original neural network, a conditional control network and a connection network, the original neural network and the conditional control network are connected through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; the first control condition data includes: a first global semantic feature, a first space vector and a first fusion feature; The data generation module 1104 is specifically used for: The first control condition data and the random noise data are input into the mining area data generation model, and the target mining area data is output; wherein, the first control condition data and the random noise data are input into the mining area data generation model, including: inputting the random noise data into the first sub-network of the original neural network, and fusing the random noise data and the first fusion feature and inputting them into the conditional control network, and inputting the first global semantic feature and the first space vector into each sub-network of the conditional control network and the original neural network, respectively.

[0099] Furthermore, the device also includes: A mining area data generation model training module is used to obtain real mining area data in a mining area scenario, the real mining area data includes a real scene image and label information of the real scene image, and extract the second control condition data of the real scene image; Perform noise processing on the real scene image to obtain noise data; Based on the second control condition data and the noise data, a mining area data generation model is trained.

[0100] The mining area data generating device 1100 in the embodiment of the present application may be a computer device, or a component in a computer device, such as an integrated circuit or a chip. The mining area data generating device 1100 provided in the embodiment of the present application can realize Figure 1 To avoid repetition, the various processes implemented in the mining area data generation method embodiment are not described here.

[0101] As a specific implementation of the above-mentioned mining area data generation model training method, the present application embodiment provides a mining area data generation model training device. Fig.12 As shown, the mining area data generation model training device 1200 includes: a second data acquisition module 1201, a second control condition extraction module 1202, a data noise addition module 1203 and a model training module 1204.

[0102] The second data acquisition module 1201 is used to acquire real mining area data in a mining area scene, where the real mining area data includes a real scene image and label information of the real scene image; A second control condition extraction module 1202, used to extract second control condition data of a real scene image; The data noise adding module 1203 is used to perform noise adding processing on the real scene image to obtain noise data; The model training module 1204 is used to train a mining area data generation model based on the second control condition data and the noise data. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.

[0103] Furthermore, the second control condition extraction module 1202 is specifically used for: Convert the label information of the real scene image into a second text description feature, and perform feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature; Performing position embedding processing on the spatial information of the real scene image to obtain a second spatial vector, wherein the spatial information includes resolution information and position information; Performing encoding processing and feature fusion processing on the segmentation mask data of the real scene image to obtain a second fusion feature; According to the second global semantic feature, the second space vector and the second fusion feature, the second control condition data is obtained, and the second control condition data is mapped to the latent space corresponding to the mining area data generation model.

[0104] Furthermore, the label information includes a category label and detection box geometry information; the second control condition extraction module 1202 is specifically used to: Convert the category label and detection frame geometry information of the real scene image into a second text description feature, where the second text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; The second text description feature is segmented by the first CLIP segmenter to obtain a third target text description feature, and the second target text description feature is encoded by the first CLIP encoder to obtain a third deep feature; The second text description feature is segmented by a second CLIP segmenter to obtain a fourth target text description feature, and the fourth target text description feature is encoded by a second CLIP encoder to obtain a fourth deep feature; Perform feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature; Performing pooling processing on the fourth target text description feature to obtain a second global feature; Obtain a second global semantic feature according to the second concatenated feature and the second global feature; The word segmentation process includes: splitting the second text description feature into sub-word units, adding text tags to the sub-word units, and performing length unification process on the sub-word units after adding the text tags to obtain a target text description feature with a unified length.

[0105] Furthermore, the data noise adding module 1203 is specifically used for: Perform dimensionality reduction processing on real scene images to obtain dimensionality reduction feature data; The dimension-reduced feature data is subjected to noise processing to obtain noise data.

[0106] Furthermore, the data noise adding module 1203 is specifically used for: The probability distribution parameters of the real scene image in the latent space corresponding to the mining area data generation model are obtained through the VAE encoder. The probability distribution parameters include mean and variance. According to the probability distribution parameters and standard Gaussian noise, the latent variables are calculated; The latent variables are adjusted based on the adjustment factors, and the adjusted latent variables are decoded to obtain the reduced dimension feature data.

[0107] Furthermore, the second control condition data includes: a second global semantic feature, a second space vector and a second fusion feature; the model training module 1204 is specifically used for: Input the second control condition data and the noise data into a preset model, and output the predicted noise data; wherein the preset model includes an original neural network, a conditional control network, and a connection network, the original neural network and the conditional control network are connected through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; input the second control condition data and the noise data into the preset model, including: inputting the noise data into the first sub-network of the original neural network, and fusing the noise data and the second fusion feature and inputting them into the conditional control network, and inputting the second global semantic feature and the second space vector into each sub-network of the conditional control network and the original neural network respectively; Calculate the loss function based on the predicted noise data and the noise data; The parameters of the preset model are iteratively optimized based on the loss function to obtain the mining area data generation model.

[0108] Furthermore, a cross-attention mechanism is introduced in each sub-network of the conditional control network and the original neural network.

[0109] Furthermore, the model training module 1204 is also used to: For a real scene image, a set of detection frames of the foreground area of ​​the real scene image at a time step is defined, as well as a loss weight matrix of the detection frame set. The initial value of the loss weight matrix is ​​set to an all-1 matrix. The loss weight matrix is ​​mapped to the latent space corresponding to the mining area data generation model using a bilinear interpolation method; During the model training process, the loss weight matrix is ​​dynamically adjusted; in the initial stage of model training, the loss weight matrix is ​​controlled to increase so as to preferentially capture the key features of the foreground area; in the later stage of model training, the loss weight matrix is ​​controlled to decrease so as to balance the features of the foreground area with the features of the background area of ​​the real scene image; The loss function is:

[0110] Where E is the expected value of the loss function, is the loss weight matrix, is the set of detection boxes, is the noise data, To predict noisy data, is the dimension reduction feature data of the real scene image, c is the second control condition data, t is the time step, is the model parameter.

[0111] The mining area data generation model training device 1200 in the embodiment of the present application can be a computer device, or a component in a computer device, such as an integrated circuit or a chip. The mining area data generation model training device 1200 provided in the embodiment of the present application can realize Figure 5 To avoid repetition, the various processes implemented in the mining area data generation model training method embodiment will not be repeated here.

[0112] The present application also provides a computer device, such as Fig.13 As shown, the computer device 1300 includes a first processor 1301 and a first memory 1302. The first memory 1302 stores programs or instructions that can be executed on the first processor 1301. When the program or instructions are executed by the first processor 1301, the various steps of the above-mentioned mining area data generation method or mining area data generation model training method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not repeated here.

[0113] The first memory 1302 can be used to store software programs and various data. The first memory 1302 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instruction required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the first memory 1302 may include a volatile memory or a non-volatile memory, or the first memory 1302 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The first memory 1302 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0114] The first processor 1301 may include one or more processing units; optionally, the first processor 1301 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the first processor 1301.

[0115] The present application also provides a readable storage medium. Fig.14 As shown, a program or instruction 1401 is stored on the readable storage medium 1400. When the program or instruction 1401 is executed by the processor, each process of the above-mentioned mining area data generation method or mining area data generation model training method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0116] The methods described in the above embodiments may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. The computer-readable storage medium 1400 may include computer storage media and communication media, and may also include any medium that can transfer a computer program from one place to another. The storage medium may be any target medium that can be accessed by a computer.

[0117] As a possible design, the computer-readable storage medium 1400 may include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM or other optical disc storage; the computer-readable storage medium may include a magnetic disk storage or other magnetic disk storage device. Moreover, any connecting line may also be appropriately referred to as a computer-readable storage medium. For example, if the software is transmitted from a website, server or other remote source using a coaxial cable, a fiber optic cable, a twisted pair, a DSL (Digital Subscriber Line) or wireless technology (such as infrared, radio and microwave), the coaxial cable, the fiber optic cable, the twisted pair, the DSL or wireless technology such as infrared, radio and microwave is included in the definition of the medium. Disks and optical discs as used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks usually reproduce data magnetically, while optical discs reproduce data optically using lasers.

[0118] The present application also provides a chip, such as Fig.15 As shown, the chip 1500 includes at least one second processor 1501 and a communication interface 1502. The communication interface 1502 is coupled to the second processor 1501. The second processor 1501 is used to run programs or instructions to implement the various processes of the above-mentioned mining area data generation method or mining area data generation model training method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0119] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0120] Preferably, the chip 1500 further includes a memory, such as a second memory 1503 , and the second memory 1503 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.

[0121] In an embodiment of the present application, the second memory 1503 may include a read-only memory and a random access memory, and provide instructions and data to the second processor 1501. A part of the second memory 1503 may further include a non-volatile random access memory (NVRAM).

[0122] In an embodiment of the present application, the second processor 1501, the communication interface 1502, and the second memory 1503 are coupled together through a bus system 1504. Among them, in addition to including a data bus, the bus system 1504 may further include a power bus, a control bus, a status signal bus, etc. For the sake of description, in Fig.15 all kinds of buses are labeled as the bus system 1504.

[0123] The above-described method for updating the mining area map according to the embodiment of the present application can be applied to the second processor 1501 or implemented by the second processor 1501. The second processor 1501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the second processor 1501 or instructions in the form of software. The above-mentioned second processor 1501 may be a general-purpose processor (for example, a microprocessor or a conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The second processor 1501 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.

[0124] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0125] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for generating mining area data, characterized in that: include: Acquire simulated mining area data for a mining area scene, the simulated mining area data including a virtual scene image and label information of the virtual scene image, and extract first control condition data of the virtual scene image; Sample random noise data, and input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.

2. The method according to claim 1, characterized in that The obtaining of simulated mining area data for a mining area scenario includes: Construct a virtual environment of the mining area based on the mining area collection data and digital environment model; Performing vehicle dynamics simulation processing and sensor simulation processing in the virtual environment of the mining area to obtain a virtual vehicle moving in the virtual environment of the mining area and sensors operating on the virtual vehicle; A variety of mining scenes are constructed based on the mining virtual environment, and the simulated mining data collected by the sensors of the virtual vehicle in the mining scenes are obtained. The mining scenes include at least one of the following: long-tail scenes, complex working condition scenes, and extreme climate scenes.

3. The method according to claim 1, characterized in that The step of extracting the first control condition data of the virtual scene image comprises: Convert the label information of the virtual scene image into a first text description feature, and perform feature extraction and pooling processing on the first text description feature to obtain a first global semantic feature; Performing position embedding processing on the spatial information of the virtual scene image to obtain a first spatial vector, the spatial information including resolution information and position information; Performing encoding processing and feature fusion processing on the segmentation mask data of the virtual scene image to obtain a first fusion feature; The first control condition data is obtained according to the first global semantic feature, the first spatial vector and the first fusion feature, and the first control condition data is mapped to a latent space corresponding to the mining area data generation model.

4. The method according to claim 3, characterized in that The label information includes a category label and detection frame geometry information; the label information of the virtual scene image is converted into a first text description feature, and feature extraction and pooling processing are performed on the first text description feature to obtain a first global semantic feature, including: Converting the category label and the detection frame geometry information of the virtual scene image into the first text description feature, where the first text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; The first text description feature is segmented by a first CLIP segmenter to obtain a first target text description feature, and the first target text description feature is encoded by a first CLIP encoder to obtain a first deep feature; The first text description feature is segmented by a second CLIP segmenter to obtain a second target text description feature, and the second target text description feature is encoded by a second CLIP encoder to obtain a second deep feature; Performing feature splicing processing on the first deep feature and the second deep feature to obtain a first splicing feature; Performing pooling processing on the second target text description feature to obtain a first global feature; Obtaining the first global semantic feature according to the first concatenation feature and the first global feature; The word segmentation process includes: splitting the first text description feature into sub-word units, adding text tags to the sub-word units, and performing length unification process on the sub-word units after adding the text tags to obtain a target text description feature with a unified length.

5. The method according to claim 1, characterized in that The mining area data generation model includes an original neural network, a conditional control network and a connection network, wherein the original neural network and the conditional control network are connected via the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; The first control condition data includes: a first global semantic feature, a first space vector and a first fusion feature; The step of inputting the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data includes: The first control condition data and the random noise data are input into the mining area data generation model, and the target mining area data is output; wherein, the step of inputting the first control condition data and the random noise data into the mining area data generation model includes: inputting the random noise data into the first sub-network of the original neural network, and fusing the random noise data and the first fusion feature and inputting them into the conditional control network, and respectively inputting the first global semantic feature and the first space vector into the conditional control network and each of the sub-networks of the original neural network.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Acquire real mining area data in the mining area scene, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image; Performing noise processing on the real scene image to obtain noise data; Based on the second control condition data and the noise data, the mining area data generation model is trained.

7. A mining area data generation model training method, characterized in that: include: Acquire real mining area data in a mining area scene, the real mining area data including a real scene image and label information of the real scene image, and extract second control condition data of the real scene image; Performing noise processing on the real scene image to obtain noise data; Based on the second control condition data and the noise data, a mining area data generation model is trained, and the mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario, and the target mining area data is used to control the automatic driving of vehicles in the mining area scenario.

8. The method according to claim 7, characterized in that The step of extracting the second control condition data of the real scene image comprises: Converting the label information of the real scene image into a second text description feature, and performing feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature; Performing position embedding processing on the spatial information of the real scene image to obtain a second spatial vector, wherein the spatial information includes resolution information and position information; Performing encoding processing and feature fusion processing on the segmentation mask data of the real scene image to obtain a second fusion feature; The second control condition data is obtained according to the second global semantic feature, the second spatial vector and the second fusion feature, and the second control condition data is mapped to the latent space corresponding to the mining area data generation model.

9. The method according to claim 8, characterized in that The label information includes a category label and detection frame geometry information; the converting the label information of the real scene image into a second text description feature, and performing feature extraction and pooling processing on the second text description feature to obtain a second global semantic feature includes: Converting the category label and the detection frame geometry information of the real scene image into the second text description feature, where the second text description feature includes the category label, the detection frame geometry information, and the color, state, and background environment of the category label; The second text description feature is segmented by a first CLIP segmenter to obtain a third target text description feature, and the second target text description feature is encoded by a first CLIP encoder to obtain a third deep feature; The second text description feature is segmented by a second CLIP segmenter to obtain a fourth target text description feature, and the fourth target text description feature is encoded by a second CLIP encoder to obtain a fourth deep feature; Performing feature splicing processing on the third deep feature and the fourth deep feature to obtain a second splicing feature; Performing pooling processing on the fourth target text description feature to obtain a second global feature; Obtaining the second global semantic feature according to the second concatenation feature and the second global feature; The word segmentation process includes: splitting the second text description feature into sub-word units, adding text tags to the sub-word units, and performing length unification process on the sub-word units after adding the text tags to obtain a target text description feature with a unified length.

10. The method according to claim 7, characterized in that The step of performing noise processing on the real scene image to obtain noise data includes: Performing dimensionality reduction processing on the real scene image to obtain dimensionality reduction feature data; The dimension-reduced feature data is subjected to noise processing to obtain the noise data.

11. The method according to claim 10, characterized in that The performing dimensionality reduction processing on the real scene image to obtain dimensionality reduction feature data includes: Obtaining probability distribution parameters of the real scene image in the latent space corresponding to the mining area data generation model through a VAE encoder, wherein the probability distribution parameters include a mean and a variance; According to the probability distribution parameters and standard Gaussian noise, latent variables are calculated; The latent variables are adjusted based on the adjustment factors, and the adjusted latent variables are decoded to obtain the dimension reduction feature data.

12. The method according to claim 7, characterized in that The second control condition data includes: a second global semantic feature, a second space vector and a second fusion feature; The training of a mining area data generation model based on the second control condition data and the noise data includes: Inputting the second control condition data and the noise data into a preset model, and outputting predicted noise data; wherein the preset model includes an original neural network, a conditional control network, and a connection network, the original neural network is connected to the conditional control network through the connection network, and the original neural network includes a plurality of sub-networks connected in sequence; inputting the second control condition data and the noise data into the preset model, including: inputting the noise data into the first sub-network of the original neural network, and fusing the noise data and the second fusion feature and inputting them into the conditional control network, and respectively inputting the second global semantic feature and the second space vector into each of the sub-networks of the conditional control network and the original neural network; Calculating a loss function based on the predicted noise data and the noise data; The parameters of the preset model are iteratively optimized based on the loss function to obtain the mining area data generation model.

13. The method according to claim 12, characterized in that A cross attention mechanism is introduced into each of the sub-networks of the conditional control network and the original neural network.

14. The method according to claim 12, characterized in that The method further comprises: For the real scene image, define a detection frame set of the foreground area of ​​the real scene image at a time step, and define a loss weight matrix of the detection frame set, wherein an initial value of the loss weight matrix is ​​set to an all-1 matrix; Using a bilinear interpolation method, the loss weight matrix is ​​mapped to a latent space corresponding to the mining area data generation model; During the model training process, the loss weight matrix is ​​dynamically adjusted; wherein, in the initial stage of the model training, the loss weight matrix is ​​controlled to increase so as to preferentially capture the key features of the foreground area; in the later stage of the model training, the loss weight matrix is ​​controlled to decrease so as to balance the features of the foreground area with the features of the background area of ​​the real scene image; The loss function is: Where E is the expected value of the loss function, is the loss weight matrix, is the detection box set, is the noise data, is the predicted noise data, is the dimension reduction feature data of the real scene image, c is the second control condition data, t is the time step, is the model parameter.

15. A mining area data generating device, characterized in that: include: A first data acquisition module, used to acquire simulated mining area data for a mining area scene, wherein the simulated mining area data includes a virtual scene image and label information of the virtual scene image; A first control condition extraction module, used to extract first control condition data of the virtual scene image; Noise data sampling module, used to sample random noise data; A data generation module is used to input the first control condition data and the random noise data into a trained mining area data generation model to obtain target mining area data; wherein the mining area data generation model is trained based on real mining area data, and the target mining area data is used to control the automatic driving of vehicles in the mining area scene.

16. A mining area data generation model training device, characterized in that: include: A second data acquisition module is used to acquire real mining area data in a mining area scene, wherein the real mining area data includes a real scene image and label information of the real scene image; A second control condition extraction module, used to extract second control condition data of the real scene image; A data noise adding module, used for performing noise adding processing on the real scene image to obtain noise data; A model training module is used to train a mining area data generation model based on the second control condition data and the noise data. The mining area data generation model is used to generate target mining area data based on the simulated mining area data in the mining area scenario. The target mining area data is used to control the automatic driving of the vehicle in the mining area scenario.

17. A computer device, characterized in that: It includes a first processor and a first memory, the first memory stores a program or instruction running on the first processor, and when the program or instruction is executed by the first processor, it implements the steps of the mining area data generation method as described in any one of claims 1 to 6, or implements the steps of the mining area data generation model training method as described in any one of claims 7 to 14.

18. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by the processor, the steps of the method for generating mining area data as described in any one of claims 1 to 6 are implemented, or the steps of the method for training a mining area data generation model as described in any one of claims 7 to 14 are implemented.

19. A chip, characterized in that: The chip includes at least one second processor and a communication interface, the communication interface is coupled to the at least one second processor, and the at least one second processor is used to run programs or instructions to implement the steps of the mining area data generation method as described in any one of claims 1 to 6, or to implement the steps of the mining area data generation model training method as described in any one of claims 7 to 14.

Citation Information

Patent Citations

  • Mine unmanned driving virtual scene generation method based on large model

    CN119442889A

  • Training method of image generation model and image generation method and device

    CN119648559A

  • Image generation method and device of driving scene, electronic equipment and storage medium

    CN119810239A

  • Vehicle autonomous driving perception self-learning method and apparatus, and electronic device

    EP4451231A1

  • Image classification method and apparatus, device, storage medium, and program product

    WO2024259886A1

Cited By

  • Mine safety training method and system based on digital twinning and virtual reality

    CN120781550A

  • Virtual sonar image generation method and system

    CN121414871A

  • Polarizer design method and device based on dual-path cooperative training

    CN121981052A

  • A polarizer design method and device based on double-path cooperative training

    CN121981052B