Multi-view video data generation method for unstructured mining area scene
By constructing a semantic terrain world model and a scene script-driven video generation network, the problem of generating multi-view video data in unstructured mining areas was solved, achieving efficient generation of mining area video data that meets safety constraints and improving the robustness of the autonomous driving system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to generate video training data that satisfies the geometric constraints and temporal continuity of multiple cameras in unstructured mining scenarios. They also lack semantic terrain representation and hazardous area modeling suitable for open-pit mines, and cannot controllably synthesize multi-view video data according to indicators such as safe distance and slope.
By constructing a semantic terrain world model and combining it with a scene script-driven video generation network, multi-view driving videos of mining areas are generated. This includes obtaining semantic terrain BEV representations of the mining environment, generating mining scene scripts, training a Diffusion Transformer model, and introducing a danger zone priority masking mechanism to generate temporal semantic terrain BEV sequences and equipment trajectory sequences that meet script constraints.
It enables efficient generation of multi-view video data in unstructured mining areas, improves the robustness of autonomous driving systems in complex terrain and dangerous conditions, reduces the cost of high-risk real-vehicle data collection, and meets the actual needs of autonomous driving training data in mining areas.
Smart Images

Figure CN121904503A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and more particularly to a method for generating multi-view video data for unstructured mining scenarios. By constructing a semantic terrain world model and combining it with a scene script-driven video generation network, this invention can efficiently synthesize labeled driving video data in environments lacking regular roads and high-precision maps. This method can be used to train and evaluate autonomous driving perception and decision-making models, improving their safety and generalization capabilities in complex unstructured scenarios such as mining areas and construction sites. Background Technology
[0002] With the widespread application of deep learning in autonomous driving perception, perception networks are placing increasingly higher demands on data scale and distribution. Traditional methods rely on real-vehicle data collection and manual annotation to obtain multi-sensor data that is highly consistent with the real environment. However, this approach suffers from long collection cycles, high costs, significant safety risks, and difficulty in covering low-frequency, long-tail conditions such as accident scenarios, extreme weather, and limited field of view. To address the problem of "insufficient data, especially for long-tail scenarios," a series of "training data generation" technologies, such as simulation rendering and world model-driven video generation, are being developed. These technologies generate video data in batches within a controllable virtual environment for training and testing perception algorithms. In particular, multi-camera synchronized video and corresponding auxiliary annotations alleviate the pressure of real-vehicle data collection and improve the model's generalization ability to complex scenarios.
[0003] Chinese invention patent application, publication number CN120147787A, entitled "Training Data Generation Method, Apparatus, Storage Medium, and Electronic Device," proposes a technical approach to generate multiple training data sets by expanding and rewriting input instructions using language models, aiming to improve the diversity of training samples for specific tasks. This solution reduces the cost of manually designing samples by automatically constructing different forms of training instructions or label information. However, it primarily uses text or simple structured data and does not perform unified geometric and temporal modeling for video data from multiple cameras and multi-frame sequences used in autonomous driving. Furthermore, it does not introduce constraints such as sensor extrinsic parameters and synchronization relationships, thus failing to meet the "multi-view, long-duration, physically consistent" video generation requirements of autonomous driving perception.
[0004] Chinese invention patent application, publication number CN120411902A, entitled "End-to-End Autonomous Driving Control Method and Device Based on Multi-Camera Fusion," achieves target recognition and trajectory prediction of the surrounding environment by fusing images from forward and side cameras, and directly outputs control commands, thereby improving control performance under complex road conditions. This type of technical solution does not address how to automatically generate multi-view video training samples in a virtual environment, nor does it involve world modeling design for the synthesis process of training data.
[0005] It should be noted that existing technologies for generating a large amount of training data and controlling multiple cameras are mainly focused on structured scenarios such as urban roads or highways. Environmental representation typically relies on high-definition map elements such as lane lines, lane boundaries, and intersection topology, with the generated or perceived objects primarily being vehicles and pedestrians within regular lanes. For typical unstructured environments like open-pit mines, roads are often temporary or naturally formed, lacking stable lane structures. Three-dimensional terrain elements such as slopes, platforms, potholes, and material stockpiles are closely related to vehicle safety and frequently change with excavation and dumping operations. Existing environment representations based on lanes and high-definition maps are difficult to directly transfer. Furthermore, large operating equipment such as mining trucks, excavators, and loaders operate alongside pedestrians in confined spaces. Typical hazardous conditions are often triggered by geometric relationships such as "vehicles being too close to the edge of a slope" or "vehicles passing between potholes and material stockpiles," rather than traditional "illegal lane changes" or "running red lights." This places entirely different demands on the scenario construction and risk control of training data compared to structured roads.
[0006] Therefore, from the perspective of data generation requirements in autonomous driving perception, existing technologies still have significant shortcomings: On the one hand, language model-based training data generation methods mainly address the expansion of text or abstract annotations, failing to directly generate mining area video data that meets the geometric constraints and temporal continuity of multiple cameras; on the other hand, end-to-end control methods based on multi-camera fusion focus on improving online control performance using existing data, without addressing how to construct semantic terrain representations suitable for unstructured mining areas, nor providing a mechanism for batch generating multi-view video training data based on this representation using a world model. Especially in open-pit mining scenarios, there is a lack of a unified environmental representation describing terrain elements such as slopes, platforms, pits, and stockpiles, as well as the operational status of mining equipment. There is also a lack of masking modeling strategies for hazardous areas and key operational entities, and a lack of a complete solution for controllably synthesizing multi-camera video data according to indicators such as safe distance and slope through scene scripts or parameter configuration. Based on this, a video data generation method for unstructured mining areas is proposed. While maintaining compatibility with existing generation technologies, it introduces semantic terrain and operational entity modeling adapted to mining areas, and combines multi-view video generation capabilities to better meet the actual needs of constructing autonomous driving training data in mining areas. Summary of the Invention
[0007] Given the irregular road morphology, significant terrain undulations, and diverse and complex movement patterns of operating equipment in unstructured mining environments, existing data collection methods relying on actual vehicle data collection struggle to cover a large number of long-tail conditions while maintaining safety and controllability. Furthermore, traditional video generation methods based on structured roads and high-definition maps are difficult to directly adapt to mining scenarios. This invention proposes a video data generation method for unstructured mining scenarios. By constructing a semantic terrain world model and introducing a hazardous area priority modeling and scene script control mechanism, multi-view mining driving videos are automatically generated based on a unified semantic terrain BEV representation and mining scene map. This provides high-quality training data and test scenarios for autonomous driving perception and simulation in mining areas.
[0008] This invention provides a method for generating multi-view video data for unstructured mining area scenarios, comprising the following steps: S1, Obtain the semantic terrain BEV representation of the mining area environment. The semantic terrain BEV representation integrates the geometric features of the mining area terrain and the semantic category of the surface on the bird's-eye view grid, and generates a danger zone mask and a mining area scene map based on this. S2, Based on operational or testing requirements, generate a parameterized mining area scenario script describing the mining area operation process and safety conditions, and convert the mining area scenario script into a structured script object; S3. Based on the semantic terrain BEV representation, multi-view image sequence and mining area scene map obtained in step S1, train a semantic terrain world model with Diffusion Transformer as the backbone. The semantic terrain world model introduces a danger area priority masking mechanism during training. S4. The structured script object obtained in step S2 is used as a conditional input to the trained semantic terrain world model to generate a temporal semantic terrain BEV sequence and a device trajectory sequence that satisfy the script constraints. This sequence is then used to drive a multi-view rendering or decoding network to generate a multi-view mining area driving video sequence.
[0009] Optionally, the acquisition of the semantic terrain (BEV) representation of the mining area environment specifically includes: The collected multimodal sensor data are synchronized and registered in the same spatiotemporal coordinate system, and geometric features are statistically analyzed and encoded on the BEV grid of the bird's-eye view. Using a semantic segmentation network trained for mining scene, the semantic categories of slopes, platforms, stockpiles, pits, water bodies, passable areas and impassable areas in the mining area are mapped to the bird's-eye view BEV grid. Based on the semantic terrain BEV representation, a traffic connectivity analysis is performed, and based on the slope, elevation difference, and distance thresholds to the edge of dangerous terrain, dangerous area masks are generated for the slope danger zone, the boundary of the pit area, and the toe of the stockpile slope area. Static terrain blocks and dynamic operation entities are abstracted as nodes, and a mining scene map is introduced into the semantic terrain BEV representation.
[0010] Optionally, the mining scenario script in step S2 includes mining scenario scripts with site-level configuration, work area-level configuration, fleet dispatch-level configuration, and individual vehicle behavior-level configuration.
[0011] Optionally, in step S3, the training of the semantic terrain world model specifically includes: The semantic terrain BEV representation and the state of the mining area scene graph nodes are encoded into a spatiotemporal token sequence. The Diffusion Transformer denoising network and the conditional vector encoded according to the mining scene script parameters in step S2 are injected into the semantic terrain world model to be trained. Combining the danger zone mask generated in step S1, construct a danger zone weighted noise prediction loss function; A semantic terrain world model is obtained by training the model by minimizing the noise prediction loss weighted by danger zones.
[0012] Optionally, in step S4, generating a multi-view mining area driving video sequence specifically includes: The temporal semantic terrain BEV sequence generated by the semantic terrain world model and the device trajectory sequence are fused and encoded into a unified latent representation; A viewpoint encoding vector is assigned to each camera viewpoint, and combined with camera intrinsic and extrinsic parameters, the unified latent representation is mapped to a sequence of conditional features under each viewpoint. For each camera viewpoint, the unified latent representation, viewpoint encoding vector, and conditional feature sequence are input into the decoding network. During the decoding process, terrain height and device information are used for geometric alignment and occlusion inference to generate a latent feature sequence. The latent feature sequence is then restored into pixel-space video frames using a decoder. A multi-view video sequence of driving in the mining area was obtained from pixel-space video frames.
[0013] Optionally, when training the semantic terrain world model, a geometric consistency constraint between viewpoints is imposed on the decoding results of different camera viewpoints at the same time step, specifically including: The decoding results from different perspectives are inversely projected into the bird's-eye view BEV space, and the key structures are aligned with the semantic terrain BEV representation output by the semantic terrain world model. Inconsistent regions are penalized.
[0014] Optionally, it also includes step S5: performing quality assessment and screening on the multi-view mining area videos generated in step S4, and constructing a training dataset; The assessment includes image quality assessment, temporal coherence assessment, terrain structure consistency assessment, and safety constraint compliance assessment.
[0015] Optionally, each discrete token in the spatiotemporal token sequence includes spatial coordinates, a time step index, and geometric and semantic features.
[0016] Optionally, the safety constraints of the mining area scenario script include a slope safety distance threshold, minimum passing distance, maximum vehicle speed, and loading cycle.
[0017] Compared with the prior art, the present invention has at least the following beneficial effects: 1. A unified representation method for semantic terrain BEV and scene graphs adapted to unstructured mining environments is proposed. Instead of relying on structured road elements such as lane lines and intersections, it models based on core mining elements such as slopes, platforms, stockpiles, potholes and passable areas. This method can accurately depict the relationship between the three-dimensional terrain and working space of the mining area, providing a unified environmental expression for the world model and scene script.
[0018] 2. Introducing a priority masking strategy centered on hazardous areas and key operational entities during world model training significantly enhances the model's generation capabilities in critical slope zones, pothole boundaries, material stockpile toe areas, and multi-vehicle intersection areas, enabling generated videos to simultaneously meet geometric rationality and safety constraints in high-risk areas.
[0019] 3. Construct scenario scripts and condition control mechanisms for mining operation processes, transforming operational or testing requirements into parameterized scripts. Generate multi-view videos that meet specified risk levels and operation modes through semantic terrain world models, enabling on-demand and controllable long-tail scene synthesis in mining areas. This reduces the cost of high-risk real-vehicle data collection and provides targeted data augmentation for scenarios with weak perception models, thereby effectively improving the robustness of the autonomous driving system in mining areas under complex terrain and dangerous conditions. Attached Figure Description
[0020] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.
[0021] Figure 1 This is a flowchart of the multi-view video data generation method for unstructured mining area scenarios according to the present invention. Detailed Implementation
[0022] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0023] A specific embodiment of the present invention, such as Figure 1 This paper provides a method for generating video data for unstructured mining areas, including the following steps: S1: Multimodal Data Acquisition and Semantic Terrain BEV Representation Construction in the Mining Area: Multimodal sensors are used to collect real mining area operation data, including multi-view images from surround-view cameras, point cloud data from vehicle-mounted LiDAR, and vehicle attitude and position information. The above data are synchronized and registered in a unified spatiotemporal coordinate system. Geometric attributes such as point cloud height, slope, and surface roughness, as well as surface type, are encoded on the bird's-eye view BEV grid. A semantic segmentation network specifically trained for mining scene is introduced into the bird's-eye view BEV grid to semantically label slopes, platforms, stockpiles, pits, water bodies, passable areas, and impassable areas, forming a semantic terrain BEV representation. Based on the semantic terrain BEV representation, a danger zone mask is obtained, and a mining scene map is introduced into the semantic terrain BEV representation.
[0024] Specifically, the following steps are included: S1.1: Collect multi-view images, point clouds, and vehicle attitude and position information of the mining area through multi-modal sensors mounted on the vehicle, and perform preprocessing and geometric statistics on the point clouds.
[0025] Specifically, the multimodal sensor includes forward and side cameras, a multi-line LiDAR, an inertial navigation unit, and a positioning module.
[0026] Furthermore, for the point cloud acquired by the multi-line lidar, the space is voxelized according to a preset resolution, dividing the point cloud into a three-dimensional voxel grid, and the number of points within each voxel is obtained. ,average height Height variance and slope obtained based on local plane fitting Furthermore, surface roughness is also calculated by statistically analyzing the height variations of the point cloud data. This invention uses these statistical measures to obtain explicit representations of slope steepness, platform elevation differences, material stockpile contours, and pit depths. Simultaneously, it utilizes height variance and point count thresholds to filter out outliers and anomalous echoes, providing a stable geometric basis for subsequent semantic segmentation.
[0027] The mining area point cloud preprocessing and geometric statistics are performed by voxelizing the lidar point cloud and statistically analyzing geometric features such as the number of points within the voxels, average height, height variance, and local slope in the 3D mesh to reflect slope steepness, platform height difference, material stockpile shape, and pit depth. Outliers and abnormal echoes are filtered out to provide a reliable geometric basis for subsequent semantic segmentation and hazardous area identification.
[0028] S1.2: Construct a semantic terrain segmentation network for mining area scenarios. Perform semantic annotation training on the original point cloud or camera images, and map categories such as "accessible compacted road surface", "loose gravel area", "platform edge", "soil dump slope", "top of material stockpile" and "potholes / water accumulation" to the bird's-eye view BEV grid to obtain semantic terrain BEV representation.
[0029] In practice, a semantic terrain segmentation network designed for mining areas is used to process new mining area data during the inference phase. Based on the registration relationship (spatial alignment) between point clouds and images, semantic results are obtained. These results are then projected or fused into a unified BEV coordinate system to form an initial semantic terrain BEV representation map. Each BEV unit contains: geometric features (height, slope, roughness, etc.), surface semantic category labels ("accessible compacted surface", "loose gravel area", "platform edge", "dump slope", "top of stockpile" and "potholes / water accumulation") and a mark indicating whether it belongs to an accessible area.
[0030] In the inference phase, the semantic segmentation network of this invention jointly analyzes point cloud and image data based on the registration relationship between the two data sources. For example, point cloud data provides geometric information about the mining area, while image data provides texture and color information. Through registration, the two can be combined to provide a more comprehensive terrain description, thereby assigning a semantic label to each BEV grid cell.
[0031] S1.3: Connectivity Analysis and Dangerous Area Mask Generation: Based on the Semantic Terrain BEV Representation Obtained in S1.2 Connectivity analysis is performed on the passable grid area to extract the main passageways and candidate parking areas. Simultaneously, based on slope thresholds, elevation difference thresholds, and distance thresholds from the platform edge, hazardous area masks are automatically generated for slope hazard zones, pothole boundaries, and the toe of the stockpile slope area. These masks will then serve as important inputs for subsequent dangerous area priority masking strategies and scenario script security constraints.
[0032] S1.4: Abstract the static terrain blocks and dynamic operation entities in the mining area into nodes to form a mining area scene graph, and introduce the mining area scene graph into the semantic terrain BEV representation.
[0033] Understandably, these nodes, by defining their spatial location, motion state, and relative relationship with other nodes, characterize the spatiotemporal dynamics and interactions of the mining scene, providing accurate spatial semantic information for subsequent world model training and scene script generation.
[0034] Furthermore, the platform is a flat area within the mining area used for operations or stockpiling materials.
[0035] Through step S1, the present invention constructs a semantic terrain representation in the BEV space that simultaneously includes geometric attributes, semantic categories, and danger zone markers, providing an input basis for the world model training in subsequent step S2 and the scene script constraints in step S3.
[0036] S2: Construction of mining scene script and structured script object input: Based on the needs of mining operators, test engineers or algorithm engineers, design mining scene script interfaces to describe mining operation processes and safety conditions. Define platform layout, slope shape, material stacking location, pothole distribution, vehicle driving route and timing, excavation and loading operation rhythm, etc. in parameterized form as mining scene scripts.
[0037] Furthermore, the mining area scenario script is automatically generated after parsing the natural language description using a large language model. The scenario script content is converted into structured conditions such as terrain parameters, hazardous area parameters, and the start and end positions, operating speed curves, meeting methods, and loading and unloading cycles of various operating entities. These are then combined with the semantic terrain BEV representation from step S1 and the mining area scenario. Figure 1 One-to-one correspondence provides high-level conditional guidance for the world model.
[0038] The specific steps are as follows: S2.1: Define the hierarchical structure of the mining scene script, dividing the mining scene script into four levels: site-level configuration, work area-level configuration, fleet scheduling level, and individual vehicle behavior level. Among them, the site-level configuration describes the platform topology, slope orientation, and material stacking area location; the work area-level configuration describes the loading area, unloading area, and temporary construction area; the fleet scheduling level describes the task allocation and departure rhythm of different vehicles; and the individual vehicle behavior level describes the detailed driving trajectory, speed curve, and passing strategy of each vehicle.
[0039] Furthermore, site-level configuration includes the number and relative positions of platforms, the elevation differences and slope orientations between platforms, and the geometric boundaries of the stockpiling area. Furthermore, the configuration at the work area level includes the spatial scope and passable directions of loading areas, unloading areas, and temporary enclosed areas; Furthermore, at the fleet dispatch level: the task cycle, departure interval, priority rules, etc. for each type of vehicle; Furthermore, at the individual vehicle behavior level: this includes each vehicle's driving path, target speed curve, oncoming traffic strategy (giving way, passing, etc.), and coordinated actions with excavators and loaders.
[0040] S2.2: Parameter template and constraint rule configuration, which defines parameter templates and safety constraints for each level in the script hierarchy of the mining area scene in step S2.1. When the user specifies the requirements of the mining area scene through the graphical interface or natural language, the script parsing module fills the parsing results into the parameter template and automatically checks whether the geometric relationship and safety constraints of the semantic terrain BEV representation are satisfied. Furthermore, the parameter template includes platform length, width, elevation difference range, slope safety distance threshold, stockpile height range, maximum vehicle speed, minimum passing distance, loading cycle, etc. S2.3: Parse the user's test requirements expressed in natural language (e.g., the requirements of the mine operator, test engineer or algorithm engineer) into a mine scene script; populate the mine scene script into S2.1 and S2.2 to generate a structured script object, and map the terrain parameters and dangerous area constraints involved in it into the conditional input of the world model.
[0041] For example, when a user describes the test requirements in natural language, such as "lay out a one-way transportation line at the edge of the platform, with two mining trucks meeting each other on a narrow section of road, and a steep slope adjacent to the road", the large language model parses the test requirement description into mining area scenario script parameters.
[0042] Furthermore, the consistency check between the mining scene script and the semantic terrain BEV representation involves using semantic terrain BEV and scene graphs to perform geometric consistency checks on the driving routes and work areas in the script before parsing the user's test requirements expressed in natural language into the mining scene script. This ensures that all planned paths fall within the passable area, all stopping points have sufficient safety distances, and automatically reverts to a feasible approximate configuration or prompts the user to adjust parameters when a conflict is detected.
[0043] S3: Based on the semantic terrain BEV representation, multi-view image sequence and mining area scene map obtained in step S1, the bird's-eye view BEV grid features, scene map nodes and their state evolution over time are encoded into a spatiotemporal token sequence. The joint distribution of the mining area scene over time is learned through a diffusion generation framework, and a semantic terrain world model with Diffusion Transformer as the backbone network is constructed.
[0044] Furthermore, a masking mechanism prioritizing hazardous areas is introduced during the training process of the semantic terrain world model. This mechanism applies high-frequency occlusion to discrete tokens near the critical distance of slopes, pothole boundaries, material pile toe, and the boundary between vehicles and hazardous terrain, and increases the corresponding masking reconstruction loss. This enables the semantic terrain world model to have stronger structural completion and motion prediction capabilities in these safety-critical areas, ensuring the generation quality and physical rationality in complex terrain and hazardous areas.
[0045] In this invention, the semantic terrain world model learns the joint distribution of the mining scene over time through a diffusion generation framework. This joint distribution is used in the inference phase to generate semantic terrain BEV sequences and device state trajectories for future multiple time steps from the initial state and the mining scene script, and serves as the semantic and geometric conditions for the multi-view video generation network. In the training phase, this joint distribution also serves as a reference frame for consistency constraints between views, guiding the multi-view video generation results to maintain consistency with the world model output in terms of spatial structure and motion relationships. S3 specifically includes the following steps: S3.1: The semantic terrain BEV representation obtained in step S1 The state of the mining scene graph nodes at the corresponding time step (including spatial location, orientation, speed, etc. of vehicles, excavators, loaders, etc.) is discretized and encoded, and each bird's-eye view BEV grid or mining scene graph node is mapped to the initial state low-dimensional vector of the spatiotemporal token sequence.
[0046] Specifically, for a length of The spatiotemporal token sequence of time steps, with an initial low-dimensional vector. The expression is:
[0047] in, It is the initial state. The low-dimensional vector representation of the n discrete tokens, including the nth Feature representation of the bird's-eye view BEV grid and mining scene graph nodes in a discrete token; For all bird's-eye view BEV mesh and scene graph nodes at time step The total number of discrete tokens. Each discrete token contains spatial coordinates, a time step index, and basic semantic features, which are used for subsequent spatiotemporal modeling.
[0048] It is understandable that the low-dimensional vector represents a clean state.
[0049] S3.2: Initial state low-dimensional vector of the spatiotemporal token sequence Apply forward diffusion to create a clean, low-dimensional vector of the initial state. The gradual perturbation results in a low-dimensional noise vector of states with different noise levels based on the diffusion time step. .
[0050] In a preferred embodiment, standard Gaussian diffusion is used, and the expression for the low-dimensional noise vector is:
[0051] in, Indicates the first A low-dimensional noise vector at each diffusion time step noise level; For the first The fidelity coefficient corresponding to each diffusion time step; Standard Gaussian noise; This represents a multivariate Gaussian distribution with a mean of 0 and a covariance equal to the identity matrix. Represents the identity matrix.
[0052] This invention allows for the selection of linear or cosine noise scheduling, enabling the focus on lower noise levels during the initial training phase and the gradual learning of reconstruction capabilities under high noise conditions during the later training phase.
[0053] S3.3: Inject the Diffusion Transformer denoising network and conditional mechanism into the semantic terrain world model to be trained; Furthermore, the Diffusion Transformer denoising network includes: Spatiotemporal token sequence embedding and position encoding: transforming noisy low-dimensional vectors Each discrete token is input into the embedding layer of the Diffusion Transformer denoising network to obtain initial features; the BEV coordinate encoding of the two-dimensional bird's-eye view and the temporal step position encoding are superimposed to enable the Diffusion Transformer denoising network to explicitly perceive spatial and temporal information.
[0054] Multilayer self-attention and feedforward network: Stack multiple layers of self-attention and feedforward sub-layers along the discrete token sequence to capture dependencies across time steps and regions through spatiotemporal self-attention; Furthermore, in multi-layer attention computation, local biases are applied to discrete tokens from the same time step or the same spatial neighborhood to enhance the modeling capability of local terrain structures.
[0055] Furthermore, the conditional mechanism is injected: the scenario script parameters (such as operation mode, vehicle scheduling strategy, hazard level requirements, etc.) in step S2 are encoded into a conditional vector cond; the conditional vector cond is mapped to a set of conditional vectors that match the number of layers of the Diffusion Transformer denoising network through the conditional encoder, and injected in each layer in the form of gating or cross-attention, so that noise prediction is explicitly controlled by the scenario script.
[0056] The final output of the Diffusion Transformer denoising network is a standard Gaussian noise. Prediction ,in, These are the trainable parameters of the Diffusion Transformer.
[0057] S3.4: Combine the danger zone mask obtained in step S1 During the training phase, a noise prediction loss weighted by the danger zone is constructed.
[0058] Specifically, discrete tokens belonging to hazardous areas are assigned higher weights, while discrete tokens in ordinary areas are assigned lower weights. This guides the Diffusion Transformer denoising network to focus on learning the fine structure and dynamics of areas such as the critical zone of the slope, the boundary of the pit, and the toe of the stockpile slope. The expression for the noise prediction loss weighted for hazardous areas is as follows:
[0059] in, This represents the total loss function, used during training to measure the difference between the model output and the real data; This represents the weighting coefficient for high-risk areas; in the loss function, discrete tokens in high-risk areas will have a larger weight. This allows the model to focus more on these high-risk areas; The weights of discrete tokens representing low-risk areas result in these areas receiving less attention. This is the real noise added to the i-th discrete token during the forward diffusion process. For the Diffusion Transformer denoising network, the noise prediction for the i-th discrete token is given. This is the real noise added to the j-th token during the forward diffusion process. The Diffusion Transformer denoising network predicts the noise for the j-th token. The pre-set weighting coefficients.
[0060] S3.5: Predicting loss by minimizing noise weighting in hazardous areas In conjunction with regularization terms, a semantic terrain world model with DiffusionTransformer denoising network and conditions is trained based on the semantic terrain BEV representation, multi-view image sequences, and mining scene map obtained in step S1; thus, a semantic terrain world model with stable performance on the overall terrain and higher reconstruction accuracy in dangerous areas is obtained.
[0061] After training, given the initial state and scene script conditions, the DiffusionTransformer denoising network can generate semantic terrain BEV sequences and device state trajectories for several future time steps through a backdiffusion process, providing a temporal semantic skeleton for multi-view video generation.
[0062] S4: Given an initial state, the structured script object obtained in step S2 and the semantic terrain world model trained in step S3 are used to perform temporal rolling prediction on the semantic terrain BEV representation and the mining scene map to generate a BEV sequence and a trajectory sequence of equipment (dynamically with a working entity) that satisfy the specified mining scene script and risk constraints (the aforementioned dangerous area mask and its weighting strategy and the safety constraints of step S2.2). Subsequently, using a multi-view rendering or decoding network, combined with camera extrinsic and intrinsic parameters and terrain height information, the generated BEV sequence and mining scene map information are mapped to multiple camera views to generate forward, lateral, and backward multi-view video frame sequences consistent with the actual equipment installation method in the mining area, ensuring the consistency of each view in terms of geometric structure and temporal continuity.
[0063] It is understandable that the initial state is: the semantic BEV layout of the main transportation road in the mining area and the unloading area, as well as the position, speed and status of the mining dump trucks, etc.
[0064] This invention generates a multi-view mining area video sequence based on a semantic terrain world model. On the basis of the "semantic terrain BEV representation + equipment status trajectory (spatiotemporal token sequence obtained in step S3)" generated in step S3, a conditional diffusion video generation network with Diffusion Transformer as the backbone is used to generate a multi-view mining area driving video sequence that is consistent with the actual sensor layout. S4 specifically includes the following steps: S4.1: Obtain the unified latent representation and view coding. The specific steps are as follows: Using a unified latent representation network, the multi-view mining video sequence generated by the semantic terrain world model (i.e., a BEV sequence and equipment trajectory sequence that meet the specified mining scene script and risk constraints, including geometric / semantic / risk information) is fused and encoded with the scene graph (nodes, edges, dynamic operation subject status) into a compact spatiotemporal feature sequence unified latent representation. Then, a view encoding vector is assigned to each camera view to describe the installation position, orientation and field of view of that view, realizing the view condition feature sequence of "one latent scene, corresponding to multiple view decoding".
[0065] It is understandable that the process of obtaining the viewpoint condition feature sequence is as follows: the potential scene is a general scene description that is the same for all cameras; viewpoint conditions are information about how the camera is installed and which direction it is looking; when the system attaches the viewpoint conditions to the general scene description, it obtains a scene description with a certain viewpoint limitation. The scene description with a certain viewpoint limitation is the viewpoint condition feature sequence because it is still the feature that changes over time, but it has been labeled / adjusted to be viewed from a certain viewpoint.
[0066] Optionally, the viewpoint encoding vector is constructed as follows: given the mining scene script and initial state, the length is obtained using the semantic terrain world model from step S2. Semantic terrain BEV representation and the corresponding time step device status set Construct the viewpoint encoding vector.
[0067] Specifically, for each time step and each camera perspective Using camera internal and external parameters 3D terrain points on semantic terrain BEV representation Projecting onto the image plane, we obtain the pixel coordinates corresponding to each camera's viewpoint. :
[0068] in, This indicates the depth of a 3D terrain point in the camera coordinate system. Indicates projection to the first The x-coordinate of pixels in a camera image; Indicates projection to the first The vertical coordinate of pixels in a camera image; express A camera intrinsic parameter matrix for each camera viewpoint is used to map 3D points in the camera coordinate system to pixel plane coordinates. Indicates the first Rotation matrix for each camera viewpoint; This represents the coordinates of the projected 3D terrain points in the corresponding 3D coordinate system of the BEV. Indicates the first Translation vector of each camera viewpoint.
[0069] It is understandable that the pixel coordinates are not the view encoding vector itself, but rather the result calculated by the semantic terrain BEV and the device state through projection under the camera geometry described by the given view encoding vector; the view encoding vector is used to determine "how to project", and the pixel coordinates are "the result after projection".
[0070] Using the aforementioned projection relationships, the height, semantic category information, device outline, and pose of the semantic terrain BEV representation are rasterized to the corresponding... Conditional feature map of each camera view In this process, a viewpoint conditional feature sequence is formed that is consistent with the camera's field of view and contains semantic and geometric priors.
[0071] S4.2: Multi-view decoding and geometric alignment. For each camera viewpoint, the unified latent representation is concatenated with its viewpoint encoding and input into the decoding network. Through a learnable projection module or an explicit geometric projection module, the BEV features are mapped onto the image plane under the camera's field of view. During the decoding process, terrain height information and device size information are explicitly used to infer the occlusion relationship between the front and back to obtain the latent features of consecutive frames in the pixel space, suppressing geometric inconsistencies such as interlacing and dangling.
[0072] Specifically, step S4.1 utilizes the unified generation of latent representations and viewpoint encoding to construct viewpoint conditional features aligned with the image plane under a given camera viewpoint. This process belongs to the conditional side's decoding and geometric alignment. The target video frame sequence used in the video diffusion process is used to train and model the latent spatial distribution, generating the target side's data source. These two belong to different functional levels and jointly support the conditional diffusion generation of multi-view videos.
[0073] Specifically, video diffusion means, for each camera From the perspective of the camera, all that is captured Original frame sequence Mapped to the latent space, the first image is obtained through a pre-trained image encoder or a simple convolutional coding network. All the perspectives of each camera Clean latent frame representation of a frame Then, forward diffusion occurs in the potential space. Indicates all The first frame frame.
[0074] Furthermore, forward diffusion in the latent space involves adding noise to the latent frame representation to obtain the camera... From the perspective of the first Time step of frame diffusion process Noise latent frame representation ; Specifically, discretization involves dividing the latent features of each noisy latent frame representation into several patch vectors, flattening these patch vectors along the time dimension, and obtaining the result that corresponds to the first frame. Time-of-frame camera Perspective-related spatiotemporal token sequence .
[0075] Furthermore, during the inference phase, an initial noisy discrete token sequence is sampled from a standard Gaussian distribution, and clean potential discrete tokens are gradually generated through back diffusion. These tokens are then restored by the decoder to pixel-space video frames to obtain the potential features of consecutive frames.
[0076] S4.3: Introduce temporal convolution or temporal attention modules into the multi-view rendering or decoding network to jointly model the latent features of consecutive frames. At the same time, appropriately simulate motion blur and illumination changes according to vehicle speed and camera exposure time to obtain temporally consistent latent feature sequences, which are then restored into pixel-space video frames by the decoder. This makes the generated multi-view video smoother and more natural in the temporal dimension and closer to the motion characteristics of real mining area video.
[0077] S4.4: Interview consistency constraint. During the training phase, structural consistency constraints are applied to the decoding results of different camera viewpoints at the same time to obtain multi-view mining area videos.
[0078] For example, the projected positions of the same obstacle are aligned in the BEV space, and the edge positions of the same stack of materials are constrained from multiple perspectives, thereby ensuring the spatial structural consistency of the generated multi-view videos. The decoding results at the same time include the latent features of consecutive frames in the pixel space in step S4.2 and the temporally consistent latent feature representation in step 4.3.
[0079] The specific steps are as follows: Geometric consistency constraints between viewpoints: For the decoding results of different camera viewpoints at the same time step, a monocular depth estimation or semantic segmentation network is used to back-project them into the BEV space and align them with the "semantic terrain BEV representation + equipment status" output by S3; compare the consistency of key structures such as slope edges, platform outlines, material stack boundaries and vehicle positions, and impose penalties on geometrically inconsistent areas.
[0080] After the above training, in the inference stage, only a new mining scene script and initial state need to be given: the semantic BEV sequence and device state trajectory are generated by the semantic terrain world model of S3; and then the multi-view DiffusionTransformer denoising network of S4 generates multi-view mining video under the constraints of the condition features of each view.
[0081] Furthermore, a multi-view Diffusion Transformer denoising network is used for each viewpoint. The noisy discrete token sequence and conditional features are input together into the view-level Diffusion Transformer denoising network: Add or concatenate the noise discrete token with the time step code and the camera view code to obtain an initial token vector containing "when and from which view". The conditional features obtained in step S4.1 The convolutional and linear projection methods map the data to a conditional token with the same dimension as the discrete token, which is then used as the key / value pair for cross-attention. In the multi-layer self-attention + cross-attention structure, the network captures the temporal consistency between frames within the same viewpoint on the one hand, and aligns the semantic terrain BEV with the device state to the image space through conditional tokens on the other hand.
[0082] S5: Video quality assessment and training dataset construction. For the multi-view mining area videos generated in step S4, image quality indicators, temporal coherence indicators, terrain structure consistency indicators, and safety constraint compliance indicators are introduced to comprehensively evaluate and screen the multi-view mining area videos generated in step S4, and construct perception for autonomous driving in mining areas.
[0083] Specifically, regarding terrain structure consistency, the system checks whether the slope contours, platform elevation differences, material stack boundaries, and pothole shapes in the video match the input terrain parameters and semantic terrain BEV representation. Regarding safety constraints, it assesses whether the minimum distance between the vehicle's trajectory and dangerous terrain features such as slopes and potholes meets the preset safety range. Multi-view video samples that reach the scoring threshold are marked as qualified samples. For scenarios requiring automatic annotation, the corresponding semantic terrain BEV sequence, equipment trajectory, and safety area mask can be exported simultaneously, constructing a high-quality dataset for autonomous driving perception and simulation training in mining areas.
[0084] Preferably, the specific steps are as follows: S5.1: Offline quality and structural consistency assessment. First, basic indicators such as image sharpness, noise level, and color stability are calculated for each generated video segment. Then, a pre-trained semantic segmentation or depth estimation network is used to reconstruct the generated frames. The reconstructed BEV structure is compared with the BEV sequence output by the original world model to check the consistency of key terrain structures such as slope contours, platform height differences, material stack boundaries, and potholes, and a comprehensive quality score is obtained. .
[0085] S5.2: Perception Model Playback and Contribution Evaluation. The generated video and a small number of real-world collected videos are input into the mining area perception model. The changes in detection accuracy, false detection rate, and false negative rate under different terrain categories and different working conditions are statistically analyzed. Based on this, the "contribution" of each type of generated sample to the model performance is evaluated and used to guide the sampling weights of the subsequent training set.
[0086] S5.3: Sample selection and construction of difficult example dataset. Based on the results of S5.1 and S5.2, the generated samples (generated multi-view mining area videos) are selected in a hierarchical manner: samples with low quality scores or negative impacts on the perception model are removed, samples with high quality and significant improvement effects on weak scenes of the model are retained, and the proportion of extreme terrain, dangerous working conditions and rare operation modes in the training data is appropriately increased to construct the "long-tail difficult example dataset of mining area".
[0087] Through the above steps, the multi-view video data generation method for unstructured mining area scenarios proposed in this invention utilizes a semantic terrain world model and a conditional diffusion video generation network with Diffusion Transformer as its backbone to achieve high-quality, multi-view, long-term sequence video synthesis under the premise of satisfying geometric and safety constraints. This effectively reduces the cost of real vehicle data collection and manual annotation, and significantly improves the robustness and generalization ability of the mining area autonomous driving system under complex long-tail conditions.
[0088] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating multi-view video data for unstructured mining area scenarios, characterized in that, Includes the following steps: S1, Obtain the semantic terrain BEV representation of the mining area environment. The semantic terrain BEV representation integrates the geometric features of the mining area terrain and the semantic category of the surface on the bird's-eye view grid, and generates a danger zone mask and a mining area scene map based on this. S2, Based on operational or testing requirements, generate a parameterized mining area scenario script describing the mining area operation process and safety conditions, and convert the mining area scenario script into a structured script object; S3. Based on the semantic terrain BEV representation, multi-view image sequence and mining area scene map obtained in step S1, train a semantic terrain world model with Diffusion Transformer as the backbone. The semantic terrain world model introduces a danger area priority masking mechanism during training. S4. The structured script object obtained in step S2 is used as a conditional input to the trained semantic terrain world model to generate a temporal semantic terrain BEV sequence and a device trajectory sequence that satisfy the script constraints. This sequence is then used to drive a multi-view rendering or decoding network to generate a multi-view mining area driving video sequence.
2. The method according to claim 1, characterized in that, In step S1, obtaining the semantic terrain (BEV) representation of the mining area environment specifically includes: The collected multimodal sensor data are synchronized and registered in the same spatiotemporal coordinate system, and geometric features are statistically analyzed and encoded on the BEV grid of the bird's-eye view. Using a semantic segmentation network trained for mining scene, the semantic categories of slopes, platforms, stockpiles, pits, water bodies, passable areas and impassable areas in the mining area are mapped to the BEV grid of the bird's-eye view. Based on the semantic terrain BEV representation, a traffic connectivity analysis is performed, and based on the slope, elevation difference, and distance thresholds to the edge of dangerous terrain, dangerous area masks are generated for the slope danger zone, the boundary of the pit area, and the toe of the stockpile slope. Static terrain blocks and dynamic operation entities are abstracted as nodes, and a mining scene map is introduced into the semantic terrain BEV representation.
3. The method according to claim 1, characterized in that, The mining scene script in step S2 includes mining scene scripts with site-level configuration, work area-level configuration, fleet dispatch-level configuration, and individual vehicle behavior-level configuration.
4. The method according to claim 1, characterized in that, In step S3, the training of the semantic terrain world model specifically includes: The semantic terrain BEV representation and the state of the mining area scene graph nodes are encoded into a spatiotemporal token sequence. The Diffusion Transformer denoising network and the conditional vector encoded according to the mining scene script parameters in step S2 are injected into the semantic terrain world model to be trained. Based on the dangerous area mask generated in step S1, construct a noise prediction loss function weighted by dangerous areas; A semantic terrain world model is obtained by training the model by minimizing the noise prediction loss weighted by danger zones.
5. The method according to claim 1, characterized in that, In step S4, generating the multi-view mining area driving video sequence specifically includes: The temporal semantic terrain BEV sequence generated by the semantic terrain world model and the device trajectory sequence are fused and encoded into a unified latent representation; A viewpoint encoding vector is assigned to each camera viewpoint, and combined with camera intrinsic and extrinsic parameters, the unified latent representation is mapped to a sequence of conditional features under each viewpoint. For each camera viewpoint, the unified latent representation, viewpoint encoding vector, and conditional feature sequence are input into the decoding network. During the decoding process, terrain height and device information are used for geometric alignment and occlusion inference to generate a latent feature sequence. The latent feature sequence is then restored into pixel-space video frames using a decoder. A multi-view video sequence of driving in the mining area was obtained from pixel-space video frames.
6. The method according to claim 5, characterized in that, When training the semantic terrain world model, a geometric consistency constraint between different camera views is applied to the decoding results at the same time step. Specifically, this includes: The decoding results from different perspectives are inversely projected into the bird's-eye view BEV space, and the key structures are aligned with the semantic terrain BEV representation output by the semantic terrain world model. Inconsistent regions are penalized.
7. The method according to claim 1, characterized in that, It also includes step S5: to conduct quality assessment and screening of the multi-view mining area videos generated in step S4, and to construct a training dataset; The assessment includes image quality assessment, temporal coherence assessment, terrain structure consistency assessment, and safety constraint compliance assessment.
8. The method according to claim 4, characterized in that, Each discrete token in the spatiotemporal token sequence contains spatial coordinates, a time step index, and geometric and semantic features.
9. The method according to claim 1, characterized in that, The safety constraints of the mining area scenario script include the slope safety distance threshold, minimum passing distance, maximum vehicle speed, and loading cycle.
Citation Information
Patent Citations
Training data generation method and device, storage medium and electronic equipment
CN120147787A
End-to-end automatic driving control method and device based on multi-camera fusion
CN120411902A
Cited By
Desert navigation method based on object recognition and terrain understanding and related device
CN122237629A