Virtual simulation data intelligent generation and enhancement method based on deep learning

By using deep learning-based generation and enhancement methods, the problems of low efficiency and insufficient diversity in simulation data generation have been solved, enabling efficient and diverse generation of simulation data and improving the performance of virtual simulation systems and AI models.

CN121859709APending Publication Date: 2026-04-14BEIJING JUNHE CHUANGXIANG TECH DEV CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, traditional simulation data generation is inefficient and lacks data diversity, and existing data augmentation techniques are difficult to significantly improve the generalization and robustness of downstream AI models.

Method used

A deep learning-based intelligent generation and enhancement method for virtual simulation data is adopted. Through multi-source heterogeneous data acquisition, deep generation model training and optimization, and data quality verification feedback, a closed loop is formed to generate new simulation data that conforms to a high-dimensional probability distribution and perform diverse enhancement operations.

Benefits of technology

It enables the efficient generation of highly realistic and diverse simulation data that meets specific requirements, significantly improving data quality and richness, reducing acquisition costs, and supporting virtual simulation systems and AI model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859709A_ABST
    Figure CN121859709A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual simulation data intelligent generation and enhancement method based on deep learning. The method comprises the following steps: step 1, collecting and converging multi-source heterogeneous simulation data; 2, simulation data preprocessing and feature extraction; step 3, constructing and configuring a depth generation model; step 4, iterative training and optimization of the depth generation model; step 5, intelligently generating and synthesizing new simulation data; step 6, generating data quality verification and enhancement feedback; step 7, carrying out controllability adjustment and directional enhancement on the generated data; and 8, model online learning and data generation closed-loop optimization are carried out. The method can automatically learn the internal law and distribution of the simulation data, efficiently generate highly-vivid and diversified new simulation data meeting specific requirements, and deeply enhance the existing limited data, thereby significantly reducing the simulation data acquisition cost, improving the data quality and richness, and improving the simulation data acquisition efficiency. And powerful data support is provided for development and test of a virtual simulation system and simulation-based AI model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual simulation technology, and more specifically to a method for intelligent generation and enhancement of virtual simulation data based on deep learning. Background Technology

[0002] High-quality, large-scale simulation data is crucial in the development, testing, and optimization of virtual simulation systems. Traditional simulation data generation relies primarily on accurate physical models and rules, but building high-fidelity models is time-consuming and labor-intensive, and it is difficult to cover all complex, extreme, or unknown conditions, resulting in insufficient data diversity and low generation efficiency. Furthermore, collecting sufficient data from specific scenarios in the real world is usually costly, time-consuming, and even poses security and ethical risks.

[0003] Existing data augmentation techniques mostly employ simple geometric transformations or noise additions, which have limited effects on enhancing the intrinsic features and complex distributions of data, making it difficult to significantly improve the generalization and robustness of downstream AI models. Summary of the Invention

[0004] To address this, the present invention provides a deep learning-based method for intelligent generation and enhancement of virtual simulation data, which solves the problem that existing technologies often employ simple geometric transformations or noise additions, resulting in limited enhancement effects on the intrinsic features and complex distributions of data, and making it difficult to significantly improve the generalization and robustness performance of downstream AI models.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A deep learning-based method for intelligent generation and enhancement of virtual simulation data includes the following steps:

[0007] Step 1, Multi-source heterogeneous simulation data acquisition and aggregation: Collect multi-modal raw simulation data, including time-series state data, environmental image data, 3D point cloud data and behavioral logic data, from the virtual simulation environment, historical simulation logs and associated physical sensing devices, and aggregate them in a unified manner;

[0008] Step 2, simulation data preprocessing and feature extraction: The aggregated multimodal raw simulation data is cleaned, aligned, normalized and labeled, and deep feature representations of the data are extracted using a feature extraction network to construct a normalized feature dataset;

[0009] Step 3, Deep Generative Model Construction and Configuration: Construct a deep generative model architecture with generative adversarial networks or variational autoencoders as the core, and configure the corresponding generator, discriminator or encoder-decoder network structure and initial parameters according to the modal characteristics of the target simulation data.

[0010] Step 4, deep generative model iterative training and optimization: The deep generative model is trained in multiple rounds using the normalized feature dataset. The model parameters are optimized by adversarial loss or reconstruction loss function so that the model learns the high-dimensional probability distribution features of the simulation data.

[0011] Step 5, Intelligent generation and synthesis of new simulation data: Using a trained deep generative model, new simulation data sequences or samples that conform to the learned data distribution are automatically generated by inputting random noise vectors or specific condition vectors.

[0012] Step 6, Data Quality Verification and Enhancement Feedback: Input the newly generated simulation data into the preset verification rule base or auxiliary discrimination model to evaluate its realism, rationality and diversity. Feedback the evaluation results to the model training stage to optimize the quality of the next round of generated data.

[0013] Preferably, in step two, cleaning and aligning the time-series state data includes: detecting and removing abnormal transition points, filling missing data segments with an interpolation algorithm based on context relevance, and unifying data streams with different sampling frequencies to the same time reference through resampling technology.

[0014] Preferably, in step three, when processing simulation data containing complex spatiotemporal correlations, the generator and discriminator of the deep generative model both adopt a hybrid network structure that combines a three-dimensional convolutional neural network with an attention mechanism and a long short-term memory network.

[0015] Preferably, step five further includes step fiveA, simulation data augmentation operation: applying various augmentation operations, including random scaling, rotation, translation, adding noise perturbation, style transfer and local feature replacement, to the original simulation data samples and the newly generated simulation data samples to expand the scale and diversity of the dataset.

[0016] Preferably, in step fiveA, the local feature replacement specifically involves: identifying specific object or scene region features in the simulation data sample, using the deep generation model to generate multiple variations of the object or region features, and replacing the corresponding parts in the original sample in a procedural manner to generate a new sample that is semantically consistent but has diverse details.

[0017] Preferably, in step six, the verification rule base includes constraint rules based on physical laws, logical rules based on domain knowledge, and distribution rules based on statistical features; the auxiliary discrimination model is a pre-trained deep classification network independent of the discriminator in the deep generation model, used to quantitatively score the authenticity of the generated data from multiple dimensions.

[0018] Preferably, the method further includes step seven, generating data controllability adjustment and targeted enhancement: providing a user interaction interface or parameter configuration interface to receive user-defined attribute constraints or target guidance instructions for the generated data; based on the instructions, adjusting the direction and weight of the input condition vector or latent space vector of the deep generative model to guide the model to generate targeted enhanced simulation data that conforms to the specific attributes, scenarios, or difficulty levels set by the user.

[0019] Preferably, the method further includes step eight, online model learning and data generation closed-loop optimization: dynamically expanding the validated high-quality generated data and new feedback data generated in actual simulation applications to the normalized feature dataset; based on the expanded dataset, periodic or triggered incremental training and fine-tuning of the deep generative model are performed so that the model can continuously adapt to new data distributions and simulation requirements, forming a self-evolving closed loop of data generation-validation-feedback-model optimization.

[0020] The present invention has the following advantages: It can automatically learn the inherent laws and distribution of simulation data, efficiently generate new simulation data that meet specific requirements, are highly realistic and diverse, and can deeply enhance existing limited data, thereby significantly reducing the cost of acquiring simulation data, improving data quality and richness, and providing strong data support for the development, testing and simulation-based AI model training of virtual simulation systems. Attached Figure Description

[0021] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).

[0022] Figure 1 A flowchart of a deep learning-based intelligent generation and enhancement method for virtual simulation data provided in an embodiment of this application. Detailed Implementation

[0023] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these embodiments are merely for further explanation of the present invention and should not be construed as limiting the scope of protection of the present invention. Technical engineers in the field can make some non-essential improvements and adjustments to the present invention based on the above-described content. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Please see Figure 1 A deep learning-based method for intelligent generation and enhancement of virtual simulation data includes the following steps:

[0025] Step 1, Multi-source heterogeneous simulation data acquisition and aggregation: Collect multi-modal raw simulation data, including time-series state data, environmental image data, 3D point cloud data and behavioral logic data, from the virtual simulation environment, historical simulation logs and associated physical sensing devices, and aggregate them in a unified manner;

[0026] Step 2, simulation data preprocessing and feature extraction: The aggregated multimodal raw simulation data is cleaned, aligned, normalized and labeled, and deep feature representations of the data are extracted using a feature extraction network to construct a normalized feature dataset;

[0027] Step 3, Deep Generative Model Construction and Configuration: Construct a deep generative model architecture with generative adversarial networks or variational autoencoders as the core, and configure the corresponding generator, discriminator or encoder-decoder network structure and initial parameters according to the modal characteristics of the target simulation data.

[0028] Step 4, deep generative model iterative training and optimization: The deep generative model is trained in multiple rounds using the normalized feature dataset. The model parameters are optimized by adversarial loss or reconstruction loss function so that the model learns the high-dimensional probability distribution features of the simulation data.

[0029] Step 5, Intelligent generation and synthesis of new simulation data: Using a trained deep generative model, new simulation data sequences or samples that conform to the learned data distribution are automatically generated by inputting random noise vectors or specific condition vectors.

[0030] Step 6, Data Quality Verification and Enhancement Feedback: Input the newly generated simulation data into the preset verification rule base or auxiliary discrimination model to evaluate its realism, rationality and diversity. Feedback the evaluation results to the model training stage to optimize the quality of the next round of generated data.

[0031] In implementing this invention, firstly, in step one, the system collects multimodal data such as time-series states, images, and point clouds from multiple sources, including virtual simulation environments, historical logs, and physical sensors. This convergence of heterogeneous data from multiple sources provides the model with comprehensive and three-dimensional learning materials, ensuring that the generated data can reflect the complex relationships of the real world and avoiding biases caused by a single data source. Subsequently, step two involves deep cleaning, alignment, and feature extraction of the raw data. This crucial preprocessing step eliminates noise and inconsistencies and extracts the essential features of the data, providing high-quality and standardized input for the subsequent deep learning model and directly determining the upper limit of the model's learning performance. In step three, based on the characteristics of the simulation data (such as the local correlation of images or the long-term dependence of time-series data), a deep generative model architecture centered on Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs) is flexibly constructed. This design enables the model to capture and reproduce high-dimensional, nonlinear data distributions.

[0032] Step four involves iteratively training the model using the processed feature dataset. Through continuous adversarial or reconstructive optimization, the model implicitly grasps the complex physical laws and statistical distributions behind the simulation data. This is a key step in achieving high-quality data generation. Step five involves inputting controllable noise or conditional signals, enabling the model to efficiently and automatically synthesize a large number of new simulation samples that are consistent with the distribution of real data but also possess diversity. This significantly overcomes the bottlenecks of traditional methods in terms of data output and scene coverage. Finally, step six introduces a quality verification and feedback mechanism. By using preset physical rules, knowledge logic, and additional discriminative models, the generated data is filtered and evaluated, and the evaluation results are fed back into the training process, forming a continuous improvement cycle. This ensures the realism and usability of the output data, effectively prevents model degradation or the generation of unreasonable data, and guarantees the reliability and practicality of the entire system.

[0033] In step two, cleaning and aligning the time-series state data includes: detecting and removing abnormal transition points, filling missing data segments with an interpolation algorithm based on context relevance, and unifying data streams with different sampling frequencies to the same time base through resampling technology.

[0034] In virtual simulations, time-series data (such as velocity and angle) output by sensors or simulation engines often suffers from anomalies and misalignments due to communication delays, packet loss, or differences in simulation step sizes. The aforementioned technical solutions effectively filter transient error signals by "detecting abnormal transition points"; employing "context-dependent interpolation algorithms" (e.g., Kalman filtering prediction and filling based on the motion trends of previous and subsequent time points) instead of simple linear interpolation allows for a more reasonable reconstruction of missing segments, maintaining physical continuity; and "resampling technology to unify the time base" ensures precise synchronization of multiple data streams on the time axis.

[0035] In step three, when processing simulation data containing complex spatiotemporal relationships, the generator and discriminator of the deep generative model both adopt a hybrid network structure that combines a three-dimensional convolutional neural network with an attention mechanism and a long short-term memory network.

[0036] Many simulation datasets (such as fluid motion videos and robotic arm motion point cloud sequences) contain both spatial structure and temporal evolution information. Simple convolutional neural networks excel at spatial feature extraction but struggle to model long-term temporal dependencies; long short-term memory networks can handle temporal relationships but are insensitive to spatial information. This solution combines both approaches and incorporates an attention mechanism, enabling the model to dynamically focus on key spatial regions at different time steps.

[0037] Step five also includes step fiveA, simulation data augmentation operation: applying various augmentation operations, including random scaling, rotation, translation, adding noise perturbation, style transfer and local feature replacement, to the original simulation data samples and the newly generated simulation data samples to expand the scale and diversity of the dataset.

[0038] This approach comprehensively utilizes geometric transformations (scaling, rotation, and translation) to alter the object's perspective and position; adds random noise to simulate sensor errors or environmental interference; leverages style transfer to change the global style of the data, such as lighting and texture; and finally, local feature replacement begins to enhance semantics. This combined method can expand the dataset at multiple levels, from pixel-level to feature-level, making it particularly suitable for situations with limited initial data. It can quickly provide more diverse samples for model training, improving the model's robustness.

[0039] In step 5A, the local feature replacement specifically involves: identifying specific object or scene region features in the simulation data sample, using the deep generation model to generate multiple variations of the object or region features, and replacing the corresponding parts in the original sample in a procedural manner to generate a new sample that is semantically consistent but has diverse details.

[0040] In practice, the system first identifies specific targets (such as a car or a pedestrian) in the original simulation scene and segments their regions. Then, using a pre-trained generative model (such as a model specifically designed for vehicle appearance deformation), it generates various variations of the target in terms of color, vehicle type, and degree of damage. Finally, these variations are "stitched" back into the original scene in a way that conforms to scene lighting, perspective, and occlusion relationships. For example, in a city scene, it can automatically replace cars on the roadside with cars or SUVs of different colors and models, thereby quickly generating street scenes containing a rich variety of vehicle types without the need to manually model each vehicle.

[0041] In step six, the verification rule base includes constraint rules based on physical laws, logical rules based on domain knowledge, and distribution rules based on statistical features; the auxiliary discrimination model is a pre-trained deep classification network independent of the discriminator in the deep generation model, used to quantitatively score the authenticity of the generated data from multiple dimensions.

[0042] Constraints based on physical laws (such as the motion of objects should conform to Newtonian mechanics and the law of conservation of energy) are hard conditions to ensure the physical rationality of data; logical rules based on domain knowledge (such as the correspondence between traffic light colors and vehicle behavior) ensure the logical consistency of the scene; and distribution rules based on statistical features (such as the size of generated objects should conform to the distribution of historical data) maintain the overall authenticity of the data from a macro perspective.

[0043] It also includes step seven, generating data controllability adjustment and targeted enhancement: providing a user interface or parameter configuration interface to receive user-defined attribute constraints or target-oriented instructions for the generated data; based on the instructions, adjusting the direction and weight of the input condition vector or latent space vector of the deep generative model to guide the model to generate targeted enhanced simulation data that conforms to the specific attributes, scenarios, or difficulty levels set by the user; this process allows for fine-grained manipulation of the generated data features without retraining the model.

[0044] Users can set parameters (such as "expect to generate driving scenarios with visibility less than 50 meters in foggy weather" or "increase the speed of the target object by 20%), and the system encodes these instructions into specific conditional vectors, or finds the vector direction of the corresponding attribute change in the model's latent space through calculation. Subsequently, when generating data, the model is guided by this conditional vector to produce simulation data that conforms to the user's instructions. This upgrades the method from a fully automated data factory to a customizable data workshop, capable of generating data specifically for testing specific extreme conditions or training AI models to solve specific tasks, greatly improving R&D and testing efficiency.

[0045] It also includes step eight, online model learning and data generation closed-loop optimization: the validated high-quality generated data and new feedback data generated in actual simulation applications are dynamically expanded to the normalized feature dataset; based on the expanded dataset, the deep generative model is periodically or triggered for incremental training and fine-tuning, so that the model can continuously adapt to new data distributions and simulation requirements, forming a self-evolving closed loop of "data generation-validation-feedback-model optimization".

[0046] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for intelligent generation and enhancement of virtual simulation data based on deep learning, characterized in that, Includes the following steps: Step 1, Multi-source heterogeneous simulation data acquisition and aggregation: Collect multi-modal raw simulation data, including time-series state data, environmental image data, 3D point cloud data and behavioral logic data, from the virtual simulation environment, historical simulation logs and associated physical sensing devices, and aggregate them in a unified manner; Step 2, simulation data preprocessing and feature extraction: The aggregated multimodal raw simulation data is cleaned, aligned, normalized and labeled, and deep feature representations of the data are extracted using a feature extraction network to construct a normalized feature dataset; Step 3, Deep Generative Model Construction and Configuration: Construct a deep generative model architecture with generative adversarial networks or variational autoencoders as the core, and configure the corresponding generator, discriminator or encoder-decoder network structure and initial parameters according to the modal characteristics of the target simulation data. Step 4, deep generative model iterative training and optimization: The deep generative model is trained in multiple rounds using the normalized feature dataset. The model parameters are optimized by adversarial loss or reconstruction loss function so that the model learns the high-dimensional probability distribution features of the simulation data. Step 5, Intelligent generation and synthesis of new simulation data: Using a trained deep generative model, new simulation data sequences or samples that conform to the learned data distribution are automatically generated by inputting random noise vectors or specific condition vectors. Step 6, Data Quality Verification and Enhancement Feedback: Input the newly generated simulation data into the preset verification rule base or auxiliary discrimination model to evaluate its realism, rationality and diversity. Feedback the evaluation results to the model training stage to optimize the quality of the next round of generated data.

2. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 1, characterized in that, In step two, cleaning and aligning the time-series state data includes: detecting and removing abnormal transition points, filling missing data segments with an interpolation algorithm based on context relevance, and unifying data streams with different sampling frequencies to the same time base through resampling technology.

3. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 1, characterized in that, In step three, when processing simulation data containing complex spatiotemporal relationships, the generator and discriminator of the deep generative model both adopt a hybrid network structure that combines a three-dimensional convolutional neural network with an attention mechanism and a long short-term memory network.

4. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 1, characterized in that, Step five also includes step fiveA, simulation data augmentation operation: applying various augmentation operations, including random scaling, rotation, translation, adding noise perturbation, style transfer and local feature replacement, to the original simulation data samples and the newly generated simulation data samples to expand the scale and diversity of the dataset.

5. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 4, characterized in that, In step 5A, the local feature replacement specifically involves: identifying specific object or scene region features in the simulation data sample, using the deep generation model to generate multiple variations of the object or region features, and replacing the corresponding parts in the original sample in a procedural manner to generate a new sample that is semantically consistent but has diverse details.

6. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 1, characterized in that, In step six, the verification rule base includes constraint rules based on physical laws, logical rules based on domain knowledge, and distribution rules based on statistical features; the auxiliary discrimination model is a pre-trained deep classification network independent of the discriminator in the deep generation model, used to quantitatively score the authenticity of the generated data from multiple dimensions.

7. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 1, characterized in that, It also includes step seven, generating data controllability adjustment and targeted enhancement: providing a user interaction interface or parameter configuration interface to receive user-defined attribute constraints or target guidance instructions for the generated data; based on the instructions, adjusting the direction and weight of the input condition vector or latent space vector of the deep generative model to guide the model to generate targeted enhanced simulation data that conforms to the specific attributes, scenarios, or difficulty levels set by the user.

8. The method for intelligent generation and enhancement of virtual simulation data based on deep learning according to claim 1, characterized in that, It also includes step eight, online model learning and data generation closed-loop optimization: the validated high-quality generated data and new feedback data generated in actual simulation applications are dynamically expanded to the normalized feature dataset; based on the expanded dataset, the deep generative model is periodically or triggered for incremental training and fine-tuning, so that the model can continuously adapt to new data distributions and simulation requirements, forming a self-evolving closed loop of data generation-validation-feedback-model optimization.

Citation Information

Cited By

  • A self-learning combustible explosive gas measuring method and system based on virtual mixed samples

    CN122157885A