Intelligent environment analogue simulation method and system based on artificial intelligence
Through multimodal data preprocessing, deep deterministic strategy gradient update, VAE and U-Net diffusion model optimization, and meta-learning internal and external loop methods, the shortcomings of multimodal data fusion and task stratification in intelligent environment simulation are solved, and high-quality, semantic-consistent edge scene image sequences are achieved efficiently, which improves the robustness and adaptability of the simulation environment.
Patent Information
- Application Number
- CN202510644293.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing intelligent environment simulation technology lacks a unified semantic alignment mechanism when fusion of multimodal data, making it difficult to adapt to edge scenarios, and lacks systematicity in task hierarchy and parameter optimization, resulting in insufficient visual and semantic consistency and robustness of the generated scenes.
Through multimodal data preprocessing, deep deterministic strategy gradient update, VAE and U-Net diffusion model optimization, meta-learning internal and external loop methods, combined with Monte Carlo simulation and conditional diffusion model, high-quality, semantic consistent scene image sequences are generated, and through task hierarchy and adaptive parameter optimization, the generated scenes meet specific task requirements.
It significantly improves the generation efficiency and fidelity of edge scenes, enhances the semantic consistency and task adaptability of the simulation environment in complex interactive scenarios, and improves the detection accuracy and semantic alignment of the generated sequences.
Smart Images

Figure CN120493747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence environment simulation technology, and in particular to an artificial intelligence-based intelligent environment simulation method and system. Background Art
[0002] In recent years, intelligent environment simulation technology has made significant progress in areas such as autonomous driving, virtual testing, and complex scene analysis. With the rapid development of artificial intelligence (AI), deep learning-based simulation methods have gradually replaced traditional rule-based modeling approaches and are widely used in scenarios such as traffic scene simulation, industrial automation testing, and virtual reality. Methods combined with reinforcement learning (such as deep deterministic policy gradients (DDPG)) can dynamically optimize scene parameters, enhancing the interactivity and realism of simulation environments. Furthermore, the introduction of meta-learning techniques enables models to quickly adapt to new tasks, improving generalization capabilities in diverse scenarios. However, the comprehensive application of these technologies is still in the exploratory stage. Existing methods are particularly limited in their ability to integrate multimodal data, maintain dynamic consistency, and adapt to task-specific scenarios, especially when dealing with complex edge scenarios (such as severe weather and accidents). Despite the progress made in simulation, existing technologies still have several shortcomings, particularly in generating high-quality and robust edge scene image sequences. First, existing methods often lack a unified semantic alignment mechanism when fusing multimodal data, resulting in difficulties in balancing visual fidelity and semantic consistency in the generated scenes. For example, generative models based on GANs or single VAEs are not very effective in processing edge scenes. When dealing with dynamic scenes (such as vehicle movement in rainy days), problems such as inter-frame incoherence or loss of details are likely to occur. Secondly, existing technologies are not adaptable enough to edge scenarios. When faced with rare accidents or extreme weather, traditional models find it difficult to quickly adjust parameters to generate sequences that meet specific task requirements. In addition, existing methods lack systematicity in task layering and parameter optimization, and it is difficult to effectively balance multi-task requirements such as object detection, dynamic modeling, and global semantic understanding, resulting in poor robustness of the generated simulation environment in complex interactive scenarios. In contrast, our invention significantly improves the fidelity, dynamic consistency, and task adaptability of the generated sequences in edge scenarios through the efficient fusion of multimodal data, meta-learning-driven adaptive parameter optimization, and a diffusion model with joint loss optimization. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides an intelligent environment simulation method and system based on artificial intelligence to solve the problems of lack of a unified semantic alignment mechanism when fusing multimodal data, insufficient adaptability of existing technologies to edge scenarios, and lack of systematicity in task stratification and parameter optimization of existing methods.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] In a first aspect, the present invention provides an artificial intelligence-based intelligent environment simulation method, which includes collecting and preprocessing multimodal data, generating an initial scene based on the LLaMA-3-8B model, and updating the initial scene using a deep deterministic policy gradient;
[0007] Based on the updated initial scene and POMDP belief state set to form a scene tensor, the optimized VAE model is used to generate the initial image sequence, the U-Net diffusion model is used to optimize the initial image sequence, the VAE encoder reconstruction loss and the U-Net diffusion model diffusion denoising loss are calculated, and the joint loss function of the VAE encoder and the U-Net diffusion model is calculated. The joint loss function is used to generate a high-quality image sequence through the Monte Carlo simulation method, and the high-quality image sequence is optimized through the conditional diffusion model to generate the final scene image sequence;
[0008] Extract visual features from the final scene image sequence and use K-means clustering analysis to stratify edge cases into tasks based on visual features and semantic embeddings;
[0009] Adaptive parameters are created using the meta-learning inner loop method, optimized using the meta-learning outer loop method, and a single output sequence is generated using the optimized parameters. The generated single output sequence is then validated for quality indicators and tested for robustness.
[0010] As a preferred solution of the artificial intelligence-based intelligent environment simulation method described in the present invention, wherein: the multimodal data is collected and preprocessed, the initial scene is generated based on the LLaMA-3-8B model, and the initial scene is updated using deep deterministic policy gradient (DDPG) to collect multimodal data including text, images, lidar, GPS, temperature, humidity, video, and preprocess them, and the preprocessed data is generated using the CLIP-ViT-L-336px multimodal pre-trained model to generate a multimodal feature vector from the preprocessed data, the feature vector list is combined into a feature matrix A using the numpy.stack function, and stored in the Faiss vector database, the multimodal feature vector is trained based on the LLaMA-3-8B pre-trained language model, the initial scene is output in JSON format, the initial scene is parsed using the JSON parsing library, and the data is updated based on the LLaMA-3-8B pre-trained language model. The CARLA simulator dynamically generates a CARLA simulation environment vector based on the parsed JSON scene, uses the Transformer method to integrate the CARLA simulation environment vector with the multimodal feature vector into a feature vector set E, takes the feature vector set E as the state input, and uses the Actor and Critic neural network to configure a deep deterministic policy gradient architecture. Based on the state input feature vector set E, it outputs a continuous action data vector. By utilizing replay buffer transformations and soft targets, it optimizes the continuous action data vector output by the Actor-Critic network, and modifies the fields in the JSON initial scene according to the optimized action data vector to update the JSON initial scene.
[0011] As a preferred solution of the artificial intelligence-based intelligent environment simulation method of the present invention, wherein: the calculation of the VAE encoder reconstruction loss and the U-Net diffusion model diffusion denoising loss, and the calculation of the joint loss function, and the use of the joint loss function to generate a high-quality image sequence through the Monte Carlo simulation method refer to generating an initial image sequence based on the potential vector z through the VAE decoder (Decoder) , using the U-Net diffusion model to process the initial image sequence through the forward process Add Gaussian noise until it becomes pure noise, and train the reverse process to generate the optimized initial image sequence , based on the initial image sequence Calculating VAE encoder reconstruction loss and U-Net diffusion model denoising loss , based on the reconstruction loss and denoising loss Calculate the joint loss function of the VAE encoder and the U-Net diffusion model , using gradient descent and AdamW optimizer based on the joint loss function Calculate the weight set of VAE encoder and U-Net , using a weighted set via Monte Carlo simulation and the initial image sequence and scene tensors to generate high-quality image sequences , the final scene image sequence is generated through the conditional diffusion model , and obtain high-quality environmental status images.
[0012] As a preferred solution of the artificial intelligence-based intelligent environment simulation method of the present invention, wherein: extracting visual features from the final scene image sequence and using K-means clustering analysis to classify edge cases into task layers according to visual features and semantic embedding refers to using a pre-trained ResNet-50 model to extract visual features from the scene image sequence. Extract the high-dimensional visual feature vector f from the scene template set using the pre-trained CLIP-ViT-L-336px model Encode and obtain the semantic embedding vector from the encoding , using K-means cluster analysis based on the visual feature vector f and the semantic embedding vector , decomposing edge cases into three task layers .
[0013] As a preferred solution of the artificial intelligence-based intelligent environment simulation method of the present invention, wherein: the use of the meta-learning inner loop method to create adaptive parameters refers to the use of meta-model parameters , initialize the Transformer-UNet model by copying the meta-model parameters , for each task layer Creating task-specific parameters , using the nearest neighbor retrieval method for each task layer Retrieving a fixed sample support set , using PyTorch's AdamW optimizer to implement the meta-learning inner loop method, in which the scene image sequence is and the scene tensor Input the Transformer model and compare the output sequence o of the Transformer model with the scene template set Compare, calculate task-specific loss, perform a fixed number of gradient descent steps, and this iterative process generates adaptive parameters .
[0014] As a preferred solution of the artificial intelligence-based intelligent environment simulation method described in the present invention, wherein: the said optimizing the adaptive parameters by the meta-learning outer loop method and generating a single output sequence from the optimized parameters refers to implementing the meta-learning outer loop method in PyTorch using the AdamW optimizer, comparing the output sequence o with the structured label data, summarizing the performance of all task layers, calculating the combined error index, and updating the meta-model parameters by the outer loop method , while optimizing the adaptive parameters , the task-specific parameters are integrated into the parameter fusion module Encode into a unified control vector, and linearly concatenate the control vector with the scene image sequence and scene tensor Combined with the transformer layer of Transformer-UNet (TransUNet) injected as conditional input, it guides cross-modal fusion and spatial-semantic modeling in the model's attention mechanism, regulates the task-specific features of the generated sequence, and finally generates a single output image sequence y.
[0015] As a preferred solution of the artificial intelligence-based intelligent environment simulation method of the present invention, wherein: the quality index verification and robustness test of the generated single output sequence refers to comparing the single output image sequence y with the structured label data through the SSIM method, obtaining the SSIM structural similarity index, setting the threshold 、 as well as , pixel-level fidelity was evaluated using the scikit-image library in Python, image sharpness was measured using PSNR (OpenCV), and CLIP similarity (CLIP-ViT-L-336px) was used to evaluate the similarity with the scene tensor The similarity score of semantic alignment is obtained if the target similarity SSIM> , generating a PSNR score> , generating similarity scores> , then pass the verification, use the quality verified output sequence y, use the albumentations library to impose perturbations and use YOLOv8 to evaluate the task Target detection accuracy, generate mAP score, set threshold , so that the mAP score> , use OpenCV's Farneback method to calculate the task Optical flow field, obtain flow variance index, set threshold , so that the variance index ≤ , the CLIP model evaluates the task again Semantic similarity of the perturbed image, setting the threshold , so that similarity ,The indicators are summarized by weighted averaging to generate a robustness report that confirms the stability.,Using the output sequence y, the indicators and the robustness report, the,h5py library packages the output sequence y into an HDF5 simulation package,,ensuring deployment compatibility with CARLA.
[0016] In a second aspect, the present invention provides an intelligent environment simulation system based on artificial intelligence, including a multimodal data acquisition and initial scene generation module for providing multimodal input for scene generation, constructing a preliminary simulation scene, and providing a basis for subsequent dynamic optimization;
[0017] An initial image sequence generation module for sampling the latent representation into an initial image sequence through the VAE decoder;
[0018] A high-quality image sequence generation module is used to generate high-quality image sequences through Monte Carlo simulation, and then generate the final scene image sequence through the conditional diffusion model;
[0019] The meta-learning inner and outer loop modules are used to initialize the TransUNet model, generate adaptive parameters based on task-specific losses, and generate the final image sequence by comparing the output sequence with the structured label data through the outer loop;
[0020] The quality verification and robustness testing module is used to ensure the high quality and robustness of the final image sequence and supports CARLA deployment.
[0021] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the intelligent environment simulation method based on artificial intelligence as described in the first aspect of the present invention is implemented.
[0022] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the intelligent environment simulation method based on artificial intelligence as described in the first aspect of the present invention.
[0023] The beneficial effects of the present invention are: through unified semantic embedding and probability state distribution, efficient fusion of multimodal data is achieved, and the defects of inconsistency between vision and semantics in the existing technology are overcome. The generated scene tensor not only retains visual details, but also captures the probability distribution of dynamic scenes through belief states, thereby improving the semantic consistency of the simulation environment in complex interactive scenarios. The inner loop quickly generates task-specific parameters, and the outer loop optimizes the generalization ability of the metamodel. This method enables the model to quickly adapt to edge scenes. Compared with the traditional GAN model, it significantly improves the generation efficiency and fidelity of rare scenes. Monte Carlo simulation is combined with the conditional diffusion model to generate high-quality image sequences. Through the conditional guidance of CLIP embedding and belief states, the sequence is ensured to adapt to the dynamic characteristics of extreme scenes. Task stratification is combined with meta-learning to achieve multi-task collaborative optimization. The control vector enhances the adaptability of the model to specific tasks, significantly improving the detection accuracy and semantic alignment of the generated sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 This is a flow chart of an intelligent environment simulation method based on artificial intelligence in Example 1.
[0026] Figure 2 This is a structural diagram of an intelligent environment simulation system based on artificial intelligence in Example 1.
[0027] Figure 3 This is a diagram of the model training process of an artificial intelligence-based intelligent environment simulation method in Example 1.
[0028] Figure 4 This is a meta-learning framework diagram of an artificial intelligence-based intelligent environment simulation method in Example 1. DETAILED DESCRIPTION
[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0031] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0032] Example 1, with reference to Figures 1 to 4 , which is the first embodiment of the present invention, provides an intelligent environment simulation method based on artificial intelligence, comprising the following steps:
[0033] S1. Collect and preprocess multimodal data, generate an initial scene based on the LLaMA-3-8B model, and update the initial scene using deep deterministic policy gradient;
[0034] Based on the updated initial scene and POMDP belief state set to form a scene tensor, the optimized VAE model is used to generate the initial image sequence, the U-Net diffusion model is used to optimize the initial image sequence, the VAE encoder reconstruction loss and the U-Net diffusion model diffusion denoising loss are calculated, and the joint loss function of the VAE encoder and the U-Net diffusion model is calculated. The joint loss function is used to generate a high-quality image sequence through the Monte Carlo simulation method, and the high-quality image sequence is optimized through the conditional diffusion model to generate the final scene image sequence;
[0035] Specifically, multimodal data is collected and preprocessed, an initial scene is generated based on the LLaMA-3-8B model, and the initial scene is updated using deep deterministic policy gradient (DDPG). This involves collecting multimodal data including text (such as traffic rules, weather patterns, ISO26262 standards, and CARLA simulation instructions), images (satellite images, traffic scenes captured by cameras), lidar, GPS, temperature, humidity, and videos (dashcam videos (720p, 30fps, capturing urban / highway scenes)) and preprocessing them. The CLIP-ViT-L-336px multimodal pre-trained model is used to generate a multimodal feature vector (dimension 512) from the preprocessed data. The numpy.stack function is used to combine the feature vector list into a feature matrix A and store it in the Faiss vector database. The multimodal feature vector is trained based on the LLaMA-3-8B pre-trained language model, and the initial scene (such as weather, traffic, events, and spatial layout) in JSON format is output.
[0036] The initial scene is parsed using a JSON parsing library. The CARLA simulator is used to dynamically generate a CARLA simulation environment vector (e.g., vehicle position, speed) based on the parsed JSON scene. The CARLA simulation environment vector is integrated with the multimodal feature vector into a feature vector set E using the Transformer method. The feature vector set E is used as the state input (the CARLA environment vector (low-dimensional structured data) is dimensionally adjusted through a linear layer to align it with the multimodal feature vector (512-dimensional CLIP embedding)). A deep deterministic policy gradient architecture is configured using an Actor and Critic neural network (3-layer MLP). The state input feature vector set E (the Actor network processes the state input to produce continuous actions (for example, acceleration, steering angle) as a vector of real-valued numbers, while the Critic network processes the state and action input to estimate values (such as defining a reward function as an estimate) to evaluate the quality of the action) outputs a continuous action data vector. By utilizing replay buffer transformation and soft targets, the continuous action data vector output by the Actor-Critic network is optimized, and the fields in the JSON initial scene (specific attributes in the JSON scene, for example, vehicle position, traffic flow, event results) are modified according to the optimized action data vector to update the JSON initial scene.
[0037] Through CLIP-ViT-L-336px, a 512-dimensional multimodal feature vector is generated and integrated into a feature matrix. Combined with LLaMA-3-8B, the JSON initial scene is output, achieving a unified semantic representation of text, image, sensor, and video data, overcoming the defect of insufficient semantic alignment. The CARLA environment vector and multimodal features are integrated using Transformer to form a feature vector set E. Continuous actions (such as acceleration and steering angle) are generated through the DDPG Actor-Critic architecture (3-layer MLP), and the action vector is optimized using the replay buffer and soft targets. The JSON scene fields (such as vehicle position and traffic) are dynamically updated. This method significantly enhances the dynamic adaptability of edge scenarios (such as sudden accidents and severe weather). Compared with traditional GAN or VAE models, the generated initial scene is more realistic and interactive. The optimized scene provides high-quality input for subsequent image sequence generation and task stratification, effectively solving the systematic deficiencies of multi-task collaborative optimization.
[0038] Furthermore, based on the updated initial scene and POMDP belief state, a scene tensor is formed. The optimized VAE model is used to generate an initial image sequence. The initial image sequence is optimized using the U-Net diffusion model. The VAE encoder reconstruction loss and the U-Net diffusion model diffusion denoising loss are calculated. The joint loss function of the VAE encoder and the U-Net diffusion model is calculated. A high-quality image sequence is generated using the joint loss function through the Monte Carlo simulation method. The high-quality image sequence is optimized through the conditional diffusion model to generate the final scene image sequence. This refers to using the CLIP model to generate an embedding vector from the updated initial scene data vector, initializing the POMDP belief state B(s), (referring to the probability distribution of the state in the scene template (such as vehicle position, weather state), obtained by statistical inference methods, for example, if the state space contains N variables, each variable is represented by a fixed dimension (such as 2-dimensional position, 2-dimensional speed, 1-dimensional weather), then B(s) can be represented as an N-dimensional probability distribution vector), caching the embedding vector and belief state to a distributed database (such as Redis, stored in the form of key-value pairs (template ID-embedding vector-belief state), and combining each embedding vector and belief state into a scene tensor , (e.g. [1, 512+N]), each tensor contains an embedding vector of a scene template and the corresponding POMDP belief state B(s);
[0039] Use the optimized VAE model to integrate the time convolution module and generate the convolutional time enhancement features of the feature matrix A. , VAE encoder (Encoder) enhances the time feature and the scene tensor The latent representation z is mapped into the VAE latent space (referring to a “compressed encoding” of the input data (e.g., sensor sequences, video frames, scene descriptions), which retains key information (e.g., vehicle trajectory patterns, weather characteristics), but has a much lower dimensionality than the original input, facilitating computation and modeling dynamics of the generative process. The belief state ensures that the latent representation is aligned with the probabilistic environment state);
[0040] Generate the initial image sequence based on the latent vector z through the VAE decoder ,(The decoder receives the potential vector z, and gradually upsamples it through the transposed convolution layer, mapping the low-dimensional vector z to a high-dimensional tensor to generate a multi-frame state sequence Initialization: Each frame is a high-resolution image (e.g., RGB format) representing the state of the environment (e.g., road, vehicle, weather);
[0041] The U-Net diffusion model is used to train the initial image sequence through the forward process Add Gaussian noise until it becomes pure noise, and train the reverse process to generate the optimized initial image sequence :
[0042] ,
[0043] in, is the Gaussian noise at diffusion step t;
[0044] Calculating VAE encoder reconstruction loss and U-Net diffusion model denoising loss :
[0045] ,
[0046] ,
[0047] in, For the decoder, is the mathematical expectation of samples and time steps, which refers to the calculation of and The process of averaging the training data samples and diffusion model time steps when is the L2 norm of the pixel-level error, For real state sequences, through data acquisition and preprocessing, the CARLA simulation platform generates scene samples, which are processed by OpenCV and Librosa to extract high-resolution (512x512) dynamic sequence features;
[0048] Calculate the joint loss function of the VAE encoder and the U-Net diffusion model :
[0049] ,
[0050] in, and is the weight coefficient, obtained through Bayesian optimization, balancing VAE and diffusion loss;
[0051] Using gradient descent and AdamW optimizer based on joint loss function Calculate the weight set of VAE encoder and U-Net , improve the quality of generated image sequences and ensure that the output environment scenes are realistic and dynamically consistent:
[0052] ,
[0053] in, is the learning rate, the hyperparameter setting is obtained, is the gradient of loss with respect to weight, obtained by automatic differentiation method;
[0054] Using weight sets via Monte Carlo simulation Generate high-quality image sequences :
[0055] ,
[0056] Among them, P is the state transition probability, which is obtained by the state transition estimation method, and Q is the time span (50 frames), which is obtained by the fixed step size allocation method. It is a Monte Carlo simulation method;
[0057] Generate the final scene image sequence through the conditional diffusion model , obtain high-quality environmental state images and output realistic scenes:
[0058] ,
[0059] ,
[0060] in, is the conditional diffusion guided method, and To guide the weight, we obtain it through experimental tuning method. The CLIP embedding of the scene tensor is obtained through CLIP encoding. Linear projection and normalization are implemented using PyTorch. The semantic features of the CLIP embedding and the probability state of B(s) are mapped to a unified dimension. LayerNorm normalization is combined to ensure numerical compatibility. Finally, a conditional vector is generated through weighted fusion to solve the problem of inconsistent dimensionality of heterogeneous vectors and enhance semantic alignment and generation quality.
[0061] By synthesizing scene tensors from updated initial scene data and POMDP confidence states, this method uses CLIP embedding (512 dimensions) and caches it in a Redis database, ensuring robust semantic alignment across multimodal inputs, overcoming the lack of unified semantic consistency in previous methods such as GAN or independent VAE; the optimized VAE is integrated with a temporal convolution module to map feature matrices and scene tensors into potential representations, generating an initial image sequence that captures dynamic features. The diffusion model is optimized in a joint loss function by adjusting weights with AdamW to ensure high visual fidelity and temporal consistency. Monte Carlo simulation utilizes state transition probabilities and optimized weights to generate high-quality sequence samples, while the conditional diffusion model generates the final sequence, which is tailored for extreme situations such as rare accidents or severe weather. This approach significantly improves adaptability to complex scenes, unlike traditional models that struggle in dynamic edge cases. The joint optimization of VAE and U-Net loss facilitates systematic multi-task collaboration, supporting downstream tasks such as target detection and dynamics modeling, and addressing the shortcomings of previous systems in generating semantically consistent, dynamically coherent, and task-adaptive edge case scenarios.
[0062] S2. Extract visual features from the final scene image sequence and use K-means cluster analysis to stratify edge cases into tasks based on visual features and semantic embeddings;
[0063] Use meta-learning inner loop methods to create adaptive parameters;
[0064] Specifically, we extract visual features from the final scene image sequence and use K-means clustering analysis to classify edge cases into task layers based on visual features and semantic embeddings. Extract the high-dimensional visual feature vector f from (using the pre-trained ResNet-50 model implemented in PyTorch to process the sequence For each frame, a 512-dimensional feature vector is extracted to capture the visual pattern. All features in 50 frames are aggregated to form a visual feature vector f). The pre-trained CLIP-ViT-L-336px model is used to train the scene template set. Encoding, the template set consists of a feature vector set E, and the (512-dimensional) semantic embedding vector is obtained from the encoding ,capturing the semantic features of the template;
[0065] K-means clustering analysis (K=3, scikit-learn) is used to cluster the visual feature vector f and the semantic embedding vector , decomposes edge cases (challenging scenarios such as sudden obstacles, bad weather (e.g., heavy rain), and complex interactions (e.g., crosswalks during collisions)) into three task layers:
[0066] ,
[0067] in, For task set For object detection tasks, For dynamics learning tasks, is the global context task, It is a clustering algorithm that divides the task into object detection layers by combining the local spatial information (such as object boundaries and textures) in the visual feature vector f with the semantic information related to the object category (such as "pedestrian" or "vehicle" labels) in the semantic embedding. The corresponding task is to use the YOLOv8 model to evaluate the target detection accuracy, and to combine the temporal information in the visual feature vector f (such as inter-frame displacement, optical flow features) with the semantic information related to dynamic events in the semantic embedding (such as "vehicle acceleration" or "pedestrian crossing") to form a dynamic learning task layer. The corresponding task is to use OpenCV's Farneback method to calculate the optical flow field, obtain the flow variance index, and embed the global semantic information in the semantic vector (such as "rainstorm weather" or "high-traffic intersection"), combined with the scene background information in the visual feature vector f (such as road structure, sky color) into the global context task layer. , the corresponding task is to use the CLIP-ViT-L-336px model to evaluate the semantic similarity of the perturbed image).
[0068] By extracting visual feature vectors and generating semantic embeddings from a set of scene templates through CLIP-ViT-L-336px, this method captures detailed visual patterns and contextual semantic information. K-means clustering analysis decomposes edge cases into three task layers: object detection, dynamic learning, and global context, corresponding to YOLOv8-based detection, Farneback optical flow analysis, and CLIP-based semantic similarity evaluation, respectively. This task layering overcomes the semantic misalignment of previous methods (such as GAN and VAE) by integrating visual and semantic features, enhances adaptability to challenging scenarios through targeted task optimization, and ensures multi-task collaboration of the system by adjusting detection accuracy (mAP), dynamic consistency (flow variance), and semantic robustness (CLIP similarity), significantly improving simulation fidelity and robustness, and being able to accurately handle edge cases and support downstream tasks such as object detection and trajectory prediction.
[0069] Furthermore, using the meta-learning inner loop method to create adaptive parameters refers to using the meta-model parameters Load the denoising U-Net model from Stable Diffusion (Rombach et al., 2022) in PyTorch, using the pre-trained weights as meta-model parameters , weights trained on large-scale image datasets, suitable for high-quality visual generation), initialize the Transformer-UNet (TransUNet) model by copying the meta-model parameters , for each task layer Creating task-specific parameters ,(parameter By copying metamodel parameters Get, such as association , association , association are related to each other by index numbers);
[0070] Use the nearest neighbor retrieval method to find the best match for each task layer. Retrieving a fixed sample support set , using PyTorch's AdamW optimizer to implement the meta-learning inner loop method, in which the scene image sequence is and the scene tensor Input the Transformer model and compare the output sequence o of the Transformer model with the scene template set Compare, calculate task-specific losses, quickly adapt to tasks and update model parameters In the inner loop, a fixed number of gradient descent steps (e.g., 5 iterations) are performed, and this iterative process generates the adaptive parameters , (adaptive parameters By copying the model parameters After 5 iterations, the gradient is calculated based on the loss of the predicted sequence and the true sequence, and the parameters are updated using the AdamW optimizer AdamW uses adaptive learning rate and weight decay to ensure efficient updates, ultimately generating task-specific parameters that can quickly adapt to edge scenarios (such as rare accidents or severe weather).
[0071] By initializing the Transformer-UNet model with pre-trained stable diffusion U-Net weights and copying the meta-model parameters of the task-specific parameters for each task layer, this method ensures robust initialization for different edge scenarios. The inner loop of meta-learning uses nearest neighbor retrieval to obtain a fixed support set for each task layer. PyTorch's AdamW optimizer is used to perform five gradient descent iterations to calculate the task-specific loss between the Transformer output sequence and the scene template set. This process iteratively updates the parameters to generate adaptive parameters tailored for edge cases such as rare accidents or severe weather, overcoming the adaptability limitations of traditional GAN or VAE models. The adaptive learning rate and weight decay of the AdamW optimizer ensure efficient convergence and enhance task-specific performance. This systematic approach supports multi-task optimization by generating parameters that meet specific simulation requirements.
[0072] S3. Optimize the adaptive parameters through the meta-learning outer loop method, generate a single output sequence with the optimized parameters, and perform quality indicator verification and robustness testing on the generated single output sequence;
[0073] Specifically, the adaptive parameters are optimized through the meta-learning outer loop method, and the optimized parameters are used to generate a single output sequence. This means using the AdamW optimizer to implement the meta-learning outer loop method in PyTorch, comparing the output sequence o with the structured label data (obtained through CARLA simulation and multimodal perception processing (OpenCV)), providing optimization direction for the outer loop, and guiding the meta-model to improve its generalization ability for new edge scenarios through outer loop optimization, summarizing the performance of all task layers (referring to the performance of each task layer on its task, such as mAP (detection), optical flow variance (dynamic), CLIP similarity (global semantics)), calculating the combined error index (the total error metric obtained by standardizing and weighting multiple performance indicators to evaluate the comprehensive performance of the entire model in multi-task scenarios), and updating the meta-model parameters through the outer loop method. , (summarize the combined error index (mAP), optical flow variance, and CLIP similarity weighted average of each task layer), calculate the global loss, use the AdamW optimizer to perform gradient descent based on the global loss, adjust the Transformer weights, reduce the aggregation error, and improve the generalization ability), and optimize the adaptive parameters Maintain task-specific performance and integrate insights from meta-model updates; integrate task-specific parameters through parameter fusion module Encoded as a unified control vector, (all task-specific parameters including ,and and ), and the control vector is linearly concatenated with the scene image sequence and scene tensor Combined with the transformer layer of Transformer-UNet (TransUNet) injected as conditional input, it guides cross-modal fusion and spatial-semantic modeling in the model's attention mechanism, regulates the task-specific features of the generated sequence (such as enhancing pedestrian detection accuracy or vehicle dynamic consistency in rainy days), and finally generates a single output image sequence y, which is suitable for extreme situations such as rare accidents and bad weather, and achieves highly consistent and high-fidelity adaptive generation.
[0074] By using PyTorch's AdamW optimizer to implement the meta-learning outer loop, this method compares the output sequence based on structured label data from CARLA simulation and OpenCV processing, and calculates a composite error metric to guide the global loss calculation, which will enable the Transformer weights, enhance the generalization of the meta-model to new edge scenarios, and surpass the adaptability limitations of traditional GAN or VAE models. At the same time, task-specific parameters are optimized to maintain the performance of object detection, dynamic learning and global context tasks. The parameter fusion module encodes these parameters into a unified control vector, linearly connected with the scene image and tensor, and injected into the Transformer layer of TransUNet to guide cross-modal fusion and spatial semantic modeling, which will generate a single output sequence through enhanced task-specific functions. It effectively alleviates the shortcomings of existing intelligent environment simulation technology, specifically solving problems such as inconsistent semantic alignment, poor adaptability to edge cases, and insufficient multi-task optimization.
[0075] Furthermore, the quality index verification and robustness test of the generated single output sequence refers to comparing the single output image sequence y with the structured label data through the SSIM method, obtaining the SSIM structural similarity index, and setting the threshold 、 as well as , pixel-level fidelity was evaluated using the scikit-image library in Python (set by perceptual fidelity threshold analysis and semantic perception threshold analysis, respectively), image clarity was measured using PSNR (OpenCV), and CLIP similarity (CLIP-ViT-L-336px) was used to evaluate the similarity with the scene tensor The similarity score of semantic alignment is obtained if the target similarity SSIM> , generating a PSNR score> , generating similarity scores> , then the quality of the output image sequence y in terms of visual fidelity (pixel-level consistency) and semantic fidelity (scene semantic consistency) is confirmed through verification;
[0076] Using the quality-verified output sequence y, we apply perturbations (e.g., Gaussian noise, brightness changes) using the albumentations library and evaluate the task using YOLOv8. The target detection accuracy is calculated to generate the mAP score, and the threshold is set by the detection performance threshold optimization method. , so that the mAP score> , (For mAP, the YOLOv8 model is fine-tuned on the CARLA edge case dataset, such as pedestrians and vehicles as targets, with weighted IoU to enhance the detection of complex objects, data augmentation through whitening (e.g., Gaussian noise, brightness changes) to improve robustness, and non-maximum suppression (NMS) to filter low-quality bounding boxes to ensure mAP scores> ), use OpenCV's Farneback method to calculate the task The optical flow field is obtained, the flow variance index is obtained, and the threshold is set by the motion consistency evaluation method. , so that the variance index ≤ (OpenCV's Farneback optical flow method obtains a smooth flow field by adjusting the window size and pyramid level optimization. The flow consistency loss is integrated into the TransUNet training to limit the frame-to-frame motion, while Gaussian filtering can reduce local noise and ensure that the variance index ≤ .), CLIP model evaluates the task again Semantic similarity of the perturbed image (threshold is set by Perturbation Semantic Stability Analysis) , through the semantic stability analysis under perturbation, the similarity , (CLIP-ViT-L-336px model evaluates perturbed images and combines fine-tuning and semantic stability loss to maintain alignment with the scene tensor, achieving similarity ), aggregates the metrics (mAP score, flow variance metric, semantic similarity) by weighted averaging, and generates a robustness report confirming stability. Using the output sequence y, metrics, and robustness report, the h5py library packages the output sequence y into an HDF5 simulation package to ensure deployment compatibility with CARLA. This step finalizes the intelligent environment simulation by ensuring that the adjusted sequence is both high-quality and robust.
[0077] By using SSIM, PSNR and CLIP similarity to verify a single output sequence against structured labeled data, the method ensures high visual fidelity (pixel-level consistency) and semantic alignment, overcoming the semantic misalignment problem that is common in previous GAN-based methods. The robustness test applies perturbations through whitening, evaluates mAP using a fine-tuned YOLOv8 model, evaluates flow variance through OpenCV's Farneback method, and evaluates semantic similarity through CLIP to ensure that each indicator meets the standards. These indicators are aggregated by weighted averaging to generate a comprehensive robustness report. This rigorous verification and robustness framework enhances the reliability of simulations in edge cases, surpasses the adaptability limitations of traditional methods, and supports simulation environment testing by ensuring consistent, high-quality outputs in different scenarios.
[0078] This embodiment also provides an artificial intelligence-based intelligent environment simulation system, including:
[0079] Multimodal data acquisition and initial scenario generation module, used to provide multimodal input for scenario generation, build preliminary simulation scenarios, and provide a basis for subsequent dynamic optimization;
[0080] An initial image sequence generation module for sampling the latent representation into an initial image sequence through the VAE decoder;
[0081] A high-quality image sequence generation module is used to generate high-quality image sequences through Monte Carlo simulation, and then generate the final scene image sequence through the conditional diffusion model;
[0082] The meta-learning inner and outer loop modules are used to initialize the TransUNet model, generate adaptive parameters based on task-specific losses, and generate the final image sequence by comparing the output sequence with the structured label data through the outer loop;
[0083] The quality verification and robustness testing module is used to ensure the high quality and robustness of the final image sequence and supports CARLA deployment.
[0084] This embodiment also provides a computer device, which is suitable for an intelligent environment simulation method based on artificial intelligence, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement an intelligent environment simulation method based on artificial intelligence as proposed in the above embodiment.
[0085] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.
[0086] This embodiment also provides a storage medium having a computer program stored thereon. When the program is executed by a processor, it implements an artificial intelligence-based intelligent environment simulation method and system as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0087] In summary, the present invention achieves efficient fusion of multimodal data through unified semantic embedding and probability state distribution, overcomes the defects of inconsistency between vision and semantics in the prior art, and the generated scene tensor not only retains visual details, but also captures the probability distribution of dynamic scenes through belief states, thereby improving the semantic consistency of the simulation environment in complex interactive scenarios. The inner loop quickly generates task-specific parameters, and the outer loop optimizes the generalization ability of the metamodel. This method enables the model to quickly adapt to edge scenes. Compared with the traditional GAN model, it significantly improves the generation efficiency and fidelity of rare scenes. Monte Carlo simulation is combined with the conditional diffusion model to generate high-quality image sequences. Through the conditional guidance of CLIP embedding and belief states, the sequence is ensured to adapt to the dynamic characteristics of extreme scenes. Task stratification is combined with meta-learning to achieve multi-task collaborative optimization. The control vector enhances the adaptability of the model to specific tasks, significantly improving the detection accuracy and semantic alignment of the generated sequence.
[0088] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An intelligent environment simulation method based on artificial intelligence, characterized by: include, Collect and preprocess multimodal data, generate an initial scene based on the LLaMA-3-8B model, and update the initial scene using deep deterministic policy gradient; Based on the updated initial scene and POMDP belief state set to form a scene tensor, the optimized VAE model is used to generate the initial image sequence, the U-Net diffusion model is used to optimize the initial image sequence, the VAE encoder reconstruction loss and the U-Net diffusion model diffusion denoising loss are calculated, and the joint loss function is calculated. The joint loss function is used to generate a high-quality image sequence through the Monte Carlo simulation method, and the high-quality image sequence is optimized through the conditional diffusion model to generate the final scene image sequence; Extract visual features from the final scene image sequence and use K-means clustering analysis to stratify edge cases into tasks based on visual features and semantic embeddings; Adaptive parameters are created using the meta-learning inner loop method, optimized using the meta-learning outer loop method, and a single output sequence is generated using the optimized parameters. The generated single output sequence is then validated for quality indicators and tested for robustness.
2. The method for simulating an intelligent environment based on artificial intelligence according to claim 1, wherein: The method of collecting multimodal data and preprocessing it, generating an initial scene based on the LLaMA-3-8B model, and updating the initial scene using deep deterministic policy gradient (DDPG) refers to collecting multimodal data including text, images, lidar, GPS, temperature, humidity, and video, and preprocessing it, using the CLIP-ViT-L-336px multimodal pre-trained model to generate a multimodal feature vector from the preprocessed data, using the numpy.stack function to combine the feature vector list into a feature matrix A, and storing it in the Faiss vector database, training the multimodal feature vector based on the LLaMA-3-8B pre-trained language model, outputting the initial scene in JSON format, parsing the initial scene using the JSON parsing library, and updating the initial scene based on the LLaMA-3-8B pre-trained language model. The CARLA simulator dynamically generates a CARLA simulation environment vector based on the parsed JSON scene, uses the Transformer method to integrate the CARLA simulation environment vector with the multimodal feature vector into a feature vector set E, takes the feature vector set E as the state input, and uses the Actor and Critic neural network to configure a deep deterministic policy gradient architecture. Based on the state input feature vector set E, it outputs a continuous action data vector. By utilizing replay buffer transformations and soft targets, it optimizes the continuous action data vector output by the Actor-Critic network, and modifies the fields in the JSON initial scene according to the optimized action data vector to update the JSON initial scene.
3. The method for simulating an intelligent environment based on artificial intelligence according to claim 2, wherein: The calculation of the VAE encoder reconstruction loss and the U-Net diffusion model diffusion denoising loss, and the calculation of the joint loss function, and the use of the joint loss function to generate a high-quality image sequence through the Monte Carlo simulation method refers to generating an initial image sequence based on the potential vector z through the VAE decoder (Decoder) , using the U-Net diffusion model to process the initial image sequence through the forward process Add Gaussian noise until it becomes pure noise, and train the reverse process to generate the optimized initial image sequence , based on the initial image sequence Calculating VAE encoder reconstruction loss and U-Net diffusion model denoising loss , based on the reconstruction loss and denoising loss Calculate the joint loss function of the VAE encoder and the U-Net diffusion model , using gradient descent and AdamW optimizer based on the joint loss function Calculate the weight set of VAE encoder and U-Net , using a weighted set via Monte Carlo simulation and the initial image sequence and scene tensors to generate high-quality image sequences , the final scene image sequence is generated through the conditional diffusion model , and obtain high-quality environmental status images.
4. The method for simulating an intelligent environment based on artificial intelligence according to claim 3, wherein: The method of extracting visual features from the final scene image sequence and using K-means clustering analysis to classify edge cases into task layers according to visual features and semantic embedding refers to using a pre-trained ResNet-50 model to extract visual features from the scene image sequence. Extract the high-dimensional visual feature vector f from the scene template set using the pre-trained CLIP-ViT-L-336px model Encode and obtain the semantic embedding vector from the encoding , using K-means cluster analysis based on the visual feature vector f and the semantic embedding vector , decomposing edge cases into three task layers .
5. The method for simulating an intelligent environment based on artificial intelligence according to claim 4, wherein: The use of the meta-learning inner loop method to create adaptive parameters refers to using the meta-model parameters , initialize the Transformer-UNet model by copying the meta-model parameters , for each task layer Creating task-specific parameters , using the nearest neighbor retrieval method for each task layer Retrieving a fixed sample support set , using PyTorch's AdamW optimizer to implement the meta-learning inner loop method, in which the scene image sequence is and the scene tensor Input the Transformer model and compare the output sequence o of the Transformer model with the scene template set Compare, calculate task-specific loss, perform a fixed number of gradient descent steps, and this iterative process generates adaptive parameters .
6. The method for simulating an intelligent environment based on artificial intelligence according to claim 5, wherein: The method of optimizing adaptive parameters through a meta-learning outer loop and generating a single output sequence from the optimized parameters refers to implementing the meta-learning outer loop method in PyTorch using the AdamW optimizer, comparing the output sequence o with structured label data, summarizing the performance of all task layers, calculating the combined error index, and updating the meta-model parameters through the outer loop method. , while optimizing the adaptive parameters , the task-specific parameters are integrated into the parameter fusion module Encode into a unified control vector, and linearly concatenate the control vector with the scene image sequence and scene tensor Combined with the transformer layer of Transformer-UNet (TransUNet) injected as conditional input, it guides cross-modal fusion and spatial-semantic modeling in the model's attention mechanism, regulates the task-specific features of the generated sequence, and finally generates a single output image sequence y.
7. The method for simulating an intelligent environment based on artificial intelligence according to claim 6, wherein: The quality index verification and robustness test of the generated single output sequence refers to comparing the single output image sequence y with the structured label data through the SSIM method, obtaining the SSIM structural similarity index, and setting the threshold 、 as well as , using the scikit-image library in Python to evaluate pixel-level fidelity, and using PSNR to measure image sharpness, and using CLIP similarity to the scene tensor The similarity score of semantic alignment is obtained if the target similarity SSIM> , generating a PSNR score> , generating similarity scores> , then pass the verification, use the quality verified output sequence y, use the albumentations library to impose perturbations and use YOLOv8 to evaluate the task Target detection accuracy, generate mAP score, set threshold , so that the mAP score> , use OpenCV's Farneback method to calculate the task Optical flow field, obtain flow variance index, set threshold , so that the variance index ≤ , the CLIP model evaluates the task again Semantic similarity of the perturbed image, setting the threshold , so that similarity ,The indicators are summarized by weighted averaging to generate a robustness report that confirms the stability.,Using the output sequence y, the indicators and the robustness report, the,h5py library packages the output sequence y into an HDF5 simulation package,,ensuring deployment compatibility with CARLA.
8. An artificial intelligence-based intelligent environment simulation system, based on the artificial intelligence-based intelligent environment simulation method according to any one of claims 1 to 7, characterized in that: include, Multimodal data acquisition and initial scenario generation module, used to provide multimodal input for scenario generation, build preliminary simulation scenarios, and provide a basis for subsequent dynamic optimization; An initial image sequence generation module for sampling the latent representation into an initial image sequence through the VAE decoder; A high-quality image sequence generation module is used to generate high-quality image sequences through Monte Carlo simulation, and then generate the final scene image sequence through the conditional diffusion model; The meta-learning inner and outer loop modules are used to initialize the TransUNet model, generate adaptive parameters based on task-specific losses, and generate the final image sequence by comparing the output sequence with the structured label data through the outer loop; The quality verification and robustness testing module is used to ensure the high quality and robustness of the final image sequence and supports CARLA deployment.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent environment simulation method based on artificial intelligence according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent environment simulation method based on artificial intelligence according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Vehicle automatic driving method and system based on large model
CN117755336A
Multi-modal scene fusion method based on large diffusion model
CN119338940A
Method for enhancing in situ adaptive artfitial intelligent model
KR102485359B1
Cited By
Simulation-to-reality calibration method based on multi-modal loss and dynamic weighting
CN120805744A
A simulation-to-reality calibration method based on multi-modal loss and dynamic weighting
CN120805744B
Invisible structure parameter reverse design method based on diffusion model and multi-modal generation
CN121328356A
Digital teacher personalized behavior modeling method based on multi-modal feature fusion
CN121392076A
A digital teacher personalized behavior modeling method based on multi-modal feature fusion
CN121392076B