Test scene style migration method and device, electronic equipment and storage medium
By combining a diffusion model with semantic information and style guidance, the semantic consistency and quality issues of style transfer images in autonomous driving scenarios were resolved, generating high-quality target style test scenario data and improving the perception capability and robustness of autonomous driving systems.
Patent Information
- Application Number
- CN202511396705.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies struggle to guarantee semantic consistency in style transfer for autonomous driving scenarios, leading to incorrect modification or occlusion of key traffic elements. Furthermore, the generated images are of insufficient quality, failing to meet the high-quality test data requirements of autonomous driving systems.
By employing a pre-trained diffusion model combined with semantic information guidance and style guidance from text and images, semantic segmentation maps are extracted from the original test scene data and target style reference data. These maps are then input into the diffusion model for style transfer, generating high-quality, semantically consistent target style test scene data.
This improved the data quality and style accuracy of the generated target style test scenario data, ensuring that key traffic elements can still be accurately represented after style transformation, thereby enhancing the recognition accuracy of the autonomous driving perception module and the robustness of the system.
Smart Images

Figure CN121503579A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of simulation testing of autonomous vehicles, and particularly relates to a style transfer method and device for test scenes, an electronic device, and a storage medium. BACKGROUND
[0002] The rapid development of autonomous driving technology puts forward higher requirements for the diversity and complexity of test scenes. In order to comprehensively evaluate the performance of autonomous driving systems under various extreme weather, lighting conditions and road states, it is necessary to construct test scenes that can cover various environmental characteristics. As an important means to improve scene diversity and enhance model generalization ability, style transfer of test videos or images is receiving more and more attention. For example, converting driving videos shot on sunny days into visual effects under different weather conditions such as rainy days and snowy days, or converting images collected during the day into versions under different lighting conditions such as night or dusk, which is of great significance for the training and testing of autonomous driving systems.
[0003] At present, style transfer methods based on deep learning, especially generative adversarial networks (GAN), have been widely used in the style transfer task of autonomous driving scenes. However, the existing technology has difficulty in guaranteeing the semantic consistency of the generated images when generating target style test scene data, which may result in incorrect modification or occlusion of key traffic elements after style transformation, affecting the recognition results of the autonomous driving perception module. At the same time, the generated images or videos have defects in visual quality, resulting in low performance in data quality and style accuracy of the generated target style test scene data, which cannot meet the needs of autonomous driving systems for high-quality test data. SUMMARY
[0004] Therefore, the embodiments of the present application provide a style transfer method and device for test scenes, an electronic device, and a storage medium, which can improve the data quality and style accuracy of the generated target style test scene data by combining the trained diffusion model with semantic information guidance and text and image style guidance.
[0005] The present application mainly includes the following aspects: In a first aspect, the embodiments of the present application provide a style transfer method for test scenes, which comprises: obtaining original test scene data and target style reference data; the original test scene data includes various scene elements in the autonomous driving test scene; the target style reference data includes at least one of a natural language description of the target style and a reference image of the target style; extracting a semantic segmentation map corresponding to the original test scene from the original test scene data; The original test scene data, the semantic segmentation graph corresponding to the original test scene, and the reference data of the target style are input into the trained test scene style transfer model to generate target test scene data with the target style; the test scene style transfer model is a diffusion model.
[0006] In a second aspect, the embodiments of the present application further provide a test scene style transfer device, which comprises: The data acquisition module is configured to acquire original test scene data and reference data of a target style; the original test scene data comprises various scene elements in an autonomous driving test scene; and the reference data of the target style comprises at least one of a natural language description of the target style and a reference image of the target style. The semantic segmentation module is configured to extract a semantic segmentation graph corresponding to the original test scene from the original test scene data. The style transfer module is configured to input the original test scene data, the semantic segmentation graph corresponding to the original test scene, and the reference data of the target style into a trained test scene style transfer model to generate target test scene data with the target style; the test scene style transfer model is a diffusion model.
[0007] In a third aspect, the embodiments of the present application further provide an electronic device, which comprises a processor, a memory, and a bus, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the processor and the memory communicate through the bus, and the machine readable instructions are executed by the processor to perform the steps of the test scene style transfer method as described above.
[0008] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor to perform the steps of the test scene style transfer method as described above.
[0009] The method, device, electronic device and storage medium provided by the embodiment of the present application are used for style migration of a test scene, and the method comprises the following steps: obtaining original test scene data and reference data of a target style; the original test scene data comprises each scene element in an autonomous driving test scene; the reference data of the target style comprises at least one of a natural language description of the target style and a reference image of the target style; a semantic segmentation map corresponding to the original test scene is extracted from the original test scene data; the original test scene data, the semantic segmentation map corresponding to the original test scene and the reference data of the target style are input into a trained test scene style migration model to generate target test scene data with the target style; and the test scene style migration model is a diffusion model. In this way, the data quality and style accuracy of the generated target test scene data with the target style can be improved by using the trained diffusion model in combination with semantic information guidance and text and image style guidance.
[0010] In order to make the above objectives, characteristics and advantages of the present application more apparent, clear and easy to understand, the following will describe the preferred embodiments in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as a limitation to the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.
[0012] Figure 1 A flowchart of a test scene style migration method provided by the embodiment of the present application is shown; Figure 2 One of the schematic diagrams of the inference process of the test scene style migration model in the embodiment of the present application is shown; Figure 3 The second schematic diagram of the inference process of the test scene style migration model in the embodiment of the present application is shown; Figure 4 One of the functional module diagrams of a test scene style migration device provided by the embodiment of the present application is shown; Figure 5 The second functional module diagram of a test scene style migration device provided by the embodiment of the present application is shown; Figure 6 The structural schematic diagram of an electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0013] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0014] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0015] Please refer to Figure 1 , Figure 1 The flowchart of the test scene style migration method provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the test scene style migration method provided by the embodiments of the present application includes the following steps: Figure 1 S101, obtaining original test scene data and reference data of a target style; the original test scene data includes various scene elements in an autonomous driving test scene; the reference data of the target style includes at least one of a natural language description of the target style and a reference image of the target style.
[0016] In the inference stage of the embodiments of the present application, the original test scene data to be migrated and the reference data of the target style specified by the user are input. The content of the original test scene data includes various scene elements in the autonomous driving test scene, which can specifically represent key elements such as roads, vehicles, pedestrians and traffic signs in the basic environment of daytime or sunny day. The original test scene data is used as the source image of style migration for subsequent processing and generation of the diffusion model. The reference data of the target style includes at least one of the natural language description of the target style and the reference image of the target style, which is used as the direction guide information of style migration.
[0017] S102, extracting a semantic segmentation map corresponding to the original test scene from the original test scene data.
[0018] Here, in order to keep the semantic structure of the traffic scene unchanged in the style transfer process, the embodiment of the application introduces a semantic segmentation map as an additional input condition of the diffusion model, and combines the semantic segmentation map information to generate an image in each step of the denoising process. This module ensures that key traffic elements such as lane lines, vehicles, pedestrians, and traffic lights can still be accurately presented after style change, avoiding semantic errors or perception interference caused by style transformation.
[0019] S103, inputting the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with a target style; the test scene style transfer model is a diffusion model.
[0020] Here, the semantic segmentation map corresponding to the original image is input at the same time, which ensures the structural integrity of traffic elements such as lane lines, traffic lights, vehicles, and pedestrians in the image during the style transfer process, avoids loss of semantic information or misidentification due to style change, and improves the usability and safety of the generated image. The reference data of the target style specified by the user is input at the same time, which can be a natural language description (such as "night + rain") or a reference image. Through a text encoder or an image encoder, it is converted into an embedding representation recognizable by the diffusion model, serving as direction guide information for style transfer. The diffusion model gradually denoises from the initial noise to generate an image with a target style according to the input original image, semantic map, and style control signal. In the generation process, semantic constraints and style guidance are combined to achieve high-quality, semantically consistent style transfer results.
[0021] Further, please refer to Figure 2 , Figure 2 is one of the schematic diagrams of the inference process of the test scene style transfer model in the embodiment of the application. As shown in Figure 2 , the test scene style transfer model 200 includes a text encoding module 210, an image encoding module 220, and a diffusion generation module 230; the inputting the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with a target style includes: Step a1, if the reference data of the target style includes a natural language description, input the natural language description of the target style into the text encoding module to obtain a natural language embedding vector of the target style.
[0022] To achieve style control based on natural language descriptions, this embodiment introduces a text-image joint embedding model such as CLIP as a text encoder. This model transforms the user-input style description (e.g., "night + rain") into a high-dimensional semantic vector, which is then used as a conditional input during the diffusion process. Through this mechanism, the model can flexibly respond to different style instructions without additional training, significantly enhancing the controllability and interactivity of style transfer.
[0023] Step a2: If the reference data for the target style includes a reference image, then the reference image for the target style is input into the image encoding module to obtain the image feature embedding vector for the target style.
[0024] Here, to meet the style transfer requirements driven by specific style reference images, embodiments of this application optionally integrate an image encoder module for extracting high-level feature representations of the reference image and fusing them into the denoising process of the diffusion model. This module enables the model to perform style alignment based on example images provided by the user, further improving the style matching accuracy and visual consistency of the generated results.
[0025] Step a3: The original test scene data is used as the data input, and the semantic segmentation map corresponding to the original test scene, the natural language embedding vector of the target style, and / or the image feature embedding vector of the target style are used as the condition input. These are then input into the diffusion generation module to obtain the style transfer result of the original test scene data.
[0026] Here, the diffusion generation module generates images with the target style by progressively denoising the initial noise based on the input original test scene data (images or videos), semantic segmentation maps, and style-controlled feature vectors. The generation process combines semantic constraints and style guidance to achieve high-quality, semantically consistent style transfer results.
[0027] Step a4: The style transfer result of the original test scene data is determined as target test scene data with the target style.
[0028] Here, the style transfer results of the original test scenario data are identified as target test scenario data with the target style. Specifically, two examples are given below to illustrate the test scenario style transfer process of this solution.
[0029] Example 1 illustrates a scenario simulating a driving situation on a rainy urban road to evaluate the performance of the perception system under low visibility conditions. First, existing high-resolution daytime urban road images are used as input. A semantic segmentation model extracts key structural information such as lane lines, traffic lights, and pedestrians to ensure the integrity of the semantic structure during the transfer process. Then, a style description of "rain + low visibility" is input in natural language and converted into an embedding vector by the CLIP text encoder, serving as the style guidance signal for the diffusion model. The diffusion model combines the original image, semantic segmentation map, and style embedding information to progressively denoise and generate images with rainy visual characteristics, including enhanced raindrop texture, reduced road surface reflection, atmospheric perspective blurring, and fogging effects for vehicle taillights. After verification by the quality assessment module, the generated images are imported into the CARLA simulation platform to construct a virtual test environment for testing the perception stability and path planning capabilities of the autonomous driving system under rainy conditions. Results show that this method can efficiently generate high-quality rainy style images, significantly improving the recognition accuracy of the perception module and reducing the false detection rate, while effectively reducing the cost of on-site data collection.
[0030] Example 2 illustrates how, to simulate a snowy rural road test scenario, the user provided a real snowy rural road image as a reference image for the target style, and input a road image of the same area taken on a sunny day as the original test scenario data. First, the semantic segmentation map of the sunny image is extracted to ensure that key traffic elements such as lane lines and signs maintain structural integrity during the style transfer process. Then, an integrated image encoder module extracts high-level style features from the reference image and integrates them into the denoising process of the diffusion model as a guiding signal for style transfer. The diffusion model generates images with snowy characteristics such as snow cover, falling snowflakes, and slippery road surfaces while preserving the semantic structure of the original image. The quality of the generated images is evaluated using metrics such as SSIM and PSNR. Simultaneously, the CLIP model is used to compare the style similarity between the generated image and the reference image to ensure that the output meets the user's expectations. Finally, the generated images and semantic information are imported into the AirSim simulation platform to construct a snowy rural road test environment for evaluating the autonomous driving system's path planning and obstacle recognition capabilities on snowy roads. This embodiment demonstrates the flexibility and practicality of the present invention in reference image-driven migration, meets the needs of style-specific customized migration, and expands the application scope of the method in complex terrain testing.
[0031] Further, please refer to Figure 3 , Figure 3 This is the second schematic diagram illustrating the inference process of the style transfer model for the test scene in this embodiment of the application. For example... Figure 3As shown, the test scene style transfer model 200 also includes a time consistency module 240. When the original test scene data is video data, the step of inputting the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style further includes: Step b1: Starting from the first frame of the style transfer result of the video data, the style transfer result of the current frame is used as the condition input, and the style transfer result of the next frame of the current frame is used as the data input. These are input into the time consistency module to obtain the style transfer result of the next frame after time consistency optimization.
[0032] Here, for the original test scene data input of video type, an inter-frame consistency mechanism is adopted, introducing time alignment constraints between adjacent frames to ensure that the generated target test scene data video maintains visual smoothness and motion coherence while changing style, meeting the needs of autonomous driving systems for dynamic test scenes. Specifically, starting from the first frame of the style transfer result of the video data, the style transfer result of the current frame is used as a conditional input, and the style transfer result of the next frame is used as a data input, which is input into the time consistency module to obtain the style transfer result of the next frame after time consistency optimization. In this embodiment, for the video-level style transfer task, a time step adjustment module is introduced into the diffusion model to maintain the continuity and motion consistency of the video sequence after style transfer by modeling the inter-frame temporal relationship. This module uses the generation result of the previous frame and the input information of the current frame for joint inference, effectively reducing inter-frame flicker and style jump problems, and improving the overall visual quality of video style transfer.
[0033] Step b2: Use the style transfer result of the next frame after time consistency optimization as a new condition input, repeat until all frames of the style transfer result of the video data have been processed, and obtain the style transfer result sequence after time consistency optimization.
[0034] Here, the style transfer result of the next frame after time consistency optimization is used as a new conditional input, and this process is repeated until all frames of the style transfer results of the video data have been processed, resulting in a sequence of style transfer results after time consistency optimization.
[0035] Step b3: The style transfer result sequence after time consistency optimization is determined as the target test scenario data with the target style.
[0036] Here, the style transfer result sequence optimized for time consistency is identified as the target test scenario data with the target style. Specifically, the following example illustrates the test scenario style transfer process of this solution.
[0037] Example 3 illustrates how, to verify the adaptability of the autonomous driving system in nighttime high-speed driving scenarios, this application embodiment performs style transfer processing on a highway video shot during the day. The video sequence undergoes preprocessing operations such as unified resolution adjustment, color space conversion, and illumination normalization, and each frame is labeled with a corresponding semantic segmentation map and a "night scene + vehicle headlight illumination" style tag. During the inference phase, the diffusion model receives the original image frames and their semantic segmentation maps as input and performs frame-by-frame style transfer in conjunction with user-specified style descriptions, generating a video sequence with features such as streetlight illumination and vehicle headlight reflection under nighttime lighting conditions. To ensure style consistency between frames, the model introduces a time step adjustment module to establish a temporal correlation mechanism between adjacent frames, avoiding flickering issues caused by style jumps. The generated video, after inter-frame continuity evaluation, is imported into the LGSVL simulation platform to reconstruct a nighttime high-speed driving scene, and an automatic playback task is set to simulate nighttime lane changes, overtaking, and other behaviors. Experimental results show that this method successfully achieves high-quality nighttime style video transfer, with natural and smooth inter-frame transitions, supports effective testing of the nighttime perception module, and improves the robustness of the autonomous driving system in complex nighttime environments.
[0038] Furthermore, before acquiring the original test scenario data and the reference data for the target style, the method further includes: Step c1: Train the style transfer model for the initial test scene based on the multimodal sample dataset.
[0039] In this embodiment, a pre-collected multimodal sample dataset is divided into a training set, a validation set, and a test set according to a certain ratio, typically 7:2:1, for model training, hyperparameter tuning, and final performance evaluation, respectively. Different training strategies are set according to specific task requirements, including single-style transfer training, multi-style joint training, video continuous training, and text / image mixed control training, to improve the model's adaptability in different application scenarios. Simultaneously, a temporal constraint loss function or inter-frame consistency mechanism is introduced to ensure the smoothness and visual stability of video-level style transfer, providing a scientific and reasonable data foundation and training guidance for the effective learning and inference of the subsequent diffusion model.
[0040] Step c2: If the objective function, perceptual loss function, structural similarity loss function, semantic consistency loss function, and style matching loss function of the initial test scene style transfer model all satisfy the corresponding preset conditions, the training ends, and the trained test scene style transfer model is obtained.
[0041] During the model training phase, this embodiment employs a denoising score matching-based objective function, training the diffusion model by progressively adding and removing Gaussian noise. Specifically, given an input image... In time step Noise is gradually added to obtain an intermediate state. The process is shown in the following formula: ; in, , , Indicates noise scheduling parameters, It is an identity matrix.
[0042] The training objective of the diffusion model is to predict noise; the objective function is... As shown in the following formula: ; in The output of the noise prediction network, The conditional inputs include semantic segmentation graphs, natural language embedding vectors, and image feature embedding vectors.
[0043] Furthermore, to enhance the realism, style controllability, and semantic fidelity of the test scenarios after style transfer, this application's embodiments introduce perceptual loss during the training process. Structural similarity loss Semantic consistency loss and style matching loss The multi-task optimization objectives are jointly monitored.
[0044] The perceptual loss function measures the high-level semantic difference between the generated image and the real image based on the VGG feature space, ensuring style consistency and texture detail restoration. The specific formula is shown below: ; in, and These represent the generated image and the real image in the VGG network, respectively. Feature mapping of layers. The generated image and the real image were calculated at the 1st... The squared Euclidean distance between the layer feature maps. This distance measures the difference between two feature maps, reflecting the perceptual similarity between the generated image and the real image.
[0045] The structural similarity loss function is used to improve the ability to preserve edges and local structures, ensuring that key elements such as lane lines and traffic lights are not distorted. The specific formula is shown below: .
[0046] The semantic consistency loss function is used to constrain the consistency between the generated image and the original semantic graph in the semantic segmentation space, preventing the loss or alteration of traffic elements. The specific formula is shown below: ; in, For the first The predicted probability graph of class semantics (which can be computed by a fixed, co-trained segmentation network).
[0047] The style matching loss function calculates the similarity between the generated result and the target style description based on the CLIP embedding space, ensuring style controllability. The specific formula is shown below: ; here, The generated image was calculated. and target style description The cosine similarity is calculated in the CLIP embedding space. Next, the cosine similarity is converted into a loss function. Since a cosine similarity closer to 1 indicates greater similarity between the two embeddings, it is converted into a loss function using 1-cosine, such that a smaller loss value indicates a greater similarity between the generated image and the target style description.
[0048] The overall loss function is as follows: ; in - Adjustable weights are used to balance different optimization objectives. During training, the AdamW optimizer is selected, and the learning rate is set to 1×10⁻⁶. -4 The system employs cosine annealing for dynamic adjustment. Key hyperparameters are optimized on the validation set using grid search and Bayesian optimization. Through this multi-task joint optimization, the embodiments of this application significantly improve the realism, style controllability, and semantic fidelity of the generated images.
[0049] Further, the multimodal sample dataset is constructed according to the following steps: Step d1: Collect raw image or video data under various traffic scenarios; the traffic scenarios include different weather conditions, lighting conditions and time periods.
[0050] Here, this application embodiment collects driving scene images or video data from real road environments and virtual simulation platforms to construct a dataset covering various traffic scenarios and environmental conditions, including various road types such as urban roads, highways, rural roads and special areas, and covering different time periods such as morning, noon, dusk, and night, as well as various weather conditions such as sunny, rainy, snowy, and foggy days, to ensure the diversity of the collected data in spatial and temporal dimensions, providing a comprehensive training foundation for subsequent style transfer models.
[0051] Step d2: Perform semantic segmentation on the original image or video data to generate a corresponding semantic segmentation map.
[0052] Here, the acquired images or video frames are finely annotated to generate pixel-level semantic segmentation maps. The annotation categories include key traffic elements such as roads, lane lines, vehicles, pedestrians, and traffic lights. Optionally, object detection boxes or instance segmentation masks are provided to support finer-grained semantic control.
[0053] Step d3: Add a natural language description and / or a reference image of the corresponding style to the original image or video data.
[0054] Here, corresponding style labels (such as "night scene" and "rainstorm") and natural language style descriptions (such as "rainy night highway") are added to each set of data, or reference images of the corresponding style are added to form a structured annotation system, providing multimodal input support for the diffusion model.
[0055] Step d4: Integrate the original image or video data, corresponding semantic segmentation maps, corresponding style natural language descriptions, and / or corresponding style reference images from various traffic scenarios to obtain the multimodal sample dataset.
[0056] First, the original images undergo uniform preprocessing operations, including resolution adjustment, color space conversion, noise reduction, contrast enhancement, and illumination normalization, to improve image quality and reduce interference from device differences or environmental factors. Simultaneously, data augmentation techniques, such as rotation, flipping, cropping, and color perturbation, are introduced to expand data diversity, enhance model generalization ability, and ensure the stability and efficiency of subsequent diffusion model training.
[0057] Next, after image annotation and preprocessing, the original images, corresponding semantic segmentation maps, style labels, natural language descriptions, and optional reference images are integrated into a unified format of multimodal samples and organized into a standard dataset that can be used for diffusion model training. Each sample contains complete visual information and control signals. Simultaneously, a corresponding metadata file is created to record auxiliary information such as sample source, acquisition time, device parameters, and environmental conditions to support subsequent model training, debugging, and result analysis. This constructs a high-quality, clearly structured style transfer training system with semantic guidance capabilities.
[0058] Furthermore, the method also includes: Step e1: Determine the image quality of the target test scene data based on the structural similarity index and peak signal-to-noise ratio of the target test scene data.
[0059] Here, the style transfer results output by the diffusion model are evaluated for image quality. Objective indicators such as structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR) are used to measure the sharpness and detail retention of the generated images, ensuring that the images have high-quality visual performance.
[0060] Step e2: Based on the semantic segmentation results of the target test scenario data and the semantic segmentation map of the original test scenario data, determine the semantic consistency between the target test scenario data and the original test scenario data.
[0061] Here, by comparing the semantic prediction results of the original semantic segmentation image and the style-transferred image, we calculate metrics such as Intersection over Union (IoU) and Pixel Accuracy to verify whether key traffic elements such as lane lines, traffic lights, and pedestrians still maintain structural integrity and positional accuracy after style change.
[0062] Step e3: Based on the target test scene data and the natural language description of the target style and / or the reference image of the target style, determine the style matching degree between the target test scene data and the target style.
[0063] Here, a pre-trained image-text encoder (such as the CLIP model) is used to score the semantic matching degree between the generated image and the target style description, and to evaluate whether the style transfer accurately reflects the user's specified style intent, such as whether descriptions like "night" and "rainy day" are correctly presented.
[0064] Step e4: Optimize the test scene style transfer model based on the image quality of the target test scene data, the semantic consistency with the original test scene data, and the style matching degree with the target style.
[0065] Here, the generation effect is judged based on the evaluation results to determine whether it meets the set standards. If it does not meet the standards, the problematic samples are fed back to the model training stage. The loss function adjustment strategy is combined to carry out local fine-tuning or iterative optimization to improve the subsequent generation quality and form a closed-loop optimization mechanism.
[0066] Specifically, substandard generated samples are classified, and the corresponding failure types (such as semantic drift, style deviation, and inter-frame flicker) are recorded. The weights of the corresponding loss functions are then adaptively adjusted. For example, if insufficient semantic consistency is detected, the weights are increased. If the style is not expressed accurately, then improve. If the clarity is insufficient, improve it. and Local retraining is performed on a small set of failed samples, freezing some parameters (such as the U-Net encoding layer) and updating only the relevant modules to reduce the risk of overfitting.
[0067] Furthermore, when the original test scene data is video data, the method further includes: Step f1: Based on the optical flow estimation results between adjacent frames of the target test scene data, determine the inter-frame motion coherence of the target test scene data.
[0068] Step f2: Based on the style consistency evaluation results between adjacent frames of the target test scene data, determine the inter-frame style consistency of the target test scene data.
[0069] Here, for video-level style transfer results, optical flow estimation methods or temporal similarity indices are introduced to evaluate the style consistency and motion coherence between adjacent frames, preventing problems such as inter-frame flickering and style jumps, and ensuring that the output video has good visual smoothness.
[0070] Specifically, for video sequences, frame-by-frame consistency enhancement training is employed, and an optical flow consistency loss function is introduced into the loss function. The specific formula is as follows: ; in, This represents the optical flow estimation function, used to maintain motion continuity between frames.
[0071] Step f3: Optimize the test scene style transfer model based on the inter-frame motion coherence and inter-frame style consistency of the target test scene data.
[0072] Here, the generation effect is judged based on the evaluation results to determine whether it meets the set standards. If it does not meet the standards, the problematic samples are fed back to the model training stage. The loss function adjustment strategy is combined to carry out local fine-tuning or iterative optimization to improve the subsequent generation quality and form a closed-loop optimization mechanism.
[0073] Specifically, a cyclic learning rate (Cyclic LR) is used based on the feedback results. A high learning rate is maintained during local optimization to achieve fast convergence, while the learning rate is reduced during global iteration to enhance stability.
[0074] Furthermore, after generating target test scenario data with the target style, the method further includes: After style transfer, the generated test images or video sequences are uniformly converted into standard data formats supported by the autonomous driving simulation platform, such as OpenDRIVE map files, CARLA custom blueprints, LGSVL road models, etc., to ensure that the generated content can be correctly parsed and loaded by the simulation engine.
[0075] Before importing into the simulation platform, the migrated visual images are fused with the original semantic segmentation map and object detection results to construct a complete test scenario containing structured data such as lane information, traffic signs, and the position of dynamic objects, providing accurate perception input for the autonomous driving system.
[0076] The processed stylized test scenarios are imported into mainstream autonomous driving simulation platforms, such as CARLA, LGSVL, and AirSim, to reconstruct the driving environment under target weather or lighting conditions in a virtual environment for functional verification of the perception, planning, and control modules of the autonomous driving system.
[0077] By utilizing the batch testing function of the simulation platform, test scenarios with multiple styles after transfer are loaded in parallel and automatically replayed to simulate driving tasks under different extreme weather, lighting changes or complex road conditions, thereby improving the adaptability and robustness of the autonomous driving system in diverse environments.
[0078] The simulation platform integrates visualization debugging tools and performance evaluation modules to display vehicle perception results, path planning trajectories, and system response behavior in style transfer scenarios in real time, and outputs key indicators (such as recognition accuracy, decision delay, and misjudgment rate) to quantitatively evaluate the effectiveness and practicality of style transfer scenarios.
[0079] This application provides a style transfer method for test scenarios, comprising: acquiring original test scenario data and reference data for the target style; the original test scenario data includes various scene elements in an autonomous driving test scenario; the reference data for the target style includes at least one of a natural language description of the target style and a reference image of the target style; extracting a semantic segmentation map corresponding to the original test scenario from the original test scenario data; inputting the original test scenario data, the semantic segmentation map corresponding to the original test scenario, and the reference data for the target style into a trained test scenario style transfer model to generate target test scenario data with the target style; the test scenario style transfer model is a diffusion model. Thus, by combining the trained diffusion model with semantic information guidance and style guidance from text and images, the data quality and style accuracy of the generated target style test scenario data can be improved.
[0080] Based on the same application concept, this application also provides a style transfer device for a test scenario corresponding to the style transfer method for the test scenario provided in the above embodiments. Since the principle of the device in this application is similar to the style transfer method for the test scenario in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0081] Please see Figure 4 , Figure 4This is one of the functional block diagrams of a style transfer device for a test scenario provided in an embodiment of this application. For example... Figure 4 As shown, the style transfer device 400 for the test scenario provided in this application embodiment includes: The data acquisition module 401 is used to acquire original test scene data and target style reference data; the original test scene data includes various scene elements in the autonomous driving test scene; the target style reference data includes at least one of the natural language description of the target style and the reference image of the target style.
[0082] The semantic segmentation module 402 is used to extract a semantic segmentation map corresponding to the original test scene from the original test scene data.
[0083] The style transfer module 403 is used to input the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style; the test scene style transfer model is a diffusion model.
[0084] Furthermore, when the style transfer module 403 inputs the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style, the style transfer module 403 is specifically used for: If the reference data for the target style includes a natural language description, then the natural language description of the target style is input into the text encoding module to obtain the natural language embedding vector of the target style; If the reference data for the target style includes a reference image, then the reference image for the target style is input into the image encoding module to obtain the image feature embedding vector for the target style; The original test scene data is used as the data input, and the semantic segmentation map corresponding to the original test scene, the natural language embedding vector of the target style, and / or the image feature embedding vector of the target style are used as the condition input. These are then input into the diffusion generation module to obtain the style transfer result of the original test scene data. The style transfer result of the original test scenario data is determined as target test scenario data with the target style.
[0085] Furthermore, when the original test scene data is video data, the style transfer module 403, in inputting the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style, is also used to: Starting from the first frame of the style transfer result of the video data, the style transfer result of the current frame is used as the condition input, and the style transfer result of the next frame of the current frame is used as the data input. These are input into the time consistency module to obtain the style transfer result of the next frame after time consistency optimization. The style transfer result of the next frame after time consistency optimization is used as a new conditional input, and this process is repeated until all frames of the style transfer result of the video data have been processed to obtain the style transfer result sequence after time consistency optimization. The style transfer result sequence optimized by time consistency is determined as the target test scenario data with the target style.
[0086] Further, please refer to Figure 5 , Figure 5 This is a second functional block diagram of a style transfer device for a test scenario provided in an embodiment of this application. For example... Figure 5 As shown, the style transfer device 400 for the test scene also includes: Model training module 404 is used to train the style transfer model for the initial test scene based on a multimodal sample dataset; The model determination module 405 is used to terminate training and obtain the trained test scene style transfer model when the objective function, perceptual loss function, structural similarity loss function, semantic consistency loss function and style matching loss function of the initial test scene style transfer model all meet the corresponding preset conditions.
[0087] Furthermore, the model training module 404 is also used to construct the multimodal sample dataset according to the following steps: Collect raw image or video data under various traffic scenarios; the traffic scenarios include different weather conditions, lighting conditions, and time periods. Perform semantic segmentation on the original image or video data to generate a corresponding semantic segmentation map; Add a natural language description in a corresponding style and / or a reference image in a corresponding style to the original image or video data; The multimodal sample dataset is obtained by integrating the original image or video data from various traffic scenarios, the corresponding semantic segmentation maps, the corresponding style of natural language descriptions, and / or the corresponding style of reference images.
[0088] Furthermore, such as Figure 5 As shown, the style transfer device 400 for the test scene also includes: The first verification module 406 is used to determine the image quality of the target test scene data based on the structural similarity index and peak signal-to-noise ratio of the target test scene data.
[0089] The second verification module 407 is used to determine the semantic consistency between the target test scene data and the original test scene data based on the semantic segmentation result of the target test scene data and the semantic segmentation map of the original test scene data.
[0090] The third verification module 408 is used to determine the style matching degree between the target test scene data and the target style based on the target test scene data and the natural language description of the target style and / or the reference image of the target style.
[0091] The model optimization module 409 is used to optimize the test scene style transfer model based on the image quality of the target test scene data, the semantic consistency with the original test scene data, and the style matching degree with the target style.
[0092] Furthermore, such as Figure 5 As shown, the style transfer device 400 for the test scene also includes: The fourth verification module 410 is used to determine the inter-frame motion coherence of the target test scene data based on the optical flow estimation results between adjacent frames of the target test scene data; and to determine the inter-frame style consistency of the target test scene data based on the style consistency evaluation results between adjacent frames of the target test scene data. The model optimization module 409 is also used to optimize the test scene style transfer model based on the inter-frame motion coherence and inter-frame style consistency of the target test scene data.
[0093] This application provides a style transfer device for test scenarios, comprising: a data acquisition module for acquiring original test scenario data and reference data for the target style; the original test scenario data includes various scene elements in an autonomous driving test scenario; the reference data for the target style includes at least one of a natural language description of the target style and a reference image of the target style; a semantic segmentation module for extracting a semantic segmentation map corresponding to the original test scenario from the original test scenario data; and a style transfer module for inputting the original test scenario data, the semantic segmentation map corresponding to the original test scenario, and the reference data for the target style into a trained test scenario style transfer model to generate target test scenario data with the target style; the test scenario style transfer model is a diffusion model. Thus, by combining the trained diffusion model with semantic information guidance and style guidance from text and images, the data quality and style accuracy of the generated target style test scenario data can be improved.
[0094] Based on the same application concept, please refer to Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 600 includes a processor 610, a memory 620, and a bus 630.
[0095] The memory 620 stores machine-readable instructions executable by the processor 610. When the electronic device 600 is running, the processor 610 and the memory 620 communicate through the bus 630. When the machine-readable instructions are executed by the processor 610, the steps of the style transfer method for the test scenario provided in the above embodiment are executed.
[0096] Based on the same concept, this application also provides a computer-readable storage medium storing a computer program, which, when run by a processor, executes the steps of the style transfer method for the test scenario provided in the above embodiments.
[0097] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.
[0098] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0100] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0101] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0102] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0103] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the technical scope disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A style transfer method for a test scenario, characterized in that, The method includes: Acquire raw test scene data and target style reference data; the raw test scene data includes various scene elements in the autonomous driving test scene; the target style reference data includes at least one of the natural language description of the target style and the reference image of the target style. Extract the semantic segmentation map corresponding to the original test scenario from the original test scenario data; The original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style are input into the trained test scene style transfer model to generate target test scene data with the target style; the test scene style transfer model is a diffusion model.
2. The style transfer method for test scenarios according to claim 1, characterized in that, The test scene style transfer model includes a text encoding module, an image encoding module, and a diffusion generation module; the step of inputting the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style includes: If the reference data for the target style includes a natural language description, then the natural language description of the target style is input into the text encoding module to obtain the natural language embedding vector of the target style; If the reference data for the target style includes a reference image, then the reference image for the target style is input into the image encoding module to obtain the image feature embedding vector for the target style; The original test scene data is used as the data input, and the semantic segmentation map corresponding to the original test scene, the natural language embedding vector of the target style, and / or the image feature embedding vector of the target style are used as the condition input. These are then input into the diffusion generation module to obtain the style transfer result of the original test scene data. The style transfer result of the original test scenario data is determined as target test scenario data with the target style.
3. The style transfer method for test scenarios according to claim 2, characterized in that, The test scene style transfer model further includes a time consistency module; when the original test scene data is video data, the step of inputting the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style further includes: Starting from the first frame of the style transfer result of the video data, the style transfer result of the current frame is used as the condition input, and the style transfer result of the next frame of the current frame is used as the data input. These are input into the time consistency module to obtain the style transfer result of the next frame after time consistency optimization. The style transfer result of the next frame after time consistency optimization is used as a new conditional input, and this process is repeated until all frames of the style transfer result of the video data have been processed to obtain the style transfer result sequence after time consistency optimization. The style transfer result sequence optimized by time consistency is determined as the target test scenario data with the target style.
4. The style transfer method for the test scenario according to claim 3, characterized in that, Before acquiring the original test scenario data and the reference data for the target style, the method further includes: The style transfer model for the initial test scene was trained based on a multimodal sample dataset. If the objective function, perceptual loss function, structural similarity loss function, semantic consistency loss function, and style matching loss function of the initial test scene style transfer model all satisfy the corresponding preset conditions, the training ends, and the trained test scene style transfer model is obtained.
5. The style transfer method for the test scenario according to claim 4, characterized in that, The multimodal sample dataset is constructed according to the following steps: Collect raw image or video data under various traffic scenarios; the traffic scenarios include different weather conditions, lighting conditions, and time periods. Perform semantic segmentation on the original image or video data to generate a corresponding semantic segmentation map; Add a natural language description in a corresponding style and / or a reference image in a corresponding style to the original image or video data; The multimodal sample dataset is obtained by integrating the original image or video data from various traffic scenarios, the corresponding semantic segmentation maps, the corresponding style of natural language descriptions, and / or the corresponding style of reference images.
6. The style transfer method for test scenarios according to claim 1, characterized in that, The method further includes: Based on the structural similarity index and peak signal-to-noise ratio of the target test scene data, the image quality of the target test scene data is determined. Based on the semantic segmentation results of the target test scenario data and the semantic segmentation map of the original test scenario data, the semantic consistency between the target test scenario data and the original test scenario data is determined. Based on the target test scenario data and the natural language description of the target style and / or the reference image of the target style, determine the style matching degree between the target test scenario data and the target style; The test scene style transfer model is optimized based on the image quality of the target test scene data, the semantic consistency with the original test scene data, and the style matching degree with the target style.
7. The style transfer method for test scenarios according to claim 6, characterized in that, When the original test scene data is video data, the method further includes: Based on the optical flow estimation results between adjacent frames of the target test scene data, the inter-frame motion coherence of the target test scene data is determined. Based on the style consistency evaluation results between adjacent frames of the target test scene data, the inter-frame style consistency of the target test scene data is determined. Based on the inter-frame motion coherence and inter-frame style consistency of the target test scene data, the test scene style transfer model is optimized.
8. A style transfer device for a test scenario, characterized in that, The style transfer device for the test scenario includes: The data acquisition module is used to acquire raw test scene data and target style reference data; the raw test scene data includes various scene elements in the autonomous driving test scene; the target style reference data includes at least one of the natural language description of the target style and the reference image of the target style. The semantic segmentation module is used to extract a semantic segmentation map corresponding to the original test scenario from the original test scenario data; The style transfer module is used to input the original test scene data, the semantic segmentation map corresponding to the original test scene, and the reference data of the target style into the trained test scene style transfer model to generate target test scene data with the target style; the test scene style transfer model is a diffusion model.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the style transfer method for the test scenario as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the style transfer method for the test scenario as described in any one of claims 1 to 7.