Sample data synthesis method and device, equipment and medium

By building and expanding intelligent driving scenarios and generating training samples, the problem that deep learning models are difficult to generate difficult samples in intelligent driving is solved, and the model performance is improved.

CN119942269APending Publication Date: 2025-05-06CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510028656.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the field of intelligent driving, deep learning models are difficult to generate enough perceptual data difficult samples, resulting in limited model generalization.

Method used

By constructing the original dynamic scene based on environment-aware data, expanding dynamic and static environment elements, building multiple candidate simulated dynamic scenes, identifying the target simulated dynamic scenes, and using scene description information to expand scene features, and finally building training samples.

Benefits of technology

It realizes efficient generation of difficult samples, improves the performance of the model in a wide range of scenarios, and solves the problem of limited performance of traditional data enhancement methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942269A_ABST
    Figure CN119942269A_ABST
Patent Text Reader

Abstract

The invention discloses a sample data synthesis method, device and equipment and a medium, which are applied to the field of intelligent driving, and comprise the following steps: constructing an original dynamic scene based on environmental perception data, the environmental perception data being detected by a plurality of data acquisition devices, the original dynamic scene comprising an environmental background and an environmental element set, the environment element set comprises a plurality of dynamic environment elements and static environment elements; simulating the motion tracks of the dynamic environment elements in the original dynamic scene to obtain a plurality of initial simulated dynamic scenes, and expanding the static environment elements to obtain expanded static environment elements; adding the expanded static environment elements to the plurality of initial simulation dynamic scenes to obtain a plurality of candidate simulation dynamic scenes; and identifying a target simulation dynamic scene in the plurality of candidate simulation dynamic scenes, and constructing a training sample by using the target simulation dynamic scene. According to the method, the problem that model generalization is limited due to the fact that enough sensing data difficult samples are difficult to generate in deep learning model training is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving, and in particular to a method, device, equipment and medium for synthesizing sample data. Background Art

[0002] In the field of artificial intelligence and data science, obtaining difficult samples in perception data is crucial to improving the performance of deep learning models during training. Especially in the research and development of intelligent driving functions, due to the need to adapt to different scenarios, obtaining diverse and challenging data has become the key. However, there are many challenges at present: the acquisition of scarce data is costly and time-consuming; at the same time, in order to solve extreme scenario problems, a large number of training samples based on scarce dynamic scenarios are required, but this often requires large-scale vehicle data collection and a large amount of manual data screening, mining and processing, and ultimately only a limited number of training samples can be obtained, resulting in high investment, long cycle and low returns.

[0003] Although there are some relevant technical solutions at present, they still have significant limitations in solving the problem of difficult sample acquisition. For example, existing technical solutions mostly involve the construction of intelligent cloud platforms or security detection of specific scenarios, and are not directly aimed at improving the efficiency of sample acquisition that is particularly needed but difficult to obtain when training models. They are mostly limited to specific fields, lack versatility and flexibility, and have difficulty in coping with diverse, complex and difficult sample requirements, thus limiting the performance improvement of the model in a wide range of scenarios. In addition, traditional data enhancement methods also appear to be limited in effectiveness in generating a sufficient number of training samples corresponding to scarce dynamic scenarios, and cannot meet the needs of further improving model performance. Summary of the invention

[0004] In view of this, embodiments of the present invention provide a method, apparatus, device and medium for synthesizing sample data to solve the problem that it is difficult to generate enough perceptual data samples in deep learning model training, resulting in limited model generalization.

[0005] In a first aspect, an embodiment of the present invention provides a method for synthesizing sample data, the method comprising:

[0006] Constructing an original dynamic scene based on the environmental perception data, wherein the original dynamic scene includes an environmental background and an environmental element set, and the environmental element set includes dynamic environmental elements and static environmental elements that affect vehicle driving;

[0007] Expanding the dynamic environment elements and the static environment elements in the original dynamic scene, and constructing a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements;

[0008] Identifying a target simulated dynamic scene from among the plurality of candidate simulated dynamic scenes, and acquiring scene description information of the preset scene;

[0009] The scene features in the target simulation dynamic scene are expanded using the scene description information to obtain an expanded target simulation dynamic scene, and the expanded target simulation dynamic scene is used to construct a training sample.

[0010] Furthermore, the dynamic environment elements and the static environment elements in the original dynamic scene are expanded, and a plurality of candidate simulated dynamic scenes are constructed using the expanded dynamic environment elements and the expanded static environment elements, including:

[0011] Simulating the original activity tracks of the dynamic environment elements in the original dynamic scene to obtain a plurality of simulated activity tracks;

[0012] In the original dynamic scene, the dynamic environment elements are moved according to a plurality of simulated activity trajectories to obtain a plurality of initial simulated dynamic scenes;

[0013] Expanding the static environment element to obtain an expanded static environment element;

[0014] The expanded static environment elements are added to the multiple initial simulated dynamic scenes to obtain multiple candidate simulated dynamic scenes.

[0015] Furthermore, the original activity tracks of the dynamic environment elements in the original dynamic scene are simulated to obtain a plurality of simulated activity tracks, including:

[0016] Obtaining the original activity track of each dynamic environment element in the original dynamic scene;

[0017] Determining potential factors that affect the movement of the dynamic environmental element according to the environmental background, and generating a plurality of candidate simulated activity trajectories based on the original activity trajectory and the potential factors;

[0018] Trajectories that do not conform to actual traffic conditions are removed from the candidate simulated activity trajectories to obtain multiple simulated activity trajectories.

[0019] Further, the static environment element is expanded to obtain an expanded static environment element, including:

[0020] Acquire a mapping relationship between different scene types and preset static environment elements, wherein the mapping relationship is used to represent a plurality of preset static environment elements matched by each scene type;

[0021] Based on the mapping relationship, a plurality of candidate static environment elements matching the scene type corresponding to the initial simulated dynamic scene are obtained, and the environment elements other than the static environment element among the plurality of candidate static environment elements are used as the expanded static environment elements.

[0022] Furthermore, before obtaining the mapping relationship between different scene types and preset static environment elements, the method further includes:

[0023] Acquire multiple preset environment elements and element attributes of the preset environment elements;

[0024] Acquire multiple preset dynamic scenes and semantic information corresponding to each preset dynamic scene;

[0025] The semantic information is matched with the element attributes to obtain a mapping relationship between the preset dynamic scene and the preset environmental element.

[0026] Furthermore, the identifying a target simulated dynamic scene from among the plurality of candidate simulated dynamic scenes comprises:

[0027] Acquire a pre-built dynamic scene set, wherein the dynamic scene set includes a plurality of pre-collected real dynamic scenes;

[0028] Calculating the similarity between the candidate simulated dynamic scene and each of the real dynamic scenes;

[0029] The candidate simulated dynamic scene whose similarity is lower than a preset similarity is used as the target simulated dynamic scene.

[0030] Furthermore, the step of using the scene description information to expand the scene features in the target simulated dynamic scene to obtain the expanded target simulated dynamic scene includes:

[0031] Inputting the scene description information and the target simulated dynamic scene into a pre-trained language model;

[0032] Analyzing the scene description information through the language model to obtain target scene semantics corresponding to the scene description information, and identifying original three-dimensional image features of the target simulated dynamic scene through the language model;

[0033] The original three-dimensional image features are expanded based on the mapping relationship between the preset scene semantics and the three-dimensional image features to obtain the target three-dimensional image features, and the expanded target simulation dynamic scene is constructed using the target three-dimensional image features.

[0034] Furthermore, the language model training method includes:

[0035] Obtaining scene description samples, scene feature samples, and three-dimensional image feature samples;

[0036] fusing the scene description sample, the scene feature sample and the three-dimensional image feature sample to obtain a fused sample;

[0037] Using the fused sample to input a language model to be trained, and obtaining predicted three-dimensional image features output by the language model to be trained based on the fused sample;

[0038] Calculating the training loss of the language model to be trained based on the difference data between the predicted three-dimensional image features and the three-dimensional image feature samples;

[0039] Furthermore, after constructing the training samples by using the expanded target to simulate the dynamic scene, the method further includes:

[0040] Get the camera parameters of a specific car model;

[0041] Inputting the camera parameters and the training samples into a trained language model to obtain a sensor configuration for the specific vehicle model;

[0042] A virtual vehicle model is built based on the sensor configuration, and the virtual vehicle model is used to perform sensor simulation on the training sample to obtain a virtual synthetic sample.

[0043] Furthermore, after performing sensor simulation on the training sample using the virtual vehicle model to obtain a virtual synthetic sample, the method further includes:

[0044] detecting image parameters of the virtual synthetic sample;

[0045] Determining whether the image parameters meet preset conditions;

[0046] If the preset condition is met, verify whether the label of the virtual synthetic sample meets the element quantity requirement;

[0047] A virtual synthetic sample that meets the requirement on the number of elements is used as a target synthetic sample.

[0048] Furthermore, after taking the virtual synthetic sample that meets the element quantity requirement as the target synthetic sample, the method further includes:

[0049] Obtaining a real training sample of the vehicle of the specific model under the sensor configuration;

[0050] Constructing a first training set, a second training set and a test set based on the target synthetic samples and the real training samples, wherein the first training set includes a plurality of the target synthetic samples and the real training samples, the second training set includes a plurality of real training samples, and the test set includes a plurality of real training samples different from the second training set;

[0051] Using the first training set to train a preset deep learning model to obtain a first model, and using the second training set to train the preset deep learning model to obtain a second model;

[0052] Inputting the test set into the first model and the second model respectively to obtain evaluation results;

[0053] If the evaluation result is that the first performance indicator output by the first model is higher than the second performance indicator output by the second model, the target synthetic sample is determined to be a valid sample; or, if the evaluation result is that the first performance indicator output by the first model is lower than the second performance indicator output by the second model, the target synthetic sample is determined to be an invalid sample.

[0054] In a second aspect, an embodiment of the present invention provides a device for synthesizing sample data, the device comprising:

[0055] A construction module, used to construct an original dynamic scene based on the environmental perception data, wherein the original dynamic scene includes an environmental background and an environmental element set, and the environmental element set includes dynamic environmental elements and static environmental elements that affect vehicle driving;

[0056] A simulation module, used to expand the dynamic environment elements and the static environment elements in the original dynamic scene, and construct a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements;

[0057] An identification module, used to identify a target simulated dynamic scene from among the plurality of candidate simulated dynamic scenes, and obtain scene description information of a preset scene;

[0058] The processing module is used to expand the scene features in the target simulation dynamic scene by using the scene description information to obtain the expanded target simulation dynamic scene, and construct training samples by using the expanded target simulation dynamic scene.

[0059] In a third aspect, an embodiment of the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0060] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method of the first aspect or any corresponding embodiment thereof.

[0061] The embodiments of the present application have the following beneficial effects:

[0062] The method provided in the embodiment of the present application constructs the original dynamic scene by detecting environmental perception data, providing rich basic materials for subsequent scene expansion and dynamic scene construction. By using different types of environmental elements to expand the original dynamic scene, it is possible to create more diverse scenes. It solves the problem of limited effectiveness of traditional data enhancement methods and lack of versatility and flexibility of existing technical solutions, and provides more possibilities for generating difficult samples. By identifying the target simulated dynamic scene and obtaining its scene description information, the key part of accurately locating difficult samples is achieved. It solves the problem of unclear acquisition of difficult samples, and provides accurate direction for subsequent targeted expansion of preset scenes. By using scene description information to expand the preset scene and construct training samples, efficient generation of difficult samples is achieved. It solves the problem that traditional methods are difficult to meet the demand for a sufficient number of difficult samples to further improve model performance, and provides strong support for improving model performance.

[0063] The method provided in the embodiment of the present application obtains the original activity trajectory and predicts multiple simulated activity trajectories, so that the movement of dynamic environmental elements is more diversified, the diversity of simulated dynamic scenes is enriched, and a basis is provided for generating more challenging and scarce scenes, which helps to improve the quality of training samples, thereby improving the adaptability of deep learning models to different dynamic situations. In addition, by obtaining the mapping relationship to determine the candidate static environmental elements and using them as expanded static environmental elements, the complexity and diversity of the static environment can be increased, making the constructed simulated dynamic scenes more realistic and rich, providing more different training scenarios for the deep learning model, and enhancing the model's perception and processing capabilities for various static environmental conditions.

[0064] The method provided in the embodiment of the present application obtains a mapping relationship by matching the semantic information of the preset environmental elements and their attributes with the preset dynamic scene, which provides an accurate basis for the subsequent expansion of static environmental elements, ensures that the expanded environmental elements match the simulated dynamic scene, improves the rationality and effectiveness of scene construction, and thus improves the quality of training samples. In addition, by obtaining scene description information and using language model analysis to determine the target three-dimensional image features to construct training samples, it is possible to more accurately capture the key features of the preset scene, making the training samples more targeted and representative, which helps the deep learning model to better learn and identify the preset scene and improve the performance of the model in complex scenes.

[0065] The method provided in the embodiment of the present application is trained by fusing scene description samples, scene feature samples and three-dimensional image feature samples, which enables the language model to better understand the scene information and accurately output the predicted three-dimensional image features, providing strong support for building high-quality training samples and improving the reliability and effectiveness of the entire sample synthesis method. In addition, the camera parameters of a specific vehicle model are obtained and the sensor configuration is obtained through the trained language model. Based on this, a virtual vehicle model is built to simulate the sensor of the training sample to obtain a virtual synthetic sample, which can make the training sample closer to the actual vehicle perception, provide more realistic training data for the deep learning model, and improve the accuracy and reliability of the model in practical applications.

[0066] The method provided in the embodiment of the present application can screen out high-quality virtual synthetic samples by detecting the image parameters of the virtual synthetic samples and judging whether they meet the preset conditions, and verifying whether the labels meet the element quantity requirements, thereby ensuring the validity and integrity of the training samples, providing better training data for the deep learning model, and improving the training effect of the model. In addition, by constructing different training sets and test sets based on the target synthetic samples and the real training samples, and by comparing the performance indicators of different models on the test sets to determine the validity of the target synthetic samples, the quality and value of the synthetic samples can be accurately evaluated, providing a reliable basis for the training of the deep learning model, and improving the performance and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0068] Figure 1 is a schematic flow chart of a method for synthesizing sample data according to some embodiments of the present invention;

[0069] Figure 2 is a schematic diagram of an original dynamic scene construction process according to an embodiment of the present invention;

[0070] Figure 3 is a schematic diagram of a three-dimensional image generation process according to an embodiment of the present invention;

[0071] Figure 4 is a flow chart of another method for synthesizing sample data according to some embodiments of the present invention;

[0072] Figure 5is a system block diagram of a sample data synthesis system according to an embodiment of the present invention;

[0073] Figure 6 is a structural block diagram of a device for synthesizing sample data according to an embodiment of the present invention;

[0074] Figure 7 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0076] According to an embodiment of the present invention, a method, apparatus, device and medium for synthesizing sample data are provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0077] In this embodiment, a method for synthesizing sample data is provided. Figure 1 is a flow chart of a method for synthesizing sample data according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0078] Step S11, constructing an original dynamic scene based on the environmental perception data, wherein the original dynamic scene includes an environmental background and an environmental element set, and the environmental element set includes dynamic environmental elements and static environmental elements that affect vehicle driving.

[0079] In the embodiment of the present application, the environmental perception data comes from devices such as mobile phones, cameras or sensors installed on collection vehicles, and contains clear images, element size information and data required for 3D reconstruction. When constructing the original dynamic scene, the environmental perception data is first used to create the environmental background through 3D reconstruction technology, and then a set of environmental elements is added to it, including dynamic elements (such as pedestrians and vehicles) and static elements (such as roads, buildings, obstacles) that affect vehicle driving. For example, Figure 2 As shown, the dynamic elements of the vehicle are added to the original data in the left figure, and the original data and the newly added dynamic elements are combined to construct a complete original dynamic scene, and the result shown in the right figure is obtained.

[0080] Step S12, expanding the dynamic environment elements and the static environment elements in the original dynamic scene, and constructing a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements.

[0081] In an embodiment of the present application, the dynamic environment elements and the static environment elements in the original dynamic scene are expanded, and multiple candidate simulated dynamic scenes are constructed using the expanded dynamic environment elements and the expanded static environment elements, including the following steps A1-A4:

[0082] Step A1, simulating the original activity tracks of the dynamic environment elements in the original dynamic scene to obtain a plurality of simulated activity tracks.

[0083] In the embodiment of the present application, simulating the activity trajectories of the dynamic environment elements in the original dynamic scene to obtain a plurality of initial simulated dynamic scenes includes the following steps A11-A13:

[0084] Step A11, obtaining the original activity track of each dynamic environment element in the original dynamic scene.

[0085] Specifically, motion data is extracted from the original dynamic scene, including the position, speed, acceleration, direction, etc. of the elements. These data describe the motion state of the elements in the scene. Subsequently, the motion data is preprocessed, such as cleaning, denoising, smoothing, etc., to obtain more accurate motion data. Then, these processed data are converted into the activity trajectory of the elements, which is usually expressed as the correspondence between timestamps and motion states. Finally, the original activity trajectory of each dynamic environment element is extracted from these data. These trajectories reflect the actual motion path and behavior pattern of the elements in the scene, providing a basis for subsequent simulation and prediction.

[0086] Step A12, determining potential factors that affect the movement of dynamic environmental elements according to the environmental background, and generating a plurality of candidate simulated activity trajectories based on the original activity trajectory and the potential factors.

[0087] Specifically, first, the environmental background is analyzed to determine the potential factors that affect the movement of dynamic environmental elements. The environmental background may include road conditions (such as the slope, flatness, and width of the road), traffic rules (such as speed limits, traffic signals, and lane divisions), surrounding obstacles (such as buildings, trees, and other stationary objects), and other dynamic elements (such as other vehicles, pedestrians, etc.). Through detailed observation and analysis of the environmental background, the specific ways in which these potential factors affect the movement of dynamic environmental elements can be identified. Next, the movement habits and laws of dynamic environmental elements are understood based on the original activity trajectory. The original activity trajectory may include average speed, acceleration, steering frequency, dwell time, etc. These features reflect the movement patterns of dynamic environmental elements in a specific environment. Then, multiple candidate simulated activity trajectories are generated by combining the potential factors and the original activity trajectory. A random generation method can be used to adjust the parameters of the trajectory, such as speed, direction, and acceleration, based on the original activity trajectory according to the potential factors. For example, if there is a curve in the environment, the steering angle and speed of the vehicle can be randomly adjusted according to the curvature of the curve and the driving speed of the vehicle to generate multiple different candidate trajectories. Or if there is a traffic signal, the vehicle's driving speed and stop time can be adjusted according to the signal status and time to generate different trajectories.

[0088] Step A13, removing the trajectories that do not conform to the actual traffic conditions from the candidate simulated activity trajectories, and obtaining a plurality of simulated activity trajectories.

[0089] Specifically, first, determine the criteria that meet the actual traffic conditions. This may include whether the trajectory collides with road boundaries and obstacles, whether the speed is within a reasonable range, whether traffic rules are followed, etc. Then, check each candidate simulated activity trajectory. For trajectories that collide with road boundaries or obstacles, remove them directly. For trajectories with speeds exceeding a reasonable range, determine whether to remove them based on the actual situation. For example, if it is a vehicle trajectory, a speed that is too high may not meet the actual traffic conditions; if it is a pedestrian trajectory, a speed that is too fast may also be unreasonable. At the same time, check whether the trajectory complies with traffic rules. For example, if the trajectory violates traffic signals, drives in the opposite direction, or drives in a prohibited area, it should also be removed. By performing such checks and screening on all candidate simulated activity trajectories, multiple simulated activity trajectories that meet the actual traffic conditions are finally obtained. These trajectories can be used to further generate initial simulated dynamic scenes to enrich the diversity and authenticity of dynamic scenes.

[0090] Step A2: In the original dynamic scene, the dynamic environment elements are moved according to a plurality of simulated activity trajectories to obtain a plurality of initial simulated dynamic scenes.

[0091] Specifically, first, multiple simulated activity trajectories are preliminarily screened according to rules such as trajectory rationality and diversity or random selection. Then, the current position and state information of the dynamic environment elements in the original dynamic scene are read. For each selected simulated activity trajectory, the position and state of the element in the original scene are gradually updated according to the timestamp and motion state in the trajectory, and the movement can be achieved with the help of animation or simulation software. As the element moves along the trajectory, the original dynamic scene is transformed into a simulated dynamic scene. Each selected trajectory generates a corresponding initial simulated dynamic scene, and records its generation process and results, including information such as element position changes and interactions with other elements.

[0092] Step A3, expanding the static environment element to obtain an expanded static environment element.

[0093] Specifically, the static environment element is expanded to obtain the expanded static environment element, including the following steps A31-A32:

[0094] Step A31, obtaining a mapping relationship between different scene types and preset static environment elements, wherein the mapping relationship is used to represent a plurality of preset static environment elements matched by each scene type.

[0095] Specifically, the preset dynamic scenes are various traffic dynamic scenes that are preset in advance, covering different traffic conditions, time points and locations, etc. The preset environmental elements include weather conditions that may appear in traffic scenes (such as sunny, cloudy, rainy, snowy, etc.), lighting conditions (changes between day and night, angle of sunlight, etc.), material properties (reflectivity and roughness of roads and buildings, etc.), and traffic participants (pedestrians, vehicles, etc.). For example, the preset specific dynamic scene is a city street during rush hour. According to the mapping relationship, it may correspond to preset environmental elements such as sunny weather, daytime lighting, road materials with a certain degree of roughness, and a large number of vehicles and pedestrians.

[0096] Step A32: based on the mapping relationship, a plurality of candidate static environment elements matching the scene type corresponding to the initial simulated dynamic scene are obtained, and the environment elements other than the static environment elements among the plurality of candidate static environment elements are used as expanded static environment elements.

[0097] Specifically, the initial simulated dynamic scene is obtained by the dynamic environment elements moving in the original dynamic scene according to multiple simulated activity trajectories, including position changes and interaction information with other elements. Based on the mapping relationship, the scene characteristics are first analyzed, such as traffic conditions, time settings, dynamic element behaviors, etc., and then matching static environment elements are found in the mapping relationship based on these characteristics. The mapping relationship covers the correspondence between weather, lighting, material properties, etc. and the traffic environment, and multiple candidate static environment elements can be found, including roads, buildings, etc. For example, when the initial simulated dynamic scene is set to a rainy day, candidate elements such as road materials with reflective characteristics of slippery roads and buildings with wet appearances can be found based on the mapping relationship.

[0098] The extended static environment elements are the elements remaining after removing the parts that overlap with the static environment elements in the original dynamic scene from the candidate static environment elements. These elements can add details and diversity to the original dynamic scene, making the scene more realistic and rich, and providing more valuable data for subsequent analysis and application. For example, in a specific traffic dynamic scene, the original static environment elements include roads and buildings, and the candidate static environment elements include special light and shadow effects under specific weather conditions in addition to roads and buildings. After removing the overlapping parts, the special light and shadow effects become extended static environment elements, which can enrich the scene and make it closer to reality.

[0099] Step A4, adding the expanded static environment elements to the multiple initial simulated dynamic scenes to obtain multiple candidate simulated dynamic scenes.

[0100] Specifically, the extended static environment elements are added to multiple initial simulated dynamic scenes to make the obtained candidate simulated dynamic scenes richer and more realistic. For example, the initial simulated dynamic scene only contains basic static elements, and after adding extended static environment elements, such as special light and shadow effects under specific weather conditions, the scene can be closer to the actual situation. These candidate simulated dynamic scenes can provide more valuable scene data for the testing and optimization of traffic systems and the research and development of intelligent driving technologies, and help solve various complex situations and problems that may be encountered in actual traffic.

[0101] In the embodiment of the present application, before obtaining the mapping relationship between different scene types and preset static environment elements, the following steps B1-B3 are also included:

[0102] Step B1, obtaining a plurality of preset environment elements and element attributes of the preset environment elements.

[0103] Specifically, the preset environmental elements come from the basic elements such as road networks, traffic signs, and traffic lights provided by OpenScenario on the one hand, and weather conditions (including specific parameters), lighting conditions (such as day and night changes, and other properties), and material properties (such as reflectivity, etc.) that affect the appearance of traffic scenes and sensors on the other hand. Each preset environmental element has specific properties, which are very important in simulating real traffic scenes and will affect the traffic environment, vehicle driving, and sensor detection. Acquiring them provides the basis for subsequent scene generation and expansion. It can be used as a reference when constructing original dynamic scenes, simulating dynamic environmental element trajectories, and expanding static environmental elements, making the scenes more realistic and diverse. For example, when simulating driving scenes in rainy days at night, the lighting conditions can be adjusted according to the weather conditions.

[0104] Step B2, obtaining multiple preset dynamic scenes and semantic information corresponding to each preset dynamic scene.

[0105] Specifically, each preset dynamic scene has corresponding semantic information, which describes the specific meaning and characteristics of the scene. For example, the preset dynamic scene is "urban main road during rush hour", and its semantic information may include heavy traffic flow, relatively slow vehicle speed, many pedestrians, frequent changes in traffic lights, etc. For another example, the preset dynamic scene of "country road at night" may have semantic information such as light traffic flow, poor lighting conditions, narrow roads, and possible presence of wild animals.

[0106] Step B3, matching the semantic information with the element attributes to obtain a mapping relationship between the preset dynamic scene and the preset environmental element.

[0107] Specifically, when matching, the semantic information of the preset dynamic scene is first analyzed to determine the required environmental features. For example, the "rainy highway at night" scene requires low visibility weather, dim lighting, and slippery road material properties. Then these requirements are compared and matched with the properties of the preset environmental elements. If the properties of a preset environmental element meet the requirements of the dynamic scene, a mapping relationship is established. The mapping relationship obtained in this way provides clear guidance for generating difficult sample scenes. When generating specific difficult sample scenes, the required preset environmental elements can be quickly determined based on this relationship, and the corresponding information can be automatically filled in using the synthetic data platform and algorithm to intelligently generate a large number of unique static and dynamic difficult sample scenes.

[0108] Step S13, identifying a target simulated dynamic scene from a plurality of candidate simulated dynamic scenes, and acquiring scene description information of the preset scene.

[0109] In an embodiment of the present application, identifying a target simulated dynamic scene from a plurality of candidate simulated dynamic scenes includes: obtaining a pre-constructed dynamic scene set, wherein the dynamic scene set includes a plurality of pre-collected real dynamic scenes; calculating the similarity between the candidate simulated dynamic scene and each real dynamic scene; and taking the candidate simulated dynamic scene whose similarity is lower than a preset similarity as the target simulated dynamic scene.

[0110] Specifically, the pre-built dynamic scene set can be constructed in a variety of ways. For example, different scene fragments can be extracted from actual traffic data, and the scene fragments cover different road conditions, traffic conditions, and environmental factors. The constructed multiple preset dynamic scenes are stored in a suitable data structure, such as a database, a file system, or a dedicated data container. A unique identifier can be assigned to each preset dynamic scene for subsequent query and comparison. At the same time, an effective indexing mechanism is established to quickly retrieve specific preset dynamic scenes.

[0111] For each candidate simulated dynamic scene and each real dynamic scene, extract key information that can characterize its features. These features may include the type, quantity, position, motion trajectory of dynamic environmental elements, features of environmental background (such as road shape, traffic signs, etc.), etc. For example, these features can be extracted from the scene using image processing technology, target detection algorithms, trajectory analysis methods, etc.

[0112] Select a similarity measurement method to calculate the similarity between the candidate simulated dynamic scene and each real dynamic scene. Common similarity measurement methods include Euclidean distance, cosine similarity, structural similarity index (SSIM), etc. Select the most suitable measurement method based on the type of extracted features and the characteristics of the scene. For example, if the feature is in vector form, cosine similarity can be used; if it is an image feature, SSIM can be considered. For each real dynamic scene, compare the features of the candidate simulated dynamic scene with the features of the real dynamic scene, and calculate a similarity value based on the selected similarity measurement method. This process can be implemented by programming, traversing each real dynamic scene in the dynamic scene set and calculating the similarity one by one.

[0113] According to the actual needs and the characteristics of the scene, a preset similarity threshold is set. This threshold can be determined through experience or adjusted through experiments and analysis. If the threshold is set too high, it may result in very few scenes being identified as target simulated dynamic scenes; if the threshold is set too low, some less unique scenes may also be identified as preset scenes. The calculated similarity between the candidate simulated dynamic scene and each real dynamic scene is compared with the preset similarity threshold. If the similarity between a candidate simulated dynamic scene and all real dynamic scenes is lower than the preset similarity threshold, then the candidate simulated dynamic scene is identified as the target simulated dynamic scene. The target simulated dynamic scene can be understood as a scarce dynamic scene in the real scene.

[0114] In the embodiments of the present application, the preset scene refers to a scarce scene, which refers to dynamic scenes that appear less frequently in actual situations but are of great significance. These scenes are rare for various reasons, such as a specific combination of environmental conditions, a specific sequence of events, or a specific user behavior pattern. The scene description information includes: ① Environmental characteristics: Physical environment: describes the geographical location, topography, climate conditions, etc. where the scene occurs. For example, in a traffic scene, it may include the type of road (highway, city street, etc.), the distribution of surrounding buildings, etc. Time characteristics: clarify the time range of the scene, including specific time periods (such as morning, noon, evening), seasons, dates, etc. For example, a specific scarce scene may only appear in a specific time period in a specific season. ② Element characteristics: If the scene involves a specific subject (such as a vehicle, pedestrian, etc.), describe the behavioral characteristics of the subject. For example, in a traffic scarce scene, it may include abnormal driving behavior of the vehicle (such as high-speed reverse driving, sudden lane change, etc.). ③ Influencing factors: Trigger conditions: describe the triggering factors or conditions that cause the scarce scene to appear. For example, specific weather conditions, special traffic flow combinations, etc. may trigger a traffic scarce scene. Consequences and impacts: Analyze the possible consequences and impacts of the scenario, such as the impact on traffic flow, damage to the surrounding environment, and threats to personnel safety.

[0115] Step S14, using the scene description information to expand the scene features in the target simulated dynamic scene to obtain an expanded target simulated dynamic scene, and using the expanded target simulated dynamic scene to construct a training sample.

[0116] In an embodiment of the present application, scene features in a target simulated dynamic scene are expanded using scene description information to obtain an expanded target simulated dynamic scene, including the following steps: inputting the scene description information and the target simulated dynamic scene into a pre-trained language model; analyzing the scene description information through the language model to obtain target scene semantics corresponding to the scene description information, and identifying original three-dimensional image features of the target simulated dynamic scene through the language model; expanding the original three-dimensional image features based on a mapping relationship between preset scene semantics and three-dimensional image features to obtain target three-dimensional image features, and using the target three-dimensional image features to construct an expanded target simulated dynamic scene.

[0117] Specifically, the scene description information may include a detailed text description of various elements and situations in the target simulated dynamic scene, such as describing the road conditions, traffic signs, signal light status, specific behaviors of vehicles and pedestrians, weather and lighting conditions, etc. These description information can provide a basis for subsequent steps. By inputting the scene description information into a pre-trained language model, it is analyzed and processed to determine the target scene semantics and the corresponding target 3D image features, and finally constructing training samples based on these features for training models in the traffic system to improve its performance and accuracy in handling complex traffic scenes.

[0118] The scene description information is input into the pre-trained language model. The language model will analyze the description information and determine the semantics of the target scene corresponding to the scene description information by understanding and processing the text content. For example, if the description information mentions "a city street on a rainy day, vehicles are driving slowly, pedestrians are walking quickly with umbrellas, and there is water on the road reflecting the lights", the language model will analyze the semantics of this scene to include "slow traffic on rainy days", "pedestrians are in a hurry", "reflective road surface", etc.

[0119] At the same time, based on the mapping relationship between the preset scene semantics and the 3D image features, the language model can determine the target 3D image features corresponding to the target scene semantics. This mapping relationship is established during the training of the language model, which enables the language model to infer the corresponding 3D image features based on the scene semantics, such as specific colors, shapes, textures, etc. The target 3D image features determined in this way can provide key information for the subsequent construction of training samples, which can be used to train models in the traffic system, so that they can better handle complex traffic scenes and improve performance and accuracy.

[0120] Finally, the target three-dimensional image features are integrated with the actual traffic data to form a target simulation dynamic scene with rich information, and the target simulation dynamic scene is used to construct training samples. For example, the target three-dimensional image features can be combined with dynamic environmental element information such as vehicle speed and pedestrian position in the scene, as well as static environmental element information such as road material, to construct a training sample containing multi-dimensional data. The training samples constructed in this way can be used to train models in the traffic system, such as intelligent driving models, traffic flow prediction models, etc. By using these training samples for training, the model can better learn the various characteristics and laws in complex traffic scenes, improve the performance and accuracy when dealing with complex traffic scenes, and enable it to more accurately respond to various rare and challenging traffic situations.

[0121] As an example, Figure 3 As shown, first, the scene description information is input into the text encoder according to the guidance prompt to obtain the scene description feature P, the preset environmental element is input into the scene encoder to obtain the scene feature Q, and the candidate simulated dynamic scene is input into the visual encoder to obtain the three-dimensional image feature O, wherein the guidance prompt includes the preset environmental element, and the guidance prompt is associated with the preset environmental element. Then, the scene description feature P, the scene feature Q and the three-dimensional image feature O are fused to obtain a fused sample I. In the training process of the language model (prior model), the fused sample I can be used as the language model (prior model) to be trained, and it is trained to obtain a trained language model. In the reasoning process of the language model (prior model) to be trained, the fused sample I can be used as the input of the trained language model to obtain the amplified fused feature, and the amplified fused feature is input into the decoder to obtain multiple three-dimensional images, which are the corresponding training samples.

[0122] In the embodiment of the present application, the language model training method includes the following steps C1-C5:

[0123] Step C1, obtaining scene description samples, scene feature samples and three-dimensional image feature samples.

[0124] It should be noted that the scene description sample is a collection of text descriptions of traffic scenes, covering road conditions, traffic signs, signal light settings, vehicle and pedestrian behavior, and weather and lighting. For example, the rainy highway scene has descriptions such as "rainy highway, slippery road surface, fast speed, poor visibility, and some cars turning on fog lights." The scene feature sample is a collection of features of different traffic scenes, including scene layout, traffic flow, and the number of participant types. For example, the scene characteristics of the city center are dense roads, heavy traffic, and many pedestrians. The three-dimensional image feature sample is a collection of three-dimensional features extracted from actual traffic scene images, including color, shape, texture, and object position and size. For example, the image features of traffic intersections include intersection shape, signal light color and position, and vehicle size.

[0125] Specifically, these samples are obtained for training the language model, so that it can learn the characteristics and semantics of different traffic scenes through fusion analysis, determine the target scene semantics and three-dimensional image features based on the scene description information, and provide a basis for constructing high-quality training samples, thereby improving the performance and accuracy of the traffic system model in handling complex scenes.

[0126] Step C2, fusing the scene description samples, the scene feature samples and the three-dimensional image feature samples to obtain a fused sample.

[0127] Specifically, the fusion process can be carried out in a variety of ways, such as vectorizing the text description and then splicing it with the scene feature vector and the three-dimensional image feature vector, or combining different types of features with a specific fusion algorithm. The purpose of fusion is to allow the language model to learn different types of information at the same time and better understand the overall picture of the traffic scene. Through fusion, a fusion sample containing rich traffic scene information is obtained, which provides a more comprehensive data basis for subsequent input into the language model to be trained, so that the language model outputs more accurate predicted three-dimensional image features. Fusion can be performed based on a certain ratio, for example, combining the text description vector, the scene feature vector and the three-dimensional image feature vector in a certain ratio.

[0128] Step C3, using the fused sample to input the language model to be trained, and obtaining the predicted three-dimensional image features output by the language model to be trained based on the fused sample.

[0129] Specifically, after receiving the fused samples, the language model processes them according to the internal algorithm and parameters, and outputs the predicted 3D image features. This feature is inferred and generated based on the fused sample information, simulating the 3D image features of the actual traffic scene, including color, shape, texture, and object position and size. Obtaining this predicted feature can provide a basis for subsequent calculation of training loss and adjustment of model parameters, so as to optimize the language model, make it predict 3D image features more accurately, and better serve the training of traffic system models and the task of processing complex traffic scenes.

[0130] Step C4, calculating the training loss of the language model to be trained based on the difference data between the predicted three-dimensional image features and the three-dimensional image feature samples.

[0131] Specifically, first, the predicted 3D image features of the language model to be trained based on the fusion sample output are obtained, and the 3D image feature samples extracted from the actual traffic scene images are also obtained. Secondly, the difference data is calculated by comparing the two, which can be measured in terms of color, shape, texture, and object position and size. Finally, the training loss of the language model is calculated based on the difference data. The training loss can measure the gap between the current prediction results of the language model and the actual results, so as to understand its performance and determine the direction of improvement and optimization.

[0132] Step C5, using the training loss to adjust the model parameters of the language model to be trained, until the difference data between the predicted three-dimensional image features output by the language model after the adjustment of the parameters and the three-dimensional image feature samples is less than a preset value, thereby obtaining a trained language model.

[0133] Specifically, the training loss is used to adjust the language model parameters. An optimization algorithm such as the gradient descent method can be used to gradually adjust the training loss according to its size and direction, so that the model output is closer to the real result. The process of fusion sample input language model, obtaining predicted three-dimensional image features, calculating difference data and training loss, and adjusting parameters is repeated continuously until the predicted three-dimensional image features output by the model after adjusting the parameters and the sample difference data are less than the preset value. At this time, the trained language model can accurately predict the three-dimensional image features based on the input information, is closely related to the actual scene features and contains rich semantic information, which can provide support for the construction of high-quality training samples and improve the performance and accuracy of the traffic system model in handling complex scenes.

[0134] In the embodiment of the present application, after constructing the training sample by using the target simulation dynamic scene, the following steps D1-D3 are also included:

[0135] Step D1, obtaining camera parameters of a specific vehicle model.

[0136] Specifically, the camera parameters of a specific vehicle model are the basis for subsequent sensor simulation and vehicle model building, including internal and external parameters, image resolution, calibration angle, vertical and horizontal FOV angles, etc. Obtaining these parameters can accurately understand the characteristics and performance of the vehicle camera and provide information input for adapting the sensor calibration status. Only by obtaining accurate parameters can effective sensor configuration optimization and vehicle model building be carried out, high-quality sensor simulation of data-difficult samples can be achieved, and virtual synthetic samples that meet the requirements can be generated.

[0137] Step D2, inputting the camera parameters and training samples into the trained language model to obtain the sensor configuration for the specific vehicle model.

[0138] It should be noted that before inputting the camera parameters and training samples into the trained language model, the trained language model needs to be retrained. The specific process includes: first, preprocessing the camera parameters corresponding to multiple pre-stored vehicle models and converting them into a format suitable for language model input. For example, the camera resolution, field of view and other parameters can be converted into numerical vector representations. At the same time, the scene description information, scene feature samples and three-dimensional image features in the training samples are sorted and standardized to ensure the consistency and validity of the input data. Secondly, the preprocessed camera parameter vector is fused with the training samples. The two can be combined by splicing, weighted summation and other methods, so that the language model can simultaneously receive the hardware information and scene information of a specific vehicle model. Finally, the fused input data is input into the trained language model, and the language model will recalculate and analyze according to the new input. In this process, optimization algorithms such as small batch stochastic gradient descent can be used to update the parameters of the language model, so that the output of the model gradually approaches the optimal sensor configuration for a specific vehicle model. By repeating the above process and through multiple iterative training, the language model will continuously optimize its ability to understand specific vehicle models and scenarios, and ultimately output a more accurate and optimized sensor configuration for specific vehicle models.

[0139] Specifically, the camera parameters of a specific vehicle model and the training samples are input into the retrained language model. At this point, the camera parameters provide the language model with the hardware characteristics of the specific vehicle model, while the training samples contain the semantic and image feature information of various scenes. By combining these two inputs, the language model can better understand the needs of a specific vehicle model in different scenarios, and thus accurately output the sensor configuration for the specific vehicle model.

[0140] Step D3, building a virtual vehicle model based on the sensor configuration, and using the virtual vehicle model to perform sensor simulation on the training samples to obtain virtual synthetic samples.

[0141] Specifically, after obtaining the optimized sensor configuration for a specific vehicle model, a virtual vehicle model is built based on this configuration, which can simulate the real vehicle structure, sensor layout and performance. The virtual vehicle model is used to simulate the sensors of the training samples, simulating the sensor's perception and collection of scene data, just like the real sensor working in the actual environment. After processing, the reconstructed data of any mode and parameter specifications are rendered, that is, virtual synthetic samples. These samples can be used for traffic system testing, optimization and model training, providing valuable data support for solving complex problems in actual traffic.

[0142] In the embodiment of the present application, after the virtual vehicle model is used to perform sensor simulation on the training sample to obtain the virtual synthetic sample, the following steps E1-E4 are also included:

[0143] Step E1, detecting image parameters of the virtual synthetic sample.

[0144] Specifically, the virtual synthetic samples obtained by the virtual vehicle model sensor simulation are tested for image parameters, including brightness, contrast, saturation, etc. Brightness reflects the brightness of the image, contrast reflects the degree of difference between different regions, and saturation indicates the vividness of the color. Testing these parameters can help us understand the sample image quality and visual effects. These parameters can be measured and evaluated through professional tools or algorithms to provide a basis for judging whether the image meets the preset conditions.

[0145] Step E2, determining whether the image parameters meet preset conditions.

[0146] Specifically, it is determined whether the image parameters of the virtual synthetic sample meet the preset conditions according to specific standards and requirements. The preset conditions are determined according to actual application needs, industry standards or specific project requirements, such as setting a suitable range for brightness, and setting a reasonable range for contrast and saturation. If the parameters of the virtual synthetic sample are within the preset range, the preset conditions are met and the next step of verification can be continued; if not, the sample needs to be adjusted or regenerated.

[0147] Step E3: If the preset conditions are met, verify whether the label of the virtual synthetic sample meets the element quantity requirement.

[0148] Specifically, if the parameters of the virtual synthetic sample image meet the preset conditions, verify whether its label meets the element quantity requirement. Virtual synthetic samples are labeled with various elements, such as traffic signs, vehicles, pedestrians, etc. The purpose of verification is to ensure that the scarce sample requirements are met. The element quantity requirement is the minimum number or specific combination of different elements that the label should contain. For example, traffic scene samples require labels to contain a certain number of different types of traffic signs and vehicles. By checking the number of elements, it can be determined whether the sample has enough information for traffic system analysis and model training.

[0149] Step E4, taking the virtual synthetic sample that meets the element quantity requirement as the target synthetic sample.

[0150] Specifically, if the label of the virtual synthetic sample meets the element quantity requirement, the sample is determined as the target synthetic sample. The target synthetic sample has been strictly tested and verified, has high quality and reliability, and can be used for various applications such as model training, performance testing, and algorithm optimization of the transportation system. Screening out target synthetic samples that meet the requirements can improve the accuracy and stability of the transportation system and provide more effective support for solving complex problems in actual transportation.

[0151] In the embodiment of the present application, after the virtual synthetic sample that meets the element quantity requirement is used as the target synthetic sample, the following steps F1-F5 are also included:

[0152] Step F1, obtaining real training samples of a specific vehicle model under sensor configuration.

[0153] Specifically, after the optimal sensor configuration is determined for a specific vehicle model through sensor simulation, in order to effectively train the deep learning model, it is necessary to obtain real training samples collected by the vehicle in the real traffic environment under this configuration. These samples contain rich information such as road conditions, traffic signs, and vehicle and pedestrian behaviors.

[0154] Step F2, constructing a first training set, a second training set and a test set based on the target synthetic samples and the real training samples, wherein the first training set includes multiple target synthetic samples and real training samples, the second training set includes multiple real training samples, and the test set includes multiple real training samples different from the second training set.

[0155] Specifically, three different sets are constructed based on target synthetic samples and real training samples. The first training set (training set X) consists of target synthetic samples and some real training samples. The target synthetic samples provide rich virtual scene data to help the model learn traffic scene characteristics, and the real samples enhance the authenticity and diversity of the data. The second training set (training set Y) is composed entirely of real training samples, providing accurate real-world information to enable the model to adapt to actual traffic conditions. The test set consists of real training samples different from the second training set, and is used to conduct objective comparative tests on models trained by different training sets.

[0156] As an example, assume that the internal model training set Y has 100,000 frames of real data. The training set X can be composed of 80,000 frames of target synthetic samples and 40,000 frames of real training samples. The target synthetic samples provide rich virtual traffic scenes, and the real training samples enhance authenticity and diversity. The training set Y consists of 100,000 frames of real data actually collected, reflecting the real traffic conditions. The test set Z uses 10,000 frames of real training samples that are different from the training sets X and Y to ensure that the comparative test results are objective and accurately evaluate the performance of models trained by different training sets.

[0157] Step F3: Use the first training set to train the preset deep learning model to obtain a first model, and use the second training set to train the preset deep learning model to obtain a second model.

[0158] Specifically, first, the first training set (training set X) is input into the preset deep learning model for training. The first training set contains virtual synthetic data and part of real data. The model adjusts parameters by learning the characteristics of virtual scenes and the actual conditions in the real data to adapt to different traffic scenes. After the training is completed, the first model is obtained. Secondly, the second training set (training set Y) is input into the preset deep learning model for training. The second training set consists of a large amount of real data. The model focuses on learning the characteristics and laws of real-world traffic scenes, and finally obtains the second model. The preset model is trained by different training sets to obtain the first model and the second model with different characteristics. The two models will be compared and tested on the test set Z to determine the validity of the synthetic data.

[0159] Step F4, input the test set into the first model and the second model respectively to obtain the evaluation results.

[0160] Specifically, the test set Z is composed of real data different from the second training set. The test set Z is input into the first model and the second model respectively. When the data of the test set Z is input into the first model, the first model will process and predict the data in the test set Z according to the characteristics of the virtual scene learned during the training process and the characteristics of some real data, and generate corresponding output results. When the data of the test set Z is input into the second model, the second model will process and predict the data in the test set Z according to the characteristics and laws of the real-world traffic scene learned from a large amount of real data, and also generate corresponding output results.

[0161] Step F5, if the evaluation result is that the first performance indicator output by the first model is higher than the second performance indicator output by the second model, the target synthetic sample is determined to be a valid sample; or, if the evaluation result is that the first performance indicator output by the first model is lower than the second performance indicator output by the second model, the target synthetic sample is determined to be an invalid sample.

[0162] Specifically, if the first performance index output by the first model is higher than the second performance index output by the second model, it means that the first model performs better when processing the test set, indicating that the combination of virtual synthetic data and a small amount of real data can effectively improve the performance of the deep learning model, and the target synthetic sample is determined to be a valid sample, which can be formally delivered for traffic system analysis and model training, etc. On the contrary, if the first performance index is lower than the second performance index, the second model performs better, indicating that the current virtual synthetic data does not play a positive role in improving the performance of the model, and the target synthetic sample is determined to be an invalid sample. The synthetic data is unqualified and needs to be modified and optimized to improve quality and effectiveness.

[0163] In the embodiments of the present application, Figure 4 As shown, a sample data synthesis method, the process is as follows:

[0164] First, multiple data sources in the data acquisition unit (mobile data acquisition equipment and vehicle-mounted data acquisition equipment) detect and obtain environmental perception data.

[0165] Secondly, the scene reconstruction module in the data processing unit constructs the original dynamic scene based on the environmental perception data. Then, the dynamic scene generation module simulates the activity trajectory of the dynamic environmental elements in the original dynamic scene to obtain multiple initial simulated dynamic scenes, and expands the static environmental elements to obtain expanded static environmental elements. After that, the sample generation module adds the expanded static environmental elements to the initial simulated scene to obtain multiple candidate simulated dynamic scenes, from which the target simulated dynamic scene is identified, and the scene description information of the preset scene is obtained. The scene description information is used to expand the scene features in the target simulated dynamic scene to obtain the expanded target simulated dynamic scene to construct the training sample. At the same time, the sensor simulation module uses the virtual vehicle model to perform sensor simulation on the training sample according to the camera parameters and sensor configuration of the specific vehicle model to generate a virtual synthetic sample. In addition, the scene migration module migrates the sensor configuration of the specific vehicle model to the real environment or conditions to obtain a real training sample.

[0166] Finally, the quality acceptance module in the data acceptance unit detects the image parameters of the virtual synthetic samples, the annotation acceptance module verifies whether the labels of the virtual synthetic samples meet the element quantity requirements, and the model acceptance module uses the virtual synthetic samples to train the first model and the real training samples to train the second model. It also uses the virtual synthetic samples and the real training samples to generate evaluation samples, and compares the evaluation results on different models based on the evaluation samples to verify the effectiveness of the target synthetic samples.

[0167] Figure 5 is a system block diagram of a sample data synthesis system according to an embodiment of the present invention. Figure 5 As shown, the system includes: a data acquisition unit 100, a data processing unit 200 and a data acceptance unit 300;

[0168] The data collection unit 100 includes: a mobile data collection device 101 and a vehicle-mounted data collection device 102;

[0169] The mobile data acquisition device 101 is used to detect environmental perception data, wherein the environmental perception data includes relevant data of dynamic environmental elements and static environmental elements.

[0170] The vehicle-mounted data acquisition device 102 is used to collect environmental perception data during the vehicle's driving process.

[0171] The data processing unit 200 includes: a scene reconstruction module 201, a dynamic scene generation module 202, a sample generation module 203, a sensor simulation module 204, and a scene migration module 205;

[0172] A scene reconstruction module 201 is used to construct an original dynamic scene based on the environmental perception data detected by multiple data acquisition devices, wherein the original dynamic scene includes an environmental background and an environmental element set, wherein the environmental element set includes multiple dynamic environmental elements and static environmental elements that affect vehicle driving;

[0173] The dynamic scene generation module 202 is used to simulate the activity tracks of the dynamic environment elements in the original dynamic scene to obtain a plurality of initial simulated dynamic scenes, and to expand the static environment elements to obtain expanded static environment elements;

[0174] The sample generation module 203 is used to add the expanded static environment elements to the initial simulated dynamic scene to obtain multiple candidate simulated dynamic scenes, then identify the target simulated dynamic scene from the multiple candidate simulated dynamic scenes, and obtain the scene description information of the preset scene, and use the scene description information to expand the scene features in the target simulated dynamic scene to obtain the expanded target simulated dynamic scene to construct a training sample;

[0175] The sensor simulation module 204 is used to perform sensor simulation on the training samples constructed by the expanded target simulation dynamic scene using a virtual vehicle model according to the camera parameters and sensor configuration of a specific vehicle model, and generate a virtual synthetic sample;

[0176] A scene migration module 205 is used to migrate the sensor configuration of a specific vehicle model to a real environment or condition to obtain a real training sample;

[0177] The data acceptance unit 300 includes: a quality acceptance module 301, a labeling acceptance module 302, and a model acceptance module 303;

[0178] The quality acceptance module 301 is used to detect the image parameters of the virtual synthetic sample generated by the sensor simulation module to ensure that the sample quality meets the preset conditions;

[0179] The labeling acceptance module 302 is used to verify whether the label of the virtual synthetic sample meets the element quantity requirement;

[0180] The model acceptance module 303 is used to verify the validity of the target synthetic sample constructed by the sample generation module by comparing the evaluation results of different samples on different models.

[0181] In this embodiment, a sample data synthesis device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0182] This embodiment provides a sample data synthesis device, such as Figure 6 As shown, including:

[0183] A construction module 61 is used to construct an original dynamic scene based on the environmental perception data, wherein the original dynamic scene includes an environmental background and an environmental element set, and the environmental element set includes dynamic environmental elements and static environmental elements that affect vehicle driving;

[0184] A simulation module 62, used to expand the dynamic environment elements and the static environment elements in the original dynamic scene, and construct a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements;

[0185] An identification module 63, used to identify a target simulated dynamic scene from a plurality of candidate simulated dynamic scenes, and obtain scene description information of a preset scene;

[0186] The processing module 64 is used to expand the scene features in the target simulated dynamic scene using the scene description information to obtain the expanded target simulated dynamic scene, and construct training samples using the expanded target simulated dynamic scene.

[0187] In an optional embodiment of the present application, the simulation module 62 includes:

[0188] The first simulation submodule is used to simulate the original activity track of the dynamic environment elements in the original dynamic scene to obtain multiple simulated activity tracks;

[0189] A processing submodule is used to move the dynamic environment elements according to a plurality of simulated activity trajectories in the original dynamic scene to obtain a plurality of initial simulated dynamic scenes;

[0190] The second simulation submodule is used to expand the static environment element to obtain an expanded static environment element;

[0191] The fusion submodule is used to add the expanded static environment elements to multiple initial simulated dynamic scenes to obtain multiple candidate simulated dynamic scenes.

[0192] In an optional embodiment of the present application, the first simulation submodule is used to obtain the original activity trajectory of each dynamic environmental element in the original dynamic scene; determine the potential factors affecting the movement of the dynamic environmental elements according to the environmental background, and generate multiple candidate simulated activity trajectories based on the original activity trajectory and the potential factors; remove the trajectories in the candidate simulated activity trajectories that do not conform to the actual traffic conditions, and obtain multiple simulated activity trajectories.

[0193] In an optional embodiment of the present application, the second simulation submodule is used to obtain a mapping relationship between different scene types and preset static environment elements, wherein the mapping relationship is used to represent multiple preset static environment elements matched by each scene type; based on the mapping relationship, multiple candidate static environment elements matching the scene type corresponding to the initial simulated dynamic scene are obtained, and the environment elements other than the static environment elements in the multiple candidate static environment elements are used as expanded static environment elements.

[0194] In an optional embodiment of the present application, the device also includes: a matching module, which is used to obtain multiple preset environmental elements and element attributes of the preset environmental elements; obtain multiple preset dynamic scenes, and semantic information corresponding to each preset dynamic scene; use the semantic information to match the element attributes to obtain a mapping relationship between the preset dynamic scene and the preset environmental element.

[0195] In an optional embodiment of the present application, the identification module 63 is used to obtain a pre-constructed dynamic scene set, wherein the dynamic scene set includes multiple pre-collected real dynamic scenes; calculate the similarity between the candidate simulated dynamic scene and each real dynamic scene; and use the candidate simulated dynamic scene with a similarity lower than a preset similarity as the target simulated dynamic scene.

[0196] In an optional embodiment of the present application, the processing module 64 is used to input the scene description information and the target simulated dynamic scene into a pre-trained language model; analyze the scene description information through the language model to obtain the target scene semantics corresponding to the scene description information, and identify the original three-dimensional image features of the target simulated dynamic scene through the language model; expand the original three-dimensional image features based on the mapping relationship between the preset scene semantics and the three-dimensional image features to obtain the target three-dimensional image features, and use the target three-dimensional image features to construct the expanded target simulated dynamic scene.

[0197] In an optional embodiment of the present application, the device also includes: a training module, which is used to obtain scene description samples, scene feature samples and three-dimensional image feature samples; fuse the scene description samples, scene feature samples and three-dimensional image feature samples to obtain fused samples; use the fused samples to input the language model to be trained to obtain the predicted three-dimensional image features output by the language model to be trained based on the fused samples; calculate the training loss of the language model to be trained based on the difference data between the predicted three-dimensional image features and the three-dimensional image feature samples; use the training loss to adjust the model parameters of the language model to be trained until the difference data between the predicted three-dimensional image features output by the language model after the adjustment of the parameters and the three-dimensional image feature samples is less than a preset value, thereby obtaining a trained language model.

[0198] In an optional embodiment of the present application, the device further includes: a simulation module for obtaining camera parameters of a specific vehicle model;

[0199] The camera parameters and training samples are input into the trained language model to obtain the sensor configuration for a specific vehicle model; a virtual vehicle model is built based on the sensor configuration, and the virtual vehicle model is used to perform sensor simulation on the training samples to obtain virtual synthetic samples.

[0200] In an optional embodiment of the present application, the device also includes: a verification module, used to detect image parameters of the virtual synthetic sample; determine whether the image parameters meet preset conditions; if the preset conditions are met, verify whether the label of the virtual synthetic sample meets the element quantity requirement; and use the virtual synthetic sample that meets the element quantity requirement as the target synthetic sample.

[0201] In an optional embodiment of the present application, the device also includes: a determination module, used to obtain real training samples of a vehicle of a specific model under a sensor configuration; constructing a first training set, a second training set and a test set based on the target synthetic samples and the real training samples, wherein the first training set includes multiple target synthetic samples and real training samples, the second training set includes multiple real training samples, and the test set includes multiple real training samples different from the second training set; using the first training set to train a preset deep learning model to obtain a first model, and using the second training set to train the preset deep learning model to obtain a second model; inputting the test set into the first model and the second model respectively to obtain an evaluation result; if the evaluation result is that the first performance indicator output by the first model is higher than the second performance indicator output by the second model, then the target synthetic sample is determined to be a valid sample; or, if the evaluation result is that the first performance indicator output by the first model is lower than the second performance indicator output by the second model, then the target synthetic sample is determined to be an invalid sample.

[0202] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 7As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system).

[0203] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0204] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0205] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0206] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0207] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0208] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0209] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for synthesizing sample data, characterized in that: The method comprises: Constructing an original dynamic scene based on the environmental perception data, wherein the original dynamic scene includes an environmental background and an environmental element set, and the environmental element set includes dynamic environmental elements and static environmental elements that affect vehicle driving; Expanding the dynamic environment elements and the static environment elements in the original dynamic scene, and constructing a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements; Identifying a target simulated dynamic scene from among the plurality of candidate simulated dynamic scenes, and acquiring scene description information of the preset scene; The scene features in the target simulation dynamic scene are expanded using the scene description information to obtain an expanded target simulation dynamic scene, and the expanded target simulation dynamic scene is used to construct a training sample.

2. The method according to claim 1, characterized in that The step of expanding the dynamic environment elements and the static environment elements in the original dynamic scene, and constructing a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements, includes: Simulating the original activity tracks of the dynamic environment elements in the original dynamic scene to obtain a plurality of simulated activity tracks; In the original dynamic scene, the dynamic environment elements are moved according to a plurality of simulated activity trajectories to obtain a plurality of initial simulated dynamic scenes; Expanding the static environment element to obtain an expanded static environment element; The expanded static environment elements are added to the plurality of initial simulated dynamic scenes to obtain a plurality of candidate simulated dynamic scenes.

3. The method according to claim 2, characterized in that The simulating the activity tracks of the dynamic environment elements in the original dynamic scene to obtain a plurality of simulated activity tracks includes: Obtaining the original activity track of each dynamic environment element in the original dynamic scene; Determining potential factors that affect the movement of the dynamic environmental element according to the environmental background, and generating a plurality of candidate simulated activity trajectories based on the original activity trajectory and the potential factors; Trajectories that do not conform to actual traffic conditions are removed from the candidate simulated activity trajectories to obtain multiple simulated activity trajectories.

4. The method according to claim 2, characterized in that: The step of expanding the static environment element to obtain an expanded static environment element includes: Acquire a mapping relationship between different scene types and preset static environment elements, wherein the mapping relationship is used to represent a plurality of preset static environment elements matched by each scene type; Based on the mapping relationship, a plurality of candidate static environment elements matching the scene type corresponding to the initial simulated dynamic scene are obtained, and the environment elements other than the static environment element among the plurality of candidate static environment elements are used as the expanded static environment elements.

5. The method according to claim 4, characterized in that Before obtaining the mapping relationship between different scene types and preset static environment elements, the method further includes: Acquire multiple preset environment elements and element attributes of the preset environment elements; Acquire multiple preset dynamic scenes and semantic information corresponding to each preset dynamic scene; The semantic information is matched with the element attributes to obtain a mapping relationship between the preset dynamic scene and the preset environmental element.

6. The method according to claim 1, characterized in that The step of identifying a target simulated dynamic scene from among the plurality of candidate simulated dynamic scenes comprises: Acquire a pre-built dynamic scene set, wherein the dynamic scene set includes a plurality of pre-collected real dynamic scenes; Calculating the similarity between the candidate simulated dynamic scene and each of the real dynamic scenes; The candidate simulated dynamic scene whose similarity is lower than a preset similarity is used as the target simulated dynamic scene.

7. The method according to claim 1, characterized in that The step of expanding the scene features in the target simulated dynamic scene by using the scene description information to obtain the expanded target simulated dynamic scene includes: Inputting the scene description information and the target simulated dynamic scene into a pre-trained language model; Analyzing the scene description information through the language model to obtain target scene semantics corresponding to the scene description information, and identifying original three-dimensional image features of the target simulated dynamic scene through the language model; The original three-dimensional image features are expanded based on the mapping relationship between the preset scene semantics and the three-dimensional image features to obtain the target three-dimensional image features, and the expanded target simulation dynamic scene is constructed using the target three-dimensional image features.

8. The method according to claim 7, characterized in that The training method of the language model includes: Obtaining scene description samples, scene feature samples, and three-dimensional image feature samples; fusing the scene description sample, the scene feature sample and the three-dimensional image feature sample to obtain a fused sample; Using the fused sample to input a language model to be trained, and obtaining predicted three-dimensional image features output by the language model to be trained based on the fused sample; Calculating the training loss of the language model to be trained based on the difference data between the predicted three-dimensional image features and the three-dimensional image feature samples; The training loss is used to adjust the model parameters of the language model to be trained until the difference data between the predicted three-dimensional image features output by the language model after the adjustment of the parameters and the three-dimensional image feature samples is less than a preset value, thereby obtaining a trained language model.

9. The method according to claim 1, characterized in that: After constructing the training samples by using the expanded target to simulate the dynamic scene, the method further includes: Get the camera parameters of a specific car model; Inputting the camera parameters and the training samples into a trained language model to obtain a sensor configuration for the specific vehicle model; A virtual vehicle model is built based on the sensor configuration, and the virtual vehicle model is used to perform sensor simulation on the training sample to obtain a virtual synthetic sample.

10. The method according to claim 9, characterized in that After performing sensor simulation on the training sample using the virtual vehicle model to obtain a virtual synthetic sample, the method further includes: detecting image parameters of the virtual synthetic sample; Determining whether the image parameters meet preset conditions; If the preset condition is met, verify whether the label of the virtual synthetic sample meets the element quantity requirement; A virtual synthetic sample that meets the requirement on the number of elements is used as a target synthetic sample.

11. The method according to claim 10, characterized in that After taking the virtual synthetic sample that meets the element quantity requirement as the target synthetic sample, the method further includes: Obtaining a real training sample of the vehicle of the specific model under the sensor configuration; Constructing a first training set, a second training set and a test set based on the target synthetic samples and the real training samples, wherein the first training set includes a plurality of the target synthetic samples and the real training samples, the second training set includes a plurality of real training samples, and the test set includes a plurality of real training samples different from the second training set; Using the first training set to train a preset deep learning model to obtain a first model, and using the second training set to train the preset deep learning model to obtain a second model; Inputting the test set into the first model and the second model respectively to obtain evaluation results; If the evaluation result is that the first performance indicator output by the first model is higher than the second performance indicator output by the second model, the target synthetic sample is determined to be a valid sample; or, if the evaluation result is that the first performance indicator output by the first model is lower than the second performance indicator output by the second model, the target synthetic sample is determined to be an invalid sample.

12. A sample data synthesis device, characterized in that: The device comprises: A construction module, used to construct an original dynamic scene based on the environmental perception data, wherein the original dynamic scene includes an environmental background and an environmental element set, and the environmental element set includes dynamic environmental elements and static environmental elements that affect vehicle driving; A simulation module, used to expand the dynamic environment elements and the static environment elements in the original dynamic scene, and construct a plurality of candidate simulated dynamic scenes using the expanded dynamic environment elements and the expanded static environment elements; An identification module, used to identify a target simulated dynamic scene from among the plurality of candidate simulated dynamic scenes, and obtain scene description information of a preset scene; The processing module is used to expand the scene features in the target simulation dynamic scene by using the scene description information to obtain the expanded target simulation dynamic scene, and construct training samples by using the expanded target simulation dynamic scene.

13. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 11 by executing the computer instructions.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Dynamic scene generation methods, electronic devices and computer-readable storage media

    CN122574174A