Intelligent driving data generation method and device and electronic equipment

By collecting image information around the vehicle and generating structured scene text, the problem of high cost and low efficiency in intelligent driving data generation is solved, achieving low-cost, high-efficiency data collection and improved privacy, and supporting rapid iterative optimization of intelligent driving algorithms.

CN121723975APending Publication Date: 2026-03-24CHINA AUTOMOTIVE INNOVATION CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing intelligent driving data generation solutions are costly, inefficient, inaccurate, and have issues with data authenticity and privacy, and cannot effectively support algorithm iteration and optimization.

Method used

Structured scene text is generated by analyzing image information around the vehicle and sent to the cloud for parsing to generate intelligent driving data. Combined with image information around the vehicle and metadata from the controller local area network, target scene data is constructed and transmitted to the cloud for processing.

Benefits of technology

It enables low-cost and high-efficiency data collection, improves the accuracy and privacy of data, and accelerates the iterative optimization cycle of intelligent driving algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121723975A_ABST
    Figure CN121723975A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an intelligent driving data generation method and device and electronic equipment, and the method comprises the steps: collecting first target image data in a target region which is a region around a vehicle; performing semantic analysis processing on the first target image data to obtain a target scene semantic information set corresponding to the first target image data; determining a target structured template according to the target scene semantic information set; constructing a target structured scene text based on the target structured template; and based on the target structured scene text, generating target scene data and sending the target scene data to the cloud server. According to the intelligent driving data generation method, semantics are analyzed according to the image information around the vehicle, the structured scene text is generated and sent to the cloud, the cloud analyzes the scene text to generate the intelligent driving data, low cost, high efficiency and accurate orientation of data acquisition are realized, data privacy is improved, and the intelligent driving data acquisition efficiency is improved. And meanwhile, the iterative optimization period of the intelligent driving algorithm is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent driving, in particular to an intelligent driving data generation method and device and electronic equipment. BACKGROUND

[0002] With the evolution of high-order intelligent driving technology, the iterative optimization of its system performance highly depends on large-scale, high-diversity and accurately labeled training data. This data-driven core logic has become an industry consensus.

[0003] The current intelligent driving data generation scheme still takes vehicle-side collection and cloud-side processing as the basic framework. By deploying a large-scale test vehicle fleet, relying on vehicle-mounted cameras, LiDAR and other sensors to record continuous video streams and point cloud data, and then uploading scenes such as identifying sudden braking and collision warning based on rules or simple AI triggers, part of the data is filtered to the cloud. Subsequently, a large amount of resources need to be invested for storage, cleaning and manual labeling, and finally a usable training data set is formed. At the same time, although cloud AIGC technologies such as diffusion models can generate multi-modal content such as images and videos to try to make up for the coverage gap of real data, such generation processes mostly rely on general semantic prompts and lack deep binding with real driving scenarios on the vehicle side, resulting in deviations between the generated content and the actual needs of intelligent driving models. The authenticity, physical consistency and usability of the data are difficult to guarantee, and high-quality real labeled data cannot be truly replaced. This approach results in high acquisition costs, generating massive amounts of raw data, and huge costs for storage, transmission and manual labeling. At the same time, the proportion of effective data such as dangerous and rare scenarios is extremely low, and the efficiency of filtering target segments from the data ocean is extremely low, with a large amount of resources wasted on processing useless data, and the data production efficiency is severely limited. In addition, continuous video streams and point cloud data contain personal sensitive information such as faces, license plates and precise geographic coordinates, and face increasingly stringent data security regulations and restrictions, cannot generate data targeted at model weaknesses, and the data diversity is limited to real events that have occurred, making it difficult to cope with complex and variable real-world environments, lacking controllability and generalization ability.

[0004] Therefore, it is particularly important to develop an intelligent driving data generation method, device and electronic equipment that can analyze semantics based on image information around the vehicle, generate structured scenario text and send it to the cloud, the cloud analyzes the scenario text to generate intelligent driving data, realizes low-cost, high-efficiency and precise targeting of data acquisition, improves data privacy, and accelerates the iterative optimization cycle of intelligent driving algorithms. SUMMARY

[0005] To address the aforementioned technical issues, this application provides a method, device, and electronic device for generating intelligent driving data. By analyzing semantics based on image information surrounding the vehicle, structured scene text is generated and sent to the cloud. The cloud then parses the scene text to generate intelligent driving data. This solution addresses the current lack of a method, device, and electronic device for generating intelligent driving data that achieves low-cost, high-efficiency, and precise data collection, improves data privacy, and accelerates the iterative optimization cycle of intelligent driving algorithms.

[0006] The technical solution provided in this application is as follows: On one hand, this application provides a method for generating intelligent driving data, which is applied in a vehicle, and the method includes: Acquire first target image data within the target area, wherein the target area is the area surrounding the vehicle; The first target image data is subjected to semantic parsing processing to obtain a set of target scene semantic information corresponding to the first target image data; Based on the set of semantic information of the target scene, determine the target structured template; Based on the target structured template, construct the target structured scene text; Based on the target structured scene text, target scene data is generated and sent to a cloud server so that the cloud server can parse the target scene data to obtain a target control command vector; and the cloud server can perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0007] In some optional implementations, the step of performing semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data includes: Obtain a first target model; the first target model is obtained by semantic parsing and training a preset model based on historical image data and scene semantic information labels corresponding to the historical image data. The first target image data is input into the first target model to obtain a set of target scene semantic information corresponding to the first target image data.

[0008] In some optional implementations, the target scene semantic information set includes several pieces of scene semantic information, and determining the target structured template based on the target scene semantic information set includes: Obtain preset mapping information, which represents the mapping relationship between preset conditions and structured templates, and the preset conditions include target preset conditions; If at least one piece of scene semantic information in the target scene semantic information set satisfies the target preset condition, a structured template corresponding to the target preset condition is determined based on the preset mapping information. Use the structured template corresponding to the target preset conditions as the target structured template.

[0009] In some optional implementations, the scene semantic information includes scene attributes and scene data corresponding to the scene attributes. The step of constructing the target structured scene text based on the target structured template includes: The target structured template is parsed to obtain several template scene attributes and the initial data corresponding to the several template scene attributes; The initial data corresponding to each template scene attribute is updated to the target scene data corresponding to the target scene attribute to obtain the target structured scene text; Wherein, the target scene attribute is the scene attribute in the target scene semantic information, and the target scene semantic information is any scene semantic information in the target scene semantic information set.

[0010] In some optional implementations, generating target scene data based on the target structured scene text and sending the target scene data to a cloud server includes: Reacquire image data within the target area to obtain second target image data; Obtain controller LAN metadata; The target structured scene text, the second target image data, and the controller local area network metadata are used as the target scene data; The target scene data is sent to the cloud server.

[0011] On the other hand, this application provides a method for generating intelligent driving data, which is applied in a cloud server and includes: The system receives target scene data sent by a vehicle. The target scene data is generated by the vehicle based on target structured scene text. The target structured scene text is constructed by the vehicle according to a target structured template. The target structured template is determined by the vehicle based on a set of target scene semantic information corresponding to the first target image data. The set of target scene semantic information is obtained by the vehicle by parsing the first target image data within the target area. The target area is the area surrounding the vehicle. The target scene data is parsed to obtain the target control command vector; The target control command vector is processed to generate target intelligent driving data.

[0012] In some optional implementations, the step of performing data generation processing on the target control command vector to obtain target intelligent driving data includes: Obtain a second target model; the second target model is obtained by generating and training a preset model based on historical control command vectors and the corresponding intelligent driving data of the historical control command vectors. The target control command vector is input into the second target model to obtain the target intelligent driving data corresponding to the target control command vector.

[0013] On the other hand, this application provides an intelligent driving data generation device, the intelligent driving data generation device comprising: The image acquisition module is used to acquire first target image data within a target area, wherein the target area is the area surrounding the vehicle; The semantic parsing module is used to perform semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data; The text generation module is used to determine a target structured template based on the target scene semantic information set; and to construct target structured scene text based on the target structured template. The data transmission module is used to generate target scene data based on the target structured scene text, and send the target scene data to a cloud server so that the cloud server receives the target scene data sent by the vehicle; and to enable the cloud server to parse the target scene data to obtain a target control command vector; and to enable the cloud server to perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0014] On the other hand, this application provides an intelligent driving data generation device, the intelligent driving data generation device comprising: The data receiving module is used to receive target scene data sent by the vehicle; the target scene data is generated by the vehicle based on target structured scene text, the target structured scene text is constructed by the vehicle according to the target structured template, the target structured template is determined by the vehicle according to the target scene semantic information set corresponding to the first target image data, the target scene semantic information set is obtained by the vehicle by parsing the first target image data in the target area, and the target area is the area around the vehicle; The data parsing module is used to parse the target scene data to obtain the target control command vector; The data generation module is used to perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0015] On the other hand, this application provides an electronic device including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the intelligent driving data generation method as described in any of the above embodiments.

[0016] On the other hand, this application provides a computer-readable storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the intelligent driving data generation method as described in any of the above embodiments.

[0017] This application provides a method for generating intelligent driving data, applied in a vehicle. The method includes: acquiring first target image data within a target area, the target area being the area surrounding the vehicle; performing semantic parsing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data; determining a target structured template based on the target scene semantic information set; constructing target structured scene text based on the target structured template; generating target scene data based on the target structured scene text, and sending the target scene data to a cloud server, so that the cloud server parses the target scene data to obtain a target control command vector; and having the cloud server perform data generation processing on the target control command vector to obtain target intelligent driving data. By analyzing the semantics of image information surrounding the vehicle, generating structured scene text, and sending it to the cloud, the cloud server parses the scene text to generate intelligent driving data. This achieves low-cost, high-efficiency, and precise data acquisition, improves data privacy, and accelerates the iterative optimization cycle of intelligent driving algorithms. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of a vehicle intelligent driving data generation method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall process of an intelligent driving data generation method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a cloud server intelligent driving data generation method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a vehicle intelligent driving data generation device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a cloud server intelligent driving data generation device according to an embodiment of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0021] The term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of this application. In the description of this application, it should be understood that the terms "upper," "lower," "top," "bottom," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Moreover, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.

[0022] When a numerical range is disclosed herein, the range is considered continuous and includes the minimum and maximum values ​​of the range, as well as every value between the minimum and maximum values. Furthermore, when the range refers to an integer, it includes every integer between the minimum and maximum values ​​of the range. Additionally, when multiple ranges are provided to describe a feature or characteristic, the ranges may be combined. In other words, unless otherwise specified, all ranges disclosed herein should be understood to include any and all subranges to which they are included. For example, a specified range from “1 to 10” should be considered to include any and all subranges between the minimum value 1 and the maximum value 10. Exemplary subranges of the range 1 to 10 include, but are not limited to, 1 to 6.1, 3.5 to 7.8, 5.5 to 10, etc.

[0023] Because current intelligent driving data generation solutions deviate from the actual needs of intelligent driving models, the authenticity, physical consistency, and usability of the data are difficult to guarantee. They cannot truly replace high-quality, real-world labeled data, resulting in high collection costs, massive amounts of raw data, and huge costs for storage, transmission, and manual annotation. Furthermore, the proportion of effective data, such as hazardous or rare scenarios, is extremely low, severely limiting data production efficiency. Moreover, they cannot generate data specifically for model weaknesses, lacking controllability and generalization ability. Therefore, to achieve low-cost, high-efficiency, and precise data collection, improve data privacy, and accelerate the iterative optimization cycle of intelligent driving algorithms, this application provides an intelligent driving data generation method and apparatus.

[0024] Please see Figure 1 , Figure 1 This is a schematic flowchart of a vehicle intelligent driving data generation method according to an embodiment of the present invention. In one aspect, this application provides an intelligent driving data generation method, which is applied in a vehicle, and the method includes: S101. Acquire first target image data within the target area, wherein the target area is the area surrounding the vehicle.

[0025] Optionally, the acquisition of the first target image data is achieved through dedicated vehicle-mounted sensing hardware, which can be an onboard camera sensor, such as a multi-view camera module pre-installed in the vehicle, rather than a single camera, including but not limited to front-view cameras, side-view cameras, rear-view cameras, and surround-view cameras. The raw image data acquired by the camera is transmitted to the vehicle domain controller in real time, where it is synchronously received by an embedded lightweight AI model deployed in the domain controller and prepared for subsequent semantic parsing.

[0026] Optionally, the target area refers to the spatial area directly related to intelligent driving decisions and effectively covered by the vehicle-mounted camera. The target area needs to accurately capture environmental elements that affect driving behavior. The acquisition frequency of the first target image data is 2-10 Hz, which can ensure real-time capture of dynamic environmental changes while avoiding overload of vehicle-side storage caused by high-frequency acquisition. For example, 2-5 Hz can be used in high-speed driving scenarios, and 5-10 Hz can be increased in complex urban road conditions. The frequency can be adaptively adjusted according to the scenario.

[0027] Optionally, when acquiring the first target image data, the vehicle's GPS location information and timestamp are simultaneously associated. Each frame of the first target image data carries a spatiotemporal tag. During subsequent semantic analysis, the road type can be determined by combining GPS data and the lighting conditions by combining timestamps, providing a foundation for constructing structured scene description text with temporal information. The first target image data is used locally on the vehicle for real-time semantic analysis. Continuous video streams are not stored. After acquisition, the data is directly input into an embedded lightweight AI model for feature extraction and semantic understanding. Only the semantic results are retained after analysis. The original first target image data is not stored long-term, further reducing the storage load on the vehicle.

[0028] S102. Perform semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data.

[0029] In an optional embodiment, the step of performing semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data includes: Obtain a first target model; the first target model is obtained by semantic parsing and training a preset model based on historical image data and scene semantic information labels corresponding to the historical image data. The first target image data is input into the first target model to obtain a set of target scene semantic information corresponding to the first target image data.

[0030] Optionally, the first target model is the aforementioned embedded lightweight AI model. The preset model serves as the basic framework for the first target model, selected based on the vehicle's computing power and storage resources. The preset model typically uses a lightweight convolutional neural network (CNN) or a lightweight variant of the visual transformer (ViT). Training the first target model relies on historical image data and corresponding scene semantic information. Both types of data must closely match the core scenario of intelligent driving. Specifically, the historical image data covers multi-dimensional scene variables, including environmental variables, road variables, and interaction variables. Environmental variables include different weather conditions, lighting, and time of day; road variables include different road types and traffic facilities; and interaction variables include different traffic participants, participant behaviors, and the driver's behavior. The images must contain multi-view data from the vehicle, consistent with the perspective of the actually collected first target image data. Label construction typically employs a combination of manual annotation and automated verification. First, annotators label the semantic information based on the images, and then a rule engine verifies the consistency of the labels to ensure accuracy.

[0031] Optionally, using historical image data as input and structured semantic labels as the supervision target, a multi-task loss function is constructed to train the model to learn the ability to map image features to semantic labels. After training, model size and inference latency are further reduced through model quantization, layer pruning, and knowledge distillation to meet the real-time requirements of the vehicle. The primary target model can be dynamically updated via over-the-air (OTA) updates. When the cloud detects that the model has low accuracy in semantic parsing for a certain type of scene, it will fine-tune the model using newly labeled data for that scene and then push it to the vehicle via OTA to continuously optimize semantic parsing capabilities.

[0032] Optionally, the first target image data is preprocessed to ensure that it is consistent with the input format during model training. The preprocessed image data is input into the first target model, and the model performs inference calculations on the embedded chip in the vehicle. After the model inference is completed, the output target scene semantic information set is in the form of structured key-value pairs, including five semantics: scene environment, road elements, traffic participants, main vehicle behavior, and driving suggestions. A structured labeling system is adopted to ensure that the model output can be directly used for subsequent template matching.

[0033] By utilizing the text prompting engineering principle, heavy image data can be transformed into lightweight semantic information, avoiding the need for vehicle-side storage and transmission of continuous video streams. Only KB-level structured text needs to be generated subsequently, greatly reducing the load on the vehicle. The output target scene semantic information set serves as the basis for subsequently constructing target control command vectors. The cloud needs to parse scene elements based on this set and then combine it with keyframe snapshots and CAN bus data to generate multimodal intelligent driving data, thereby improving accuracy.

[0034] S103. Determine the target structured template based on the target scene semantic information set.

[0035] In an optional embodiment, the target scene semantic information set includes several pieces of scene semantic information, and determining the target structured template based on the target scene semantic information set includes: Obtain preset mapping information, which represents the mapping relationship between preset conditions and structured templates, and the preset conditions include target preset conditions; If at least one piece of scene semantic information in the target scene semantic information set satisfies the target preset condition, a structured template corresponding to the target preset condition is determined based on the preset mapping information. Use the structured template corresponding to the target preset conditions as the target structured template.

[0036] Optionally, the preset mapping information is a rule base corresponding to the scene triggering conditions and structured text templates pre-configured on the vehicle side. It is an association standard pre-configured and distributed to the vehicle side by the cloud according to the training requirements of the intelligent driving model. The preset mapping information is not a one-to-one list, but a structured association system that covers multiple scenarios and can be dynamically adjusted. The preset conditions are the prompting engineering strategy, that is, the scene judgment criteria for triggering template matching. It can be a single condition, which can be triggered as long as a certain semantic information is satisfied, adapting to simple scenarios, or it can be a combination of conditions, which can be triggered as long as multiple semantic information is satisfied simultaneously, adapting to complex interaction scenarios. The parameter thresholds of all preset conditions can be flexibly configured by the cloud. The structured template is a scene semantic text framework corresponding to the preset conditions, and reserves real-time parameter placeholders.

[0037] Optionally, the preset mapping information is not static data fixed on the vehicle side, but has the ability to be configured in the cloud and dynamically updated via OTA. The cloud predefines the mapping relationship between conditions and templates according to the current training needs of the intelligent driving model, and distributes it to the storage module of each vehicle through the vehicle-side gateway. When the cloud finds that the generated data for a certain type of scenario is insufficient, or the accuracy of the text parsing generated by the model for a certain type of template is low, it will adjust the corresponding preset conditions or optimize the template structure, and then push the updated mapping information to the vehicle side via OTA to ensure that the template and model requirements are always synchronized.

[0038] Optionally, the target preset conditions are high-value scene trigger conditions marked in the preset mapping information by the cloud based on the current model training priority. The vehicle will check whether the target preset conditions are met one by one in order of priority from high to low. The determination of a single condition only requires that the value of a certain scene semantic information completely matches the condition. The determination of combined conditions requires that the semantic information corresponding to all sub-conditions be satisfied. The determination process is carried out synchronously with the semantic parsing of the first target model, without the need to call additional heavy algorithms.

[0039] Optionally, simple scenarios use a single condition and template to ensure strong template targeting, while complex scenarios use multiple conditions and a single template to avoid template redundancy. Regardless of the correspondence, as long as a certain target's preset condition is met, a unique structured template can be located. Transforming semantic data into matchable judgment criteria can prevent semantic information from being rendered useless, ensuring that subsequent real-time parameters can be organized in a unified format to generate structured text that can be parsed in the cloud; templates also prevent the generation of useless text and increase data value density.

[0040] S104. Based on the target structured template, construct the target structured scene text.

[0041] In an optional embodiment, the scene semantic information includes scene attributes and scene data corresponding to the scene attributes, and the step of constructing target structured scene text based on the target structured template includes: The target structured template is parsed to obtain several template scene attributes and the initial data corresponding to the several template scene attributes; The initial data corresponding to each template scene attribute is updated to the target scene data corresponding to the target scene attribute to obtain the target structured scene text; Wherein, the target scene attribute is the scene attribute in the target scene semantic information, and the target scene semantic information is any scene semantic information in the target scene semantic information set.

[0042] Optionally, template scene attributes refer to the semantic dimensions pre-defined in the target structured template to describe the core features of the scene, while initial data refers to the placeholders corresponding to the template scene attributes. These are reserved blank spaces in the template, used to subsequently fill in the real scene data extracted from the semantic information set. Specifically, the semantic information corresponding to the target structured template is first identified. Under each semantic level, the template scene attributes explicitly listed in the template are extracted. For each extracted template scene attribute, the corresponding placeholder initial data is matched. Template parsing must ensure that the extracted template scene attributes can be found in the output target scene semantic information set, avoiding invalid text with frames but no data.

[0043] Optionally, the target scene attribute refers to the semantic dimension in the target scene semantic information set that has the same name and semantic meaning as the template scene attribute. The target scene data refers to the specific value corresponding to the target scene attribute, which can be quantified using CAN bus data.

[0044] Optionally, data updates are only performed when the template scene attributes are completely consistent with the target scene attributes. The target scene data must be updated according to the preset format requirements of the template to avoid format confusion that could lead to cloud parsing failure. If the template attributes involve multiple semantic levels, semantic information from multiple levels must be selected from the set to complete the data update of all template attributes. A scene typically lasts 10-30 seconds, usually measured in clips. The target structured scene text contains temporal information.

[0045] It transforms scattered semantic information into structured, machine-readable text with a file size of only KB, eliminating the need to store or transmit continuous video streams and drastically reducing the load on the vehicle. The cloud can accurately construct control command vectors based on the explicit attributes and quantified data in the text, avoiding data distortion caused by semantic ambiguity.

[0046] S105. Based on the target structured scene text, generate target scene data and send the target scene data to the cloud server so that the cloud server can parse the target scene data to obtain a target control command vector; and enable the cloud server to perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0047] In an optional embodiment, generating target scene data based on the target structured scene text and sending the target scene data to a cloud server includes: Reacquire image data within the target area to obtain second target image data; Obtain controller LAN metadata; The target structured scene text, the second target image data, and the controller local area network metadata are used as the target scene data; The target scene data is sent to the cloud server.

[0048] Optionally, the second target image data is consistent with the first target image, and is a keyframe snapshot taken by the vehicle-mounted camera module. A front-view camera is preferred because its viewpoint covers the core direction of vehicle movement, and the scene details it captures are crucial for generating multimodal driving perspective data in the cloud. The acquisition timing is spatiotemporally synchronized with the generation of the target structured scene text, and must completely correspond to the visual image captured by the second target image to ensure that the appearance generated in the cloud is consistent with the real scene. When generating videos and point clouds in the cloud, the second target image is used as the initial frame appearance reference to ensure that the textures, colors, and lighting directions of the generated vehicles, roads, and pedestrians are consistent with the real scene.

[0049] Optionally, the controller area network metadata, i.e., the vehicle CAN bus metadata, is used to provide physical dynamic constraints for cloud-based multimodal generation. The CAN bus is the communication network of various electronic control units (ECUs) within the vehicle, and its metadata consists of parameters that quantify the vehicle's state and behavior. Specifically, it can be divided into vehicle motion state, vehicle control signals, and spatiotemporal synchronization identifiers. The CAN bus metadata is synchronized with the acquisition of the second target image and the generation of structured text. The vehicle domain controller (DCU) reads the CAN bus data in real time, and the timestamp of the reading must be completely consistent with the timestamp of the second target image capture and the timestamp of the text generation to ensure that the three correspond to the same instantaneous scene. The CAN bus metadata provides physical constraints for the cloud-generated model, avoiding the generation of contradictory scenes, quantifying vehicle state information, and making the cloud-generated scene more accurate. After the cloud generates multimodal data, it uses CAN data for reverse verification to ensure the reliability of the generated data.

[0050] Optionally, the target structured scene text, the second target image data, and the controller LAN metadata data packet are constructed into a very small data packet using a structured format. The collection timestamps of the three data modules must be consistent, and the total timestamp of the data packet is uniformly marked in the header information. The total size of the data packet is controlled in the KB to MB range, which greatly reduces the storage pressure on the vehicle and the bandwidth usage. After the data packet is assembled, the vehicle does not store the data, but only temporarily caches it in memory. It is deleted after the transmission is completed, which reduces the storage resource usage on the vehicle and avoids the risk of privacy leakage caused by continuous data retention.

[0051] Optionally, 5G cellular networks are prioritized for transmitting target scene data. Data packets are transmitted immediately after being assembled without batch caching, ensuring that the cloud can quickly obtain high-value scene data for timely model training and accelerate algorithm iteration. If the vehicle generates multiple scene data packets simultaneously, they can be sorted by risk level, with high-risk scene data transmitted first, ensuring that the cloud obtains the training data most urgently needed by the model.

[0052] Text controls the content generation, images control the style generation, and CAN bus metadata controls the dynamic generation. Data packets are only in the KB to MB range. Compared with existing technologies, the vehicle-side storage and transmission costs are reduced by several orders of magnitude, making it possible to deploy on a large-scale fleet. Furthermore, there is no continuous video stream, only single-frame images, text, and CAN data, avoiding the problem of continuous recording of privacy information. No data is retained on the vehicle side after transmission, meeting compliance requirements from the source and simplifying the de-identification process.

[0053] Please see Figure 2 , Figure 3 , Figure 2 This is a schematic diagram of the overall process of an intelligent driving data generation method according to an embodiment of the present invention. Figure 3 This is a schematic flowchart of a cloud server-based intelligent driving data generation method according to an embodiment of the present invention. On the other hand, this application provides an intelligent driving data generation method applied in a cloud server, the method comprising: S301. Receive target scene data sent by the vehicle; the target scene data is generated by the vehicle based on target structured scene text, the target structured scene text is constructed by the vehicle according to the target structured template, the target structured template is determined by the vehicle according to the target scene semantic information set corresponding to the first target image data, the target scene semantic information set is obtained by the vehicle by parsing the first target image data in the target area, and the target area is the area around the vehicle.

[0054] Optionally, the cloud can adopt a distributed receiving architecture, which can be configured with integrity and consistency checks to prevent invalid data from entering the parsing process, such as data packets missing CAN data. Even if parsed, these packets cannot generate physically consistent multimodal data, thus saving subsequent computing power in the cloud.

[0055] S302. Analyze the target scene data to obtain the target control command vector.

[0056] Optionally, after receiving the data packet from the vehicle, the cloud performs target scene data decomposition, accurately deconstructing scene elements including the aforementioned scene environment, road elements, traffic participants, vehicle behavior, driving suggestions, spatiotemporal relationships, CAN bus metadata, and keyframe snapshots. The structured text, second target image, and CAN data in the vehicle-side data packet are separated into independent modules to avoid format confusion. Natural Language Processing (NLP) technology is used to transform the semantic description in the text into machine-recognizable structured scene elements, i.e., key-value pair format scene elements. The appearance features of the second target image and the physical parameters of the CAN bus metadata are extracted as supplements and verifications of the scene elements. The text parsing elements, image appearance features, and CAN physical parameters are integrated into a numerical control command vector to ensure that the generated model can be directly read.

[0057] S303. Perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0058] In an optional embodiment, the step of performing data generation processing on the target control command vector to obtain target intelligent driving data includes: Obtain a second target model; the second target model is obtained by generating and training a preset model based on historical control command vectors and the corresponding intelligent driving data of the historical control command vectors. The target control command vector is input into the second target model to obtain the target intelligent driving data corresponding to the target control command vector.

[0059] Optionally, obtaining the second target model is the foundation for data generation. This model is a heavy-duty, conditionally controllable multimodal generative model specifically designed for intelligent driving scenarios. The preset model is the basic architecture of the second target model and can be a diffusion model or a generative Transformer architecture. The historical control command vectors used to train the second target model originate from the parsing results of data uploaded from the vehicle in the past, that is, the historical control command vectors generated by parsing historical target scenario data. The historical intelligent driving data corresponding to the historical control command vectors comes from a large-scale cleaned and labeled real driving dataset, which needs to be spatiotemporally aligned with the corresponding historical control command vectors to ensure that the model learns the mapping relationship between commands and data.

[0060] Optionally, historical control command vectors are used as model inputs. During model inference, the information from the control command vectors is integrated into the generation process. For example, in the denoising step of the diffusion model, the semantic, physical, and appearance dimensions of the vectors are converted into attention weights for the denoising network, constraining the denoising direction. In the generative Transformer, the vectors are used as global conditions, participating in self-attention calculations along with temporal data to ensure that each frame's generation meets the command requirements. A multi-task loss function is employed, and iterative training is performed using mini-batch gradient descent. After each training round, the generation quality is evaluated using a validation set, and the model's hyperparameters are adjusted. In the later stages of training, the data generated by the model is used to train the intelligent driving model. If the model performs poorly in a certain scenario, the vectors and data from that scenario are used to further fine-tune the second target model.

[0061] Optionally, the target control command vector is input into the second target model, and the resulting target intelligent driving data includes, but is not limited to, synthesized video sequences, synthesized LiDAR point cloud sequences, automatically labeled ground truth generation, and vehicle signal reconstruction. Specifically, the diffusion model receives the target control command vector and generates a multi-view, high-fidelity dynamic video stream with a duration of 10-30 seconds; generates a point cloud sequence that is spatiotemporally synchronized with the video frames and conforms to physical reflection characteristics; automatically generates pixel-level semantic segmentation maps, instance-level 2D / 3D bounding boxes, depth maps, and other annotation files; and generates a complete CAN bus signal time series that matches the dynamics of the synthesized scene.

[0062] Optionally, physical consistency enhancement and quality verification are performed on the generated data. Keyframe snapshots uploaded from the vehicle are used to calibrate the appearance of the starting frame of the synthesized sequence. A physical rule engine is used to perform dynamic rationality verification and post-processing on the generated data to ensure its usability. Specifically, collision detection can be used to check for penetration between objects, such as vehicles passing through walls or pedestrians passing through vehicles; motion rationality rules can be used to verify whether the object's trajectory conforms to physical laws, such as matching vehicle speed and displacement, and ensuring natural pedestrian gait; appearance consistency verification compares the generated data with the second target image on the vehicle, checking for lighting and texture, such as whether rainwater marks penetrate all video frames. If the verification fails, post-processing is triggered. The high-fidelity multimodal scene data that finally passes verification is stored in the intelligent driving training dataset, and a semantic-based index is established for easy subsequent retrieval and use.

[0063] Optionally, the method further includes a closed-loop optimization process, using the performance improvement effect of cloud-synthesized data after model training as a reward signal. Reinforcement learning (RL) is used to dynamically optimize the trigger threshold of the vehicle-side prompting strategy and the generation parameters of the cloud-generated model. The vehicle-side prompting strategy refers to filtering high-value scenarios based on trigger thresholds. Reinforcement learning optimization aims to adjust the thresholds to make the scenarios collected by the vehicle more closely match the model. If the scenario data generated by a certain template brings high rewards, its trigger threshold is lowered to encourage more vehicle-side collection; if the reward is low, it indicates that the scenario is worthless or redundant, so the threshold is raised to reduce collection. The generation parameters of the cloud-generated model directly determine the data quality. Reinforcement learning optimization aims to adjust the parameters to make the generated data better improve model performance. If the generated data corresponding to a certain type of parameter brings high rewards, it indicates that the current parameter setting is reasonable or needs strengthening, so the parameters are adjusted to amplify the advantages; if the reward is low, the parameters are corrected to improve the shortcomings. The optimized thresholds and parameters result in a greater improvement in model performance and a higher reward signal. Through continuous iteration, vehicle-side collection becomes increasingly accurate, cloud generation becomes increasingly high-quality, and model iteration becomes increasingly rapid, ultimately achieving the goal of self-evolution.

[0064] The second target model can generate intelligent driving data based on control command vectors. The generated data comes with high-quality automatic annotation, eliminating the need for manual annotation and greatly shortening the data production cycle. Through control command vector constraints and physics engine verification, the generated data not only meets appearance conditions but also physical conditions and is semantically correct. It can be directly used for intelligent driving model training, avoiding the problem of unusable generated data. In addition, the generated target intelligent driving data can be used for intelligent driving model training. The performance of the model in specific scenarios can serve as a feedback signal to back-optimize the parameters of the second target model and the vehicle-side template triggering condition strategy, thereby achieving system self-evolution.

[0065] On the other hand, please see Figure 4 , Figure 4 This is a schematic diagram of a vehicle intelligent driving data generation device according to an embodiment of the present invention. The present application provides an intelligent driving data generation device, which includes: Image acquisition module 401 is used to acquire first target image data within a target area, wherein the target area is the area around the vehicle; The semantic parsing module 402 is used to perform semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data; The text generation module 403 is used to determine a target structured template based on the target scene semantic information set; and to construct target structured scene text based on the target structured template. The data sending module 404 is used to generate target scene data based on the target structured scene text, and send the target scene data to a cloud server so that the cloud server receives the target scene data sent by the vehicle; and to enable the cloud server to parse the target scene data to obtain a target control command vector; and to enable the cloud server to perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0066] In an optional embodiment, the semantic parsing module 402 includes: The first target model acquisition unit is used to acquire a first target model; the first target model is obtained by semantic parsing and training a preset model based on historical image data and scene semantic information labels corresponding to the historical image data. The target scene semantic information set generation unit is used to input the first target image data into the first target model to obtain the target scene semantic information set corresponding to the first target image data.

[0067] In an optional embodiment, the text generation module 403 includes: A preset mapping information acquisition unit is used to acquire preset mapping information, wherein the preset mapping information represents the mapping relationship between preset conditions and structured templates, and the preset conditions include target preset conditions; The target structured template determination unit is configured to, when at least one piece of scene semantic information in the target scene semantic information set satisfies the target preset condition, determine the structured template corresponding to the target preset condition based on the preset mapping information; and use the structured template corresponding to the target preset condition as the target structured template. The target structured template parsing unit is used to parse the target structured template to obtain several template scene attributes and the initial data corresponding to the several template scene attributes; The target structured scene text generation unit is used to update the initial data corresponding to each template scene attribute to the target scene data corresponding to the target scene attribute, thereby obtaining the target structured scene text; wherein, the target scene attribute is the scene attribute in the target scene semantic information, and the target scene semantic information is any scene semantic information in the target scene semantic information set.

[0068] In an optional embodiment, the data sending module 404 includes: The second target image data acquisition unit is used to re-acquire image data within the target area to obtain second target image data; and to acquire controller local area network metadata. The target scene data sending unit is used to send the target structured scene text, the second target image data, and the controller local area network metadata as target scene data; and to send the target scene data to the cloud server.

[0069] On the other hand, please see Figure 5 , Figure 5 This is a schematic diagram of a cloud server-based intelligent driving data generation device according to an embodiment of the present invention. The present application provides an intelligent driving data generation device, which includes: The data receiving module 501 is used to receive target scene data sent by the vehicle; the target scene data is generated by the vehicle based on target structured scene text, the target structured scene text is constructed by the vehicle according to the target structured template, the target structured template is determined by the vehicle according to the target scene semantic information set corresponding to the first target image data, the target scene semantic information set is obtained by the vehicle by parsing the first target image data in the target area, and the target area is the area around the vehicle; Data parsing module 502 is used to parse the target scene data to obtain the target control command vector; The data generation module 503 is used to perform data generation processing on the target control command vector to obtain target intelligent driving data.

[0070] In an optional embodiment, the data generation module 503 includes: The second target model acquisition unit is used to acquire the second target model; the second target model is obtained by generating and training a preset model based on intelligent driving data corresponding to control command vectors and historical control command vectors. The target intelligent driving data generation unit is used to input the target control command vector into the second target model to obtain the target intelligent driving data corresponding to the target control command vector.

[0071] The intelligent driving data generation method provided in this application includes: acquiring first target image data within a target area, wherein the target area is the area surrounding the vehicle; performing semantic parsing processing on the first target image data to obtain a target scene semantic information set corresponding to the first target image data; determining a target structured template based on the target scene semantic information set; constructing target structured scene text based on the target structured template; generating target scene data based on the target structured scene text, and sending the target scene data to a cloud server so that the cloud server parses the target scene data to obtain a target control command vector; and having the cloud server perform data generation processing on the target control command vector to obtain target intelligent driving data. The intelligent driving data generation method provided in this application has the following beneficial effects: (1) The vehicle only generates and uploads KB-level text and MB-level snapshots, minimizing the consumption of computing, storage and communication resources, making it possible to deploy data collection at low cost for ultra-large fleets; (2) Through the scene semantic parsing and structured template triggering mechanism of the vehicle-side lightweight AI model, each data upload accurately corresponds to a high-value scene that has been semantically filtered. The backend can directly carry out subsequent processing based on the accurate semantic information uploaded, and the data processing efficiency is improved by orders of magnitude. (3) The cloud-based multimodal generation model can generate tens of thousands of scene variations based on a single semantic text prompt uploaded by the vehicle. It can not only expand a massive number of training samples from a real triggering event, but also accurately locate model defects and generate data in a targeted manner, providing comprehensive and diversified training support for model optimization. (4) The vehicle only uploads descriptive structured text and single-frame key snapshots, rather than continuous video streams, which avoids the privacy leakage risk caused by continuous recording, greatly simplifies the data compliance process, and reduces compliance risks and desensitization costs; (5) An automatic iterative closed loop has been formed. After the multimodal data generated in the cloud is used to train the intelligent driving model, the performance of the model in a specific scenario will serve as a feedback signal. Through reinforcement learning, the trigger threshold of the vehicle and the parameters of the cloud-generated model are dynamically optimized, so that the system becomes more and more accurate with use. The vehicle's ability to identify high-value scenarios continues to improve, and the quality and fit of the cloud-generated data are continuously optimized. This realizes an automated virtuous cycle of data production and algorithm iteration, accelerating the evolution of intelligent driving technology.

[0072] In an optional embodiment, this application provides an electronic device including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the intelligent driving data generation method as described in any of the above embodiments.

[0073] In an optional embodiment, this application provides a computer-readable storage medium storing at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the intelligent driving data generation method as described in any of the above embodiments.

[0074] The above description is only an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating intelligent driving data, characterized in that, The method is applied in a vehicle, and the method includes: Acquire first target image data within the target area, wherein the target area is the area surrounding the vehicle; The first target image data is subjected to semantic parsing processing to obtain a set of target scene semantic information corresponding to the first target image data; Based on the set of semantic information of the target scene, determine the target structured template; Based on the target structured template, construct the target structured scene text; Based on the target structured scene text, target scene data is generated and sent to a cloud server so that the cloud server can parse the target scene data to obtain a target control command vector; and the cloud server can perform data generation processing on the target control command vector to obtain target intelligent driving data.

2. The intelligent driving data generation method according to claim 1, characterized in that, The step of performing semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data includes: Obtain a first target model; the first target model is obtained by semantic parsing and training a preset model based on historical image data and scene semantic information labels corresponding to the historical image data. The first target image data is input into the first target model to obtain a set of target scene semantic information corresponding to the first target image data.

3. The intelligent driving data generation method according to claim 1, characterized in that, The target scene semantic information set includes several pieces of scene semantic information. The step of determining the target structured template based on the target scene semantic information set includes: Obtain preset mapping information, which represents the mapping relationship between preset conditions and structured templates, and the preset conditions include target preset conditions; If at least one piece of scene semantic information in the target scene semantic information set satisfies the target preset condition, a structured template corresponding to the target preset condition is determined based on the preset mapping information. Use the structured template corresponding to the target preset conditions as the target structured template.

4. The intelligent driving data generation method according to claim 3, characterized in that, The scene semantic information includes scene attributes and scene data corresponding to the scene attributes. The step of constructing the target structured scene text based on the target structured template includes: The target structured template is parsed to obtain several template scene attributes and the initial data corresponding to the several template scene attributes; The initial data corresponding to each template scene attribute is updated to the target scene data corresponding to the target scene attribute to obtain the target structured scene text; Wherein, the target scene attribute is the scene attribute in the target scene semantic information, and the target scene semantic information is any scene semantic information in the target scene semantic information set.

5. The intelligent driving data generation method according to claim 1, characterized in that, The step of generating target scene data based on the target structured scene text and sending the target scene data to the cloud server includes: Reacquire image data within the target area to obtain second target image data; Obtain controller LAN metadata; The target structured scene text, the second target image data, and the controller local area network metadata are used as the target scene data; The target scene data is sent to the cloud server.

6. A method for generating intelligent driving data, characterized in that, The method is applied to a cloud server, and the method includes: The system receives target scene data sent by a vehicle. The target scene data is generated by the vehicle based on target structured scene text. The target structured scene text is constructed by the vehicle according to a target structured template. The target structured template is determined by the vehicle based on a set of target scene semantic information corresponding to the first target image data. The set of target scene semantic information is obtained by the vehicle by parsing the first target image data within the target area. The target area is the area surrounding the vehicle. The target scene data is parsed to obtain the target control command vector; The target control command vector is processed to generate target intelligent driving data.

7. The intelligent driving data generation method according to claim 6, characterized in that, The process of generating target intelligent driving data by processing the target control command vector includes: Obtain a second target model; the second target model is obtained by generating and training a preset model based on historical control command vectors and the corresponding intelligent driving data of the historical control command vectors. The target control command vector is input into the second target model to obtain the target intelligent driving data corresponding to the target control command vector.

8. An intelligent driving data generation device, characterized in that, The intelligent driving data generation device includes: The image acquisition module is used to acquire first target image data within a target area, wherein the target area is the area surrounding the vehicle; The semantic parsing module is used to perform semantic parsing processing on the first target image data to obtain a set of target scene semantic information corresponding to the first target image data; The text generation module is used to determine a target structured template based on the target scene semantic information set; and to construct target structured scene text based on the target structured template. The data transmission module is used to generate target scene data based on the target structured scene text, and send the target scene data to a cloud server so that the cloud server receives the target scene data sent by the vehicle; and to enable the cloud server to parse the target scene data to obtain a target control command vector; and to enable the cloud server to perform data generation processing on the target control command vector to obtain target intelligent driving data.

9. An intelligent driving data generation device, characterized in that, The intelligent driving data generation device includes: The data receiving module is used to receive target scene data sent by the vehicle; the target scene data is generated by the vehicle based on target structured scene text, the target structured scene text is constructed by the vehicle according to the target structured template, the target structured template is determined by the vehicle according to the target scene semantic information set corresponding to the first target image data, the target scene semantic information set is obtained by the vehicle by parsing the first target image data in the target area, and the target area is the area around the vehicle; The data parsing module is used to parse the target scene data to obtain the target control command vector; The data generation module is used to perform data generation processing on the target control command vector to obtain target intelligent driving data.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or at least one program being loaded and executed by the processor to implement the intelligent driving data generation method as described in any one of claims 1-7.