Model training method, information processor and platform

Through the collaboration between the vehicle-end system and the cloud system, pre-trained models and loss functions are used to automatically label data, forming similar scene data sets and training deep neural network models, solving the problem of low efficiency in self-driving data acquisition and labeling, and achieving efficient and accurate model training.

CN120296423APending Publication Date: 2025-07-11MINGSHANG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510427181.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the acquisition and labeling of autonomous driving data is inefficient, and it relies on a large amount of manual resources and is difficult to achieve reliable model training.

Method used

Through the collaboration between the vehicle-end system and the cloud system, pre-trained models and loss functions are used to automatically label data, form similar scene data sets and train deep neural network models to provide accurate and reliable training sample data.

Benefits of technology

It realizes efficient and accurate data acquisition and automatic labeling without manual annotation, and improves the training efficiency and reliability of the autonomous driving model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296423A_ABST
    Figure CN120296423A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, an information processor and a platform. Relates to the technical field of automatic driving, and adopts the specific scheme that corresponding labeling information and corresponding decision control information errors in a similar scene data set are respectively calculated by utilizing a loss function, the decision control information with the minimum error is selected, corresponding original data in the similar scene data set is labeled, and a training data set is obtained. And the model reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of autonomous driving technology and artificial intelligence technology, and particularly to data collection and automatic data annotation, aiming to solve the problems of low efficiency in processing sample data and difficulty in reliably training an autonomous driving model in autonomous driving technology. Specifically, the present invention relates to a model training method, an information processor, and a device. Background Art

[0002] In "Automated Driving Classification for Motor Vehicles" (GB / T 40429-2021), driving automation is divided into levels 0 to 5, each level having corresponding technical requirements. As a necessary condition for the application and implementation of autonomous driving technology, data collection and data annotation have become an essential part of supporting autonomous driving technology.

[0003] Among them, data collection mainly includes two methods: manual and simulation. The manual method is to collect raw data such as space, vision, object shape, speed, etc. of the driving environment through a large number of diverse sensors installed on the collection vehicle, and then a large number of computer engineers organize, classify, cut, outline, describe, etc. these raw data to achieve data annotation, thereby obtaining the process of basic sample data. The simulation method is a process of automatically generating an environment perception data set with annotations based on a certain model algorithm or through the operation of professional engineers.

[0004] The manual method requires a large amount of human resources and time costs, and it is difficult to guarantee the accuracy of data annotation; the simulation method largely depends on the technical strength of the simulation software provider and its own data accumulation, and also requires a large amount of human resources and time costs. Moreover, it is difficult for the simulation environment to cover or reflect the latest and timely actual physical operating environment, so it is difficult to achieve reliable training of the autonomous driving algorithm model.

[0005] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of the present invention is to collect and automatically annotate a large amount of driving data, provide accurate, reliable, and end-to-end training sample data for the autonomous driving algorithm model, and eliminate the high cost and low efficiency of relying on a large number of manual annotations.

[0007] In a first aspect disclosed by the present invention, there is provided a model training method, which includes:

[0008] Obtain vehicle state information and vehicle operating environment information and perform alignment processing to obtain raw data;

[0009] Based on a pre-trained model, annotate the raw data to obtain corresponding annotation information;

[0010] Obtain the driving actions of the vehicle at the corresponding moment to obtain the corresponding decision control information;

[0011] Classify them into a similar scenario dataset based on the original data, the corresponding annotation information, and the corresponding decision control information;

[0012] Use the loss function to calculate the errors between the corresponding annotation information and the corresponding decision control information in the similar scenario dataset respectively, select the decision control information with the smallest error, and label the corresponding original data in the similar scenario dataset to obtain a training dataset;

[0013] Train the pre-trained model based on the training dataset.

[0014] According to the second aspect of the present disclosure, an information processor is provided. The information processor is configured to execute electronic instructions stored in a storage device, and when the electronic instructions are executed, the information processor executes the method of the first aspect.

[0015] Adopting the solution of the present disclosure can realize data collection and automatic data annotation, provide training data samples for the artificial intelligence algorithm of the autonomous driving system for unmanned driving, and provide a precise and efficient method for collecting, annotating, and model training of vehicle driving behavior data for the research and development of unmanned driving technology, and improve the reliability of the model. Brief Description of the Drawings

[0016] Figure 1 It is a schematic structural diagram of the vehicle-end system according to an embodiment of the present invention.

[0017] Figure 2 It is a flowchart of the dataset according to an embodiment of the present invention.

[0018] Figure 3 It is a schematic structural diagram of the cloud system according to an embodiment of the present invention.

[0019] Figure 4 It is a flowchart of the model training method according to an embodiment of the present invention.

[0020] Figure 5 It is a schematic diagram of the reinforcement learning model according to an embodiment of the present invention.

[0021] Figure 6 It is a flowchart of the reinforcement learning method according to an embodiment of the present invention.

[0022] Reference Numerals and Names in the Drawings:

[0023] 10. Vehicle-end system; 101. Data collection module; 102. Pre-trained model; 103. Decision perception module;

[0024] 20. Fleet; 30. Similar scenario dataset; 40. Similar dataset; 50. Training dataset;

[0025] 60. Cloud system; 601. Central database; 602. Automatic annotation module; 603. Pre-training system;

[0026] 701. Real-time environmental data; 702. Decision-aware label; 703. Reward function; 704. Deep neural network.

[0027] The realization, functional features and advantages of the purpose of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0028] To make the purposes, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0029] The present invention provides a vehicle terminal system, as Figure 1 shown. The vehicle terminal system 10 includes a data acquisition module 101, a pre-training model 102, and a decision-aware module 103.

[0030] The vehicle terminal system 10 refers to a system installed on a vehicle, mainly an automobile, and it should be understood that it also includes other types of land vehicles, which are equipped with terminal devices such as sensors, control units, and communication modules, and can obtain vehicle status information and vehicle operation environment information in real time, and perceive and make judgments on the information through artificial intelligence algorithms and control the vehicle to execute or not execute actions.

[0031] The vehicle terminal system 10 can be a computer system installed on a vehicle or a unit with computer processing functions, which is responsible for transmitting the data collected by the sensors to the control unit, and through data processing, aligning the data in time, converting the format and saving the data to obtain the original data at a certain moment, and annotating the original data, and performing inference operations and decision-making perceptions by model algorithms.

[0032] Multiple vehicles installed with the vehicle terminal system 10 form a fleet 20, and information interaction between vehicles and information interaction between vehicles and the cloud or traffic infrastructure can be realized through the communication module of the vehicle terminal system 10. Among them, the communication module at least includes a V2X communication module, an Internet communication module, a 4G / 5G module, a WiFi module, etc.

[0033] The data acquisition module 101 obtains vehicle status information and vehicle operating environment information, including space, vision, object shape, speed, etc. through sensors, at least including cameras, high-definition cameras, lidar, ultrasonic radars, and infrared sensors. For example, point cloud data is obtained through lidar, image data is obtained through cameras, and spatial distance data is obtained through ultrasonic radars.

[0034] The data acquisition module 101 collects data (including point cloud data, image data, and distance data) in chronological order, aligns the data in time, performs format conversion and data storage to obtain the original data corresponding to a certain moment.

[0035] The pre-trained model 102 fuses the obtained information, performs sorting and semantic understanding in time and space. For example, at a certain spatial position (relative to the vehicle) at a certain moment, there is an object or pedestrian with a certain motion state and shape.

[0036] The pre-trained model 102 is a deep neural network model, including a recurrent neural network model, a convolutional neural network model, a multi-layer neuron model, and a deep neural network model based on a Transformer. It receives the original data collected by the data acquisition module 101 and performs inference operations on the original data through the pre-trained model 102 to obtain the corresponding annotation information.

[0037] The decision-making and perception module 103 is connected to the control interface of the vehicle, obtains the driving actions at the current moment, makes real-time judgments and issues decision control information for the vehicle actions based on the current driving actions and the understanding of the perception data (annotation information), and converts the decision control information into actual control of the vehicle's physical components, such as braking, steering, accelerating, decelerating, etc.; it can also obtain the driving actions of the vehicle at the current moment as decision control information and form a corresponding relationship between the annotation information and the decision control information.

[0038] The vehicle-end system 10 aligns the collected original data, annotation information, and decision control information in time and packs or batches them into data packets for storage.

[0039] It can be known that for the autonomous driving system, the pre-trained model 102 of the vehicle-end system in this application is crucial for understanding the perception data and making correct decisions. The pre-trained model 102 requires a large amount of sample data for training. The sample data refers to the original data information of the vehicle operating environment obtained through sensors and is annotated to give explanations of features, semantics, or final decisions, which is a data set containing the original data and the corresponding annotation information.

[0040] In some embodiments, such as Figure 2As shown in the figure, the data set of the original data, annotation information, and decision control information of the vehicle end system 10 on a certain vehicle or different vehicles in the vehicle fleet 20 is classified as a similar scenario data set 30

[0041]

[0042] Among them, g{} is the similar scenario data set; is a data set containing the original data D, annotation information PT, and decision control information T; D(n) is the original data at the nth moment; PT(n) is the corresponding annotation information of the original data at the nth moment; T(n) is the decision control information at the nth moment; n = 1, 2, 3, ….

[0043] The classification methods include the same scenario, such as intersections and road conditions; and similar scenarios, such as classification according to similar intersections and similar road conditions; similar actions, such as actions like overtaking, lane changing, and sudden braking; for the original data (D(1), D(2), …, D(n)) in the similar scenario data set 30 obtained by the data acquisition module 101, the annotation information PT obtained through inference by the pre-trained model 102 is equal or the same, that is, the original data corresponding to the same annotation information PT can be classified as the similar data set 40. Therefore, Equation (1) is simplified to:

[0044]

[0045] Among them, g{} is the similar scenario data set; is a data set containing the original data D, annotation information PT, and decision control information T; D(n) is the original data at the nth moment; PT is the annotation information corresponding to the similar data set; T(n) is the decision control information at the nth moment; n = 1, 2, 3, ….

[0046] Use the loss function to calculate the error between the annotation information PT and the corresponding decision control information T of the similar data set 40 respectively, and select the decision control information with the smallest error as a label T * .

[0047]

[0048] Among them, T * is the decision control information with the smallest error between the annotation information and the corresponding decision control information in the similar scenario data set; represents the independent variable that obtains the minimum value within the range of T; δ(PT, T(n)) is the loss function of the error between the annotation information PT and the corresponding decision control information T in the similar data set; n = 1, 2, 3, ….

[0049] The loss function of the error between the annotation information PT and the corresponding decision control information T in the above includes mean square error, cross entropy, absolute error, etc. for different data types.

[0050] Relabel the label T * with the original data of the similar data set to form a new training data set 50.

[0051] Γ(D, T * ) → Γ(D(n), T * ) (4)

[0052] where Γ(D(n), T * ) is the training data set; D(n) is the original data at the nth moment in the similar data set.

[0053] Train the pre-trained model with the formed training data set 50 to ensure the stability of training and improve the inference performance and accuracy of the model.

[0054] The present invention provides a cloud system for model training, as Figure 3 shown, the cloud system 60 includes a central database 601, an automatic annotation module 602 and a pre-training system 603.

[0055] The cloud system 60 can be a cloud server or a general server, where the computer-readable instructions for model training are stored in the memory.

[0056] Central database 601: Used to store the similar scenario data set 30 and the iteratively updated data information. The similar scenario data set 30 is a data set of the original data, annotation information and decision control information from the vehicle-end systems 10 on different vehicles in the vehicle fleet 20.

[0057] The similar scenario data set is:

[0058]

[0059] where g{} is the similar scenario data set; is a data set containing the original data D, the annotation information PT and the decision control information T; D(n) is the original data at the nth moment; PT(n) is the corresponding annotation information of the original data at the nth moment; T(n) is the decision control information at the nth moment; n = 1, 2, 3,....

[0060] The similar scenario data set 30 is a data set obtained by acquiring the original data, the corresponding annotation information and the corresponding decision control information based on the same scenario, similar scenarios and / or similar actions.

[0061] Automatic annotation module 602: Calculate the errors between the annotation information and the corresponding decision and control information in the similar scenario dataset 40 respectively using a loss function, select the decision and control information with the smallest error, and annotate the original data in the similar scenario dataset to obtain a training dataset 50; Based on the training dataset 50, annotate the mapping relationship between the original data (input) and the decision and control information (model inference result).

[0062] The training dataset is as follows:

[0063] Γ(D, T * ) → Γ(D(n), T * ) (4)

[0064] where Γ(D(n), T * ) is the training dataset; D(n) is the original data at the nth moment in the similar data set; T * is the decision and control information with the smallest error between the annotation information and the corresponding decision and control information in the similar scenario dataset; n = 1, 2, 3, ….

[0065] The loss function includes mean squared error, cross entropy, absolute error, etc. for different data types.

[0066] Pre-training system 603: Train the model in the pre-training system 603 based on the annotated training dataset to obtain a pre-trained model. The model can be a deep neural network model, including a recurrent neural network model, a convolutional neural network model, a multi-layer neuron model, and a deep neural network model based on a Transformer.

[0067] Deploy the trained pre-trained model to the vehicle terminal system 10 to form a continuously iterative vehicle networking platform, and realize the management of existing vehicles and vehicle fleets, as well as the collection and automatic annotation of a large amount of driving data, providing accurate, reliable, and end-to-end training sample data for the autonomous driving model, training the model reliably, without manual annotation, and improving efficiency, accuracy, and reliability.

[0068] The present invention provides a model training method, as Figure 4 shown,[[]]END]]

[0069] S101. Obtain vehicle state information and vehicle operating environment information and perform alignment processing to obtain original data.

[0070] Obtaining vehicle state information and vehicle operating environment information includes information such as space, vision, object form, speed, etc.; for example, obtaining point cloud data through a lidar, obtaining image data through a camera, and obtaining spatial distance data through an ultrasonic radar.

[0071] Alignment processing refers to aligning the acquired information in time, that is, a certain moment corresponds to vehicle state information and vehicle operating environment information. Format conversion and storage are required for different formats of information.

[0072] S102. Label the original data based on the pre-trained model to obtain the corresponding labeled information.

[0073] The pre-trained model is a deep neural network model, including a recurrent neural network model, a convolutional neural network model, a multi-layer neuron model, and a deep neural network model based on a Transformer. It receives the original data and performs inference operations on the original data through the pre-trained model to obtain the corresponding labeled information.

[0074] S103. Obtain the driving actions of the vehicle at the corresponding moment to obtain the corresponding decision control information.

[0075] The driving actions of the vehicle at a certain moment correspond to the vehicle operating environment information at that moment and the labeled information of the original data at that moment.

[0076] S104. Classify the original data, the corresponding labeled information, and the corresponding decision control information into a similar scenario data set.

[0077] The similar scenario data set is:

[0078]

[0079] where g{} is the similar scenario data set; is a data set containing the original data D, the labeled information PT, and the decision control information T; D(n) is the original data at the nth moment; PT(n) is the corresponding labeled information of the original data at the nth moment; T(n) is the corresponding decision control information at the nth moment; n = 1, 2, 3,....

[0080] The similar scenario data set is a data set that obtains the original data, the corresponding labeled information, and the corresponding decision control information based on the same scenario, similar scenarios, and / or similar actions.

[0081] S105. Use the loss function to calculate the errors between the corresponding labeled information and the corresponding decision control information in the similar scenario data set respectively, select the decision control information with the smallest error, and label the corresponding original data in the similar scenario data set to obtain the training data set.

[0082]

[0083] where T * is the decision control information with the smallest error between the labeled information and the corresponding decision control information in the similar scenario data set; The independent variable that obtains the minimum value within the range of T; δ(PT, T(n)) is the loss function of the error between the labeled information PT in the similar dataset and the corresponding decision control information T; n = 1, 2, 3, ….

[0084] The loss function includes mean squared error, cross entropy, absolute error, etc. for different data types.

[0085] The training dataset is:

[0086] Γ(D, T * ) → Γ(D(n), T * ) (4)

[0087] Where Γ(D(n), T * ) is the training dataset; D(n) is the original data at the nth moment in the similar dataset; T * is the decision control information with the minimum error between the labeled information and the corresponding decision control information in the similar scenario dataset; n = 1, 2, 3, ….

[0088] S106. Train the pre-trained model based on the training dataset.

[0089] The pre-trained model can be a deep neural network model, including a recurrent neural network model, a convolutional neural network model, a multi-layer neuron model, and a deep neural network model based on Transformer.

[0090] Deploy the trained pre-trained model into the original model to form a continuously iterative vehicle networking platform, and realize the management of existing vehicles and fleets, as well as the collection and automatic annotation of a large amount of driving data, providing accurate, reliable, and end-to-end training sample data for the autonomous driving model, reliably training the model, without manual annotation, improving efficiency, accuracy, and reliability.

[0091] This method is a training method based on continuous data iteration of fleet collaboration. This method avoids time-consuming and laborious processing work such as manual data sorting, classification, cutting, outlining, and description, collects actual data to reflect the real physical operation environment, and ensures the reliability of model training.

[0092] Reinforcement learning is an effective method for training deep neural network models. This application proposes a reinforcement learning method, as Figure 5 and Figure 6 shown:

[0093] S201. Obtain real-time environmental data and corresponding decision perception labels to form a training dataset to train the deep neural network model.

[0094] Real-time environmental data 701: It includes obtaining vehicle running environment information through sensors, at least including cameras, high-definition cameras, lidars, ultrasonic radars, and infrared sensors, including information such as space, vision, object form, speed, etc.; for example, obtaining point cloud data through lidars, obtaining image data through cameras, and obtaining spatial distance data through ultrasonic radars.

[0095] Corresponding decision perception label 702: Based on the real-time environmental data, the corresponding actions made by the vehicle, and this information obtains the corresponding driving actions through the vehicle connection interface, and annotates this information to obtain the corresponding decision perception label.

[0096] Use the real-time environmental data as the input and the corresponding decision perception label as the output to form a training data set and train the deep neural network model to obtain a pre-trained model.

[0097] S202. Use the reward function to calculate the reward values of the real-time environmental data and the corresponding decision perception labels in the training data set, select the decision perception label with the largest reward value and annotate the real-time environmental data in the training data set to obtain a new training data set, and based on the new training data set, train the training model.

[0098] The reward function 703 refers to the incentive for the optimal action of the decision perception label under the real-time environmental data.

[0099] The reward value can reflect the safety of the vehicle during driving (including but not limited to the distance from the vehicle in front, whether it deviates from the lane, prediction and avoidance of surrounding obstacles, etc.), comfort (including but not limited to vehicle speed, acceleration change, driving jerks, steering angular acceleration, etc.), and economy (including but not limited to fuel / electricity loss, wear of vehicle components, etc.).

[0100] For example, in an emergency, the moving distance a of a driving vehicle from starting to brake until it stops, and the distance between the vehicle and the obstacle is b. To avoid the vehicle from colliding, it is necessary to satisfy a < b. Braking in an emergency will bring discomfort and even potential harm. How to brake within a limited braking distance, such as intermittent braking and controlling the length of the interval time, forms different strategies, and each strategy is scored and rewarded. The higher the score, the better the comfort.

[0101] For a detailed introduction to the reward function, please refer to the relevant description in Patent CN 115107767A.

[0102] The decision-aware label with the maximum reward value is the optimal decision among the decisions made by different drivers in similar scenarios at different times for different vehicles in the fleet. Therefore, this decision-aware label is the best decision obtained by comparing the vehicle driving in time and space. Moreover, through continuous iteration of the reinforcement learning model, it is a dynamic optimization process and becomes the consensus mechanism or the autonomous awareness of the fleet.

[0103] The deep neural network 704 includes a recurrent neural network model, a convolutional neural network model, a multi-layer neuron model, and a deep neural network model based on a Transformer.

[0104] An information processor configured to execute electronic instructions stored in a storage device, which, when executed, cause the information processor to execute the above model training method and / or reinforcement learning method.

[0105] It should be noted that the vehicle-side system 10 and the cloud-side system 60 cooperate with each other to execute the model training method.

[0106] In the above embodiments, it can be implemented in computer hardware, software, or a combination thereof.

[0107] For example, to build an Internet of Vehicles platform for fleet management, a vehicle-side system 10 installed on several vehicles and a cloud-side system 60 connected to the vehicle-side system through the Internet. All vehicles connected to the same cloud-side system 60 through the vehicle-side system 10 form a fleet 20. The vehicle-side system 10 is a set of computer programs running on multiple vehicles. The vehicle-side system 10 is connected to the cloud-side system 60 through the Internet to upload the data, status perception, and decision signals collected by the vehicle-side. The cloud-side system 60 mainly includes a central database 601, an automatic data annotation module 602, and a pre-training system 603. The central database is used to store the data sent from the vehicle-side system and the data iteratively updated in the cloud-side system; the automatic annotation module is responsible for automatically tagging the original data. The pre-training system is responsible for training a deep neural network based on the labeled data samples to obtain a pre-trained model and upgrading and deploying it to the vehicle-side system.

Claims

1. A model training method, characterized in that, Including: Obtain vehicle status information and vehicle operating environment information and perform alignment processing to obtain original data; Based on a pre-trained model, label the original data to obtain corresponding annotation information; Obtain the driving actions of the vehicle at the corresponding moment to obtain corresponding decision control information; Classify the original data, the corresponding annotation information, and the corresponding decision control information into a similar scenario dataset; Use a loss function to calculate the errors between the corresponding annotation information and the corresponding decision control information in the similar scenario dataset respectively, select the decision control information with the smallest error, and label the corresponding original data in the similar scenario dataset to obtain a training dataset; Train the pre-trained model based on the training dataset.

2. The method according to claim 1, wherein The alignment processing is to align the obtained information in time.

3. The method according to claim 1, wherein The pre-trained model is a deep neural network model.

4. The method according to claim 1, wherein The similar scenario dataset is a data set that obtains original data, corresponding annotation information, and corresponding decision control information based on the same scenario.

5. The method according to claim 1, wherein The similar scenario dataset is a data set that obtains original data, corresponding annotation information, and corresponding decision control information based on similar scenarios.

6. The method according to claim 5, wherein Similar scenarios include similar intersections and similar road conditions.

7. The method according to claim 1, characterized in that The similar scenario dataset is a data set that obtains original data, corresponding annotation information, and corresponding decision control information based on similar actions.

8. The method according to claim 7, wherein Similar actions include overtaking, lane changing, and emergency braking.

9. The method according to claim 1, characterized in that, The loss function is selected from mean square error, cross entropy, and absolute error.

10. An information processor, characterized in that, The information processor is configured to execute electronic instructions stored in a storage device, and when the electronic instructions are executed, the information processor executes the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Automatic driving braking and anti-collision control method based on artificial intelligence

    CN115107767A