Model training method, driving right distribution method, equipment and storage medium

By obtaining flight sample data for driving action prediction and driver evaluation, training the driving rights allocation model, solving the safety hazards caused by manual allocation of drivers, realizing the accurate allocation of autonomous driving rights, and improving the safety of flying cars.

CN120348304AActive Publication Date: 2025-07-22CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510492843.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the allocation of driving rights between the intelligent driving system of a flying car and the driver mainly depends on the manual allocation of the driver, which is easily affected by the driver's mental state and technical level, and poses safety hazards.

Method used

By obtaining flight sample data, driving action prediction and driver evaluation are carried out, driving rights allocation model is obtained using pre-trained model training, driving rights are automatically allocated, and driver status and environmental risks are considered.

Benefits of technology

The flight safety of flying cars is improved, and the driving rights are accurately allocated by identifying the driver's current driving ability, and the impact of human factors is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120348304A_ABST
    Figure CN120348304A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving, and discloses a model training method, a driving right distribution method, equipment and a storage medium. The model training method comprises the following steps: acquiring flight sample data; performing driving action prediction according to the flight sample data to obtain action prediction sample data; evaluating the driver according to the driver state sample data, the driver action sample data and the action prediction sample data to obtain driver evaluation sample data; inputting the flight sample data and the driver evaluation sample data into a preset to-be-trained model to obtain driving right distribution sample data; and training a to-be-trained model according to the driving right distribution sample data to obtain a driving right distribution model. According to the embodiment of the invention, the driving right can be automatically allocated, so that the flight safety of the hovercar is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular, to a model training method, a driving right allocation method, a device, and a storage medium. Background Art

[0002] Flying cars, as an emerging means of transportation, have received increasing attention in recent years.

[0003] For flying cars, there is a process of sharing the human-machine driving right in the intelligent driving function at the current development stage. Due to the sharing of the human-machine driving right, there must be a coupling and restrictive relationship between the actions of the intelligent driving system and the actions of the driver, and there will be a struggle for the driving right, which has a game nature.

[0004] In related technologies, the allocation of the driving right between the intelligent driving system and the driver mainly depends on the driver's manual allocation. During the process of the driver's manual allocation of the driving right, it is easily affected by the driver's mental state and driving skills, and thus there are potential safety hazards. Summary of the Invention

[0005] The purpose of this application is to provide a model training method, a driving right allocation method, a device, and a storage medium, which solve the technical problem of the need for the driver to manually allocate the driving right and thus there are potential safety hazards, and realize the automatic allocation of the driving right to improve the flight safety of the flying car.

[0006] An embodiment of this application provides a model training method, including: Obtain flight sample data; the flight sample data includes flight environment sample data, driver state sample data, driver action sample data, and vehicle condition sample data; Perform driving action prediction according to the flight sample data to obtain action prediction sample data; Evaluate the driver according to the driver state sample data, the driver action sample data, and the action prediction sample data to obtain driver evaluation sample data; Input the flight sample data and the driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data; Train the model to be trained according to the driving right allocation sample data to obtain a driving right allocation model.

[0007] In some embodiments, the performing driving action prediction according to the flight sample data to obtain action prediction sample data includes: Input the flight sample data into a pre-trained driving action prediction model to obtain the action prediction sample data; the driving action prediction model is trained based on the action prediction loss information of the action prediction sample data, and the action prediction loss information is obtained by fitting the first policy loss information, the first value loss information, and the first entropy loss information. The first policy loss information represents the quality of the policy for generating the action prediction sample data, the first value loss information represents the value deviation between the action prediction sample data and the real action data, and the first entropy loss information represents the uncertainty of the action prediction sample data.

[0008] In some embodiments, the evaluating the driver based on the driver state sample data, the driver action sample data, and the action prediction sample data to obtain driver evaluation sample data includes: Generate first evaluation sample data representing the driving state according to the driver state sample data; Generate second evaluation sample data representing driving skills according to the action prediction sample data and the driver action sample data; Evaluate the driver based on the first evaluation sample data and the second evaluation sample data to obtain the driver evaluation sample data.

[0009] In some embodiments, the inputting the flight sample data and the driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data includes: Predict the driving scenario at a future moment according to the flight sample data and the driver evaluation sample data to obtain multiple future scenario feature sample data; Calculate the scenario reward data corresponding to each of the future scenario feature sample data; Output the corresponding driving right allocation sample data according to the scenario reward data.

[0010] In some embodiments, the predicting the driving scenario at a future moment according to the flight sample data and the driver evaluation sample data to obtain multiple future scenario feature sample data includes: Perform feature extraction on the flight sample data and the driver evaluation sample data through continuous convolution to obtain current scenario feature sample data; Predict the driving scenario at a future moment according to the transformation probability of the current scenario feature sample data at the future moment to obtain multiple future scenario feature sample data.

[0011] In some embodiments, the scenario reward data is obtained by fitting first reward sample data, second reward sample data, third reward sample data, fourth reward sample data, and fifth reward sample data. The first reward sample data represents the flight efficiency of the driving scenario corresponding to the future scenario feature sample data. The second reward sample data represents the safety level of the driving scenario corresponding to the future scenario feature sample data. The third reward sample data represents the smoothness of the driving scenario corresponding to the future scenario feature sample data. The fourth reward sample data represents the success rate of human-machine collaboration in the driving scenario corresponding to the future scenario feature sample data. The fifth reward sample data represents the quality of the strategy for generating the future scenario feature sample data.

[0012] In some embodiments, training the to-be-trained model according to the driving right allocation sample data to obtain a driving right allocation model includes: Determining model loss information corresponding to the driving right allocation sample data; the model loss information is obtained by fitting second policy loss information, second value loss information, and second entropy loss information. The second policy loss information represents the quality of the strategy for generating the driving right allocation sample data. The second value loss information represents the value deviation between the driving right allocation sample data and the true driving right allocation data. The second entropy loss information represents the uncertainty of the driving right allocation sample data; Iteratively adjusting the network parameters in the to-be-trained model according to the action prediction loss information until the training end condition is met, to obtain the driving right allocation model.

[0013] An embodiment of the present application further provides a driving right allocation method, including: Obtaining flight data; the flight data includes flight environment data, driver state data, driver action data, and vehicle condition data; Performing driving action prediction according to the flight data to obtain action prediction data; Evaluating the driver according to the driver state data, the driver action data, and the action prediction data to obtain driver evaluation data; Inputting the flight data and the driver evaluation data into a driving right allocation model to obtain driving right allocation data; the driving right allocation model is trained by the above model training method; Performing a driving right allocation action according to the driving right allocation data.

[0014] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented.

[0015] The embodiments of the present application further provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0016] Advantages of the present application: Predict driving actions based on flight sample data to obtain action prediction sample data, and then evaluate the driver based on driver state sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data as the basis for evaluating the driving ability of the driver. Input the flight sample data and driver evaluation sample data into a preset model to be trained, use the driving right allocation sample data obtained by the model to be trained to train the model to be trained, and obtain a driving right allocation model after the training is completed. Since the driver evaluation sample data is obtained by evaluating the driver based on driver state sample data, driver action sample data, and action prediction sample data, inputting the flight sample data and driver evaluation sample data into the model to be trained and using the driving right allocation sample data to train the model to be trained can enable the model to be trained to learn the degree of dependence of different driver states and driver operation actions on the driving right allocation result during the training process, so that the trained driving right allocation model can accurately identify the current driving ability of the driver and determine the driving right allocation data. When automatically allocating the driving right according to the driving right allocation data, the flight safety of the flying car can be improved. Description of the Drawings

[0017] Figure 1 It is an application environment diagram of the model training method provided by the embodiments of the present application.

[0018] Figure 2 It is a flowchart of the model training method provided by the embodiments of the present application.

[0019] Figure 3 It is a flowchart of the driving right allocation method provided by the embodiments of the present application.

[0020] Figure 4 It is a hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments

[0021] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] It should be noted that although functional modules are divided during the demonstration of the device's actions and the logical order is shown in the flowchart, in some cases, the steps shown can be executed in a different order from the module division in the device or the flowchart. Terms such as "first" and "second" in the specification, claims, and drawings are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0024] The model training method and driving right allocation method provided by the embodiments of this application can be executed by a computer device, which can be a terminal device or a server. Among them, the terminal device includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, a distributed system, or a cloud server.

[0025] In addition, the information, data, and signals involved in the embodiments of this application are all authorized by the relevant objects or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0026] To facilitate understanding of the model training method provided by the embodiments of this application, the application scenario of this model training method will be exemplarily introduced below by taking the execution entity of this model training method as a server.

[0027] Figure 1 This is the application environment diagram of the model training method provided by the embodiments of this application. Refer to Figure 1, the model training method is applied to a model training system. The model training system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to send flight sample data including flight environment sample data, driver status sample data, driver action sample data, and vehicle condition sample data to the server 120. The server 120 is used to obtain the flight sample data, perform driving action prediction based on the flight sample data to obtain action prediction sample data, evaluate the driver based on the driver status sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data, input the flight sample data and driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data, and train the model to be trained based on the driving right allocation sample data to obtain a driving right allocation model.

[0028] It should be understood that Figure 1 The application scenarios shown are only examples. In actual applications, the model training method provided in the embodiments of the present application can also be applied to other scenarios. For example, the above model training method can be directly applied to the terminal 110. The terminal 110 is used to obtain flight sample data including flight environment sample data, driver status sample data, driver action sample data, and vehicle condition sample data, perform driving action prediction based on the flight sample data to obtain action prediction sample data, evaluate the driver based on the driver status sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data, input the flight sample data and driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data, and train the model to be trained based on the driving right allocation sample data to obtain a driving right allocation model.

[0029] Figure 2 is a flowchart of a model training method provided in an embodiment of the present application. Refer to Figure 2 , in one embodiment, the method includes but is not limited to steps S201 to S205.

[0030] Step S201, obtain flight sample data.

[0031] The flight sample data can be obtained by retrieving the historical flight data of the flying car.

[0032] Among them, the flight sample data includes flight environment sample data, driver state sample data, driver action sample data, and vehicle condition sample data. The flight environment sample data can include environmental sample data such as obstacle positions, dynamic states, climate environments, no-fly zone positions, and air traffic densities. The driver state sample data can include physiological sample data such as the saccade frequency and heart rate of the driver. The driver action sample data can include action sample data such as speed input, altitude input, and direction input of the driver's action input. The vehicle condition sample data can include sample data such as the flight speed, flight altitude, pitch angle, heading angle, roll angle, and angular velocity of the flying car.

[0033] Step S202: Perform driving action prediction based on the flight sample data to obtain action prediction sample data.

[0034] Among them, the action prediction sample data refers to the driver action sample data obtained by predicting the driving actions performed by the driver at a later time based on the flight sample data at a previous time.

[0035] In one embodiment, performing driving action prediction based on the flight sample data can be to use a pre-trained driving action prediction model to predict the driving actions performed by the driver at a later time based on the flight sample data at a previous time, and then obtain the corresponding action prediction sample data.

[0036] Step S203: Evaluate the driver based on the driver state sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data.

[0037] Among them, the driver evaluation sample data refers to the evaluation sample data for predicting the driving ability of the driver at the prediction time based on the driver state sample data, driver action sample data, and the corresponding action prediction sample data.

[0038] In one embodiment, evaluating the driver based on the driver state sample data, driver action sample data, and action prediction sample data can be to use a preset driver evaluation model to predict the driving ability of the driver at the current time based on the driver state sample data, driver action sample data, and action prediction sample data at the same time, and evaluate the driving ability of the driver at the current time to obtain driver evaluation sample data.

[0039] Step S204: Input the flight sample data and driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data.

[0040] Among them, the driving right allocation sample data refers to the sample data used to indicate the allocation ratio of the driving right to the driver and the autonomous driving system in the flying car.

[0041] In one embodiment, after inputting flight sample data and driver evaluation sample data into a preset model to be trained, the model to be trained determines the driver's current driving ability and the risks of the current flight environment based on the flight sample data and the driver evaluation sample data, predicts whether the driver's current driving ability can cope with the risks of the current flight environment, and obtains corresponding driving right allocation sample data according to the prediction result. If the driver can cope, driving right allocation sample data for allocating the driving right to the driver is obtained; if the driver cannot cope, driving right allocation sample data for allocating the driving right to the autopilot system is obtained. The driving right allocation sample data may include a driving response plan predicted based on the flight sample data.

[0042] Step S205: Train the model to be trained according to the driving right allocation sample data to obtain a driving right allocation model.

[0043] In one embodiment, the flight sample data and the driver evaluation sample data are input into a preset model to be trained, and the network parameters of the model to be trained are adjusted each time the driving right allocation sample data is obtained, and finally a driving right allocation model is obtained. In a specific implementation, multiple groups of flight sample data are traversed, action prediction sample data is generated based on the flight sample data, driver evaluation sample data is generated based on the driver status sample data, the driver action sample data, and the action prediction sample data, and the model to be trained is used to generate the driving right allocation sample data based on the flight sample data and the driver evaluation sample data to perform iterative training on the model to be trained until the model loss information meets the training end condition. After the training is completed, a driving right allocation model is obtained.

[0044] In summary, the model training method provided by the embodiments of the present application predicts driving actions based on flight sample data to obtain action prediction sample data, and then evaluates the driver based on the driver state sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data as the basis for evaluating the driving ability of the driver. The flight sample data and the driver evaluation sample data are input into a preset model to be trained, and the model to be trained is trained using the driving right allocation sample data obtained by the model to be trained. After the training is completed, a driving right allocation model is obtained. Since the driver evaluation sample data is obtained by evaluating the driver based on the driver state sample data, driver action sample data, and action prediction sample data, inputting the flight sample data and the driver evaluation sample data into the model to be trained and training the model to be trained using the driving right allocation sample data can enable the model to be trained to learn the degree of dependence of different driver states and driver operation actions on the driving right allocation result during the training process, so that the trained driving right allocation model can accurately identify the current driving ability of the driver and determine the driving right allocation data. When automatically allocating the driving right according to the driving right allocation data, the flight safety of the flying car can be improved.

[0045] In one embodiment, step S202 above includes: inputting the flight sample data into a pre-trained driving action prediction model to obtain action prediction sample data.

[0046] Among them, the driving action prediction model is trained according to the action prediction loss information of the action prediction sample data. The action prediction loss information is obtained by fitting the first policy loss information, the first value loss information, and the first entropy loss information. The first policy loss information represents the quality of the policy for generating the action prediction sample data, the first value loss information represents the value deviation between the action prediction sample data and the real action data, and the first entropy loss information represents the uncertainty of the action prediction sample data.

[0047] Specifically, the driving action prediction model is a model with a convolutional neural network and a long short-term memory network. By continuously convolving the action prediction sample data to extract the local spatial features in the action prediction sample data, and then inputting the local spatial features and the temporal features into the long short-term memory network together to capture the temporal dependence relationship between the local spatial features and the temporal features, and further predicting the operation actions of the driver to obtain the action prediction sample data.

[0048] The driving action prediction model is trained according to the action prediction loss information of the action prediction sample data. The calculation formula of the action prediction loss information is: , Among them, is the first policy loss information, is the first value loss information, is the first entropy loss information, and are hyperparameters.

[0049] The calculation formula for the first policy loss information is: , The calculation formula for the first value loss information is: ,

[0050] , , = , , , , The calculation formula for the first entropy loss information is: , , Where: / is the ratio between the old and new policies, represents the state, represents the action, is the advantage function, is the clipping threshold, is the estimated value predicted by the action, is the response reward return value, is the policy taking the driving action under the state value, is the driver action prediction reward function, is the driver action prediction accuracy reward information, is the flying car safety reward information, is the driver comfort reward information, 、 and are weight coefficients, is the collision risk reward information, is the stability reward information, is the number of sampling points within the time window T, is the action deviation, is the average action deviation within the time window is the minimum distance between the flying car and surrounding obstacles, is the safety distance, is the attitude angle deviation of the flying car, is the stability threshold, 、 are the weight coefficients, is the maneuvering acceleration, is the smoothing operation threshold.

[0051] In one embodiment, the above step S203 includes: generating first evaluation sample data representing the driving state according to the driver state sample data; generating second evaluation sample data representing the driving skill according to the action prediction sample data and the driver action sample data; and evaluating the driver according to the first evaluation sample data and the second evaluation sample data to obtain the driver evaluation sample data.

[0052] The calculation formula for the first evaluation sample data is: , , , , where, is the first evaluation sample data, is the attention score, is the tension score, is the fatigue score, 、 、 are the scoring weights, is the deviation between the measured value of the blink frequency and the normal value of the blink frequency, is the deviation between the measured value of the fixation time and the normal value of the fixation time, is the deviation between the measured value of the fixation deviation and the normal value of the fixation deviation, is the deviation between the measured value of the heart rate and the normal value of the heart rate, is the deviation between the measured value of the heart rate variability and the normal value of the heart rate variability, is the maximum allowable deviation between the measured value of the blink frequency and the normal value of the blink frequency, is the maximum allowable deviation between the measured value of the fixation time and the normal value of the fixation time, is the maximum allowable deviation between the measured value of the fixation deviation and the normal value of the fixation deviation, is the maximum allowable deviation between the measured value of the heart rate and the measured normal value of the heart rate, is the maximum allowable deviation between the measured value of the heart rate variability and the normal value of the heart rate variability.

[0053] The calculation formula for the second evaluation sample data is as follows: , , , , , , , , , , , Wherein, is the second evaluation sample data, is the driver skill score data, is the driving environment complexity index, R is the airspace condition parameter, T is the traffic density parameter, and W is the meteorological condition parameter. , , , , and are the scoring weights, is the accuracy score, is the stability score, is the sensitivity score, is the predicted action value at the t-th moment, is the actual operation value at the t-th moment, is the number of sampling points within the time window T, is the maximum allowable deviation, is the average action deviation within the time window, is the action deviation, is the action fluctuation within the time window, is the set maximum allowable fluctuation, is the penalty coefficient for the fluctuation on the score, is the deviation change rate after standardization processing, is the deviation change rate, is the designed maximum allowable deviation change rate, is the sliding step size of the sliding time window.

[0054] The calculation formula for the driver evaluation sample data is as follows: , Wherein, is the driver evaluation sample data, and is the scoring weight.

[0055] In one embodiment, the above step S204 includes: predicting the driving scenario at a future moment according to the flight sample data and the driver evaluation sample data to obtain multiple future scenario feature sample data; calculating the scenario reward data corresponding to each future scenario feature sample data; and outputting the corresponding driving right allocation sample data according to the scenario reward data.

[0056] Specifically, the model to be trained determines the current driving ability of the driver and the risk of the current flight environment according to the flight sample data and the driver evaluation sample data, so as to predict the driving operation of the driver when dealing with the current flight environment and the driving scenario at a future moment. Furthermore, multiple future scenario feature sample data representing the possible driving scenarios at a future moment are generated according to the predicted driving operation of the driver and the driving scenario at a future moment. Then, the scenario reward data corresponding to each future scenario feature sample data is calculated, and whether the driver has the ability to cope with the risk of the most likely future scenario is predicted according to the calculated scenario reward data. And the corresponding driving right allocation sample data is obtained according to the prediction result. If the driver can cope, the driving right allocation sample data for allocating the driving right to the driver is obtained. If the driver cannot cope, the driving right allocation sample data for allocating the driving right to the automatic driving system is obtained. The driving right allocation sample data may include a driving response plan predicted according to the flight sample data.

[0057] In one embodiment, predicting the driving scenario at a future moment according to the flight sample data and the driver evaluation sample data to obtain multiple future scenario feature sample data includes: extracting features from the flight sample data and the driver evaluation sample data through continuous convolution to obtain the current scenario feature sample data; and predicting the driving scenario at a future moment according to the transformation probability of the current scenario feature sample data at a future moment to obtain multiple future scenario feature sample data.

[0058] Specifically, the model to be trained is a model with a convolutional neural network and a probability prediction network. By performing continuous convolution on the flight sample data and the driver evaluation sample data to extract the local parameter features of both the flight sample data and the driver evaluation sample data, the current scenario feature sample data is obtained. Then, the current scenario feature sample data and the time series feature are input into the probability prediction network together to capture the time series dependence relationship between the current scenario feature sample data and the time series feature. Furthermore, the driving scenario at a future moment is predicted according to the transformation probability at a future moment to obtain multiple future scenario feature sample data.

[0059] In one embodiment, the scenario reward data is obtained by fitting the first reward sample data, the second reward sample data, the third reward sample data, the fourth reward sample data, and the fifth reward sample data.

[0060] Among them, the first reward sample data represents the flight efficiency of the driving scenario corresponding to the future scenario feature sample data, the second reward sample data represents the safety of the driving scenario corresponding to the future scenario feature sample data, the third reward sample data represents the smoothness of the driving scenario corresponding to the future scenario feature sample data, the fourth reward sample data represents the success rate of human-machine collaboration in the driving scenario corresponding to the future scenario feature sample data, and the fifth reward sample data represents the quality of the strategy for generating the future scenario feature sample data.

[0061] The calculation formula for the scenario reward data is: , Among them, is the scenario reward data, , , , and are weight coefficients, is the first reward sample data, is the second reward sample data, is the third reward sample data, is the fourth reward sample data, is the fifth reward sample data.

[0062] The calculation formula for the first reward sample data is: , , , Among them, and are weight coefficients, is the estimated shortest flight time based on the map and traffic conditions, is the cumulative flight time during the flight of the flying car, is the average energy consumption of the flying car under standard working conditions, is the energy consumption rate during the flight of the flying car.

[0063] The calculation formula for the second reward sample data is: , , , Among them, and is the weight coefficient, is the minimum distance between the flying car and surrounding obstacles, is the safety distance, is the attitude angle deviation of the flying car, is the stability threshold.

[0064] The calculation formula for the third reward sample data is: , where, is the weight coefficient, is the maneuvering acceleration, is the smoothing operation threshold.

[0065] The calculation formula for the fourth reward sample data is: , where, is the weight coefficient, is the success rate of human-machine collaboration.

[0066] The calculation formula for the fifth reward sample data is: , , , where, and are the weight coefficients, is the improvement ratio of the performance index under the current policy compared with the previous iteration after iteration, is the success rate of the monitoring system in making emergency responses to emergencies.

[0067] In one embodiment, step S205 above includes: determining the model loss information corresponding to the driving right allocation sample data; iteratively adjusting the network parameters in the model to be trained according to the action prediction loss information until the training end condition is met, and obtaining the driving right allocation model. Among them, the model loss information is obtained by fitting the second policy loss information, the second value loss information, and the second entropy loss information. The second policy loss information represents the quality of the policy for generating the driving right allocation sample data, the second value loss information represents the value deviation between the driving right allocation sample data and the true driving right allocation data, and the second entropy loss information represents the uncertainty of the driving right allocation sample data.

[0068] Specifically, iteratively adjusting the network parameters in the model to be trained according to the model loss information may involve presetting a loss threshold range and a reset threshold as the training end conditions. When the model loss information is within the loss threshold range and the training reset times reach the reset threshold, the training is ended, and the model to be trained in the last iteration obtained is the driving right allocation model. When the training reset times do not reach the reset threshold, adjust the network parameters of the model to be trained according to the deviation degree of the model loss information from the loss threshold range, so that the model loss information gradually approaches the loss threshold range and finally falls within the loss threshold range during the iteration process. Then, obtain new flight sample data, predict driving actions based on the flight sample data to obtain action prediction sample data, and then evaluate the driver according to the driver state sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data as the basis for evaluating the driving ability of the driver. Input the flight sample data and driver evaluation sample data into the preset model to be trained to obtain new driving right allocation sample data until the model loss information is within the loss threshold range. Repeat the above steps until the training reset times reach the reset threshold, and then end the training to obtain the driving right allocation model, which can enable the driving right allocation model to learn the dependence degree of different driver states and driver operation actions on the driving right allocation result.

[0069] Figure 3 It is a flowchart of a driving right allocation method provided by an embodiment of the present application. Refer to Figure 3 In one embodiment, the method includes but is not limited to steps S301 to S305.

[0070] Step S301, obtain flight data.

[0071] Among them, the flight data includes flight environment data, driver state data, driver action data, and vehicle condition data. The flight environment data may include environmental data such as obstacle positions, dynamic states, climate environments, no-fly zone positions, and air traffic densities. The driver state data may include physiological data such as the driver's saccade frequency and heart rate. The driver action data may include action data such as speed input, altitude input, and direction input of the driver's actions. The vehicle condition data may include data such as the flight speed, flight altitude, pitch angle, heading angle, roll angle, and angular velocity of the flying car.

[0072] Step S302, predict driving actions based on the flight data to obtain action prediction data.

[0073] Among them, the action prediction data refers to the driver action data obtained by predicting the driving actions performed by the driver at a later time based on the flight data at a previous time.

[0074] In one embodiment, predicting driving actions based on flight data may be using a pre-trained driving action prediction model to predict the driving actions performed by the driver at a later time based on the flight data at a previous time, thereby obtaining corresponding action prediction data.

[0075] Step S303: Evaluate the driver based on the driver status data, driver action data, and action prediction data to obtain driver evaluation data.

[0076] Among them, the driver evaluation data refers to the evaluation data for predicting the driving ability of the driver at the prediction time based on the driver status data, driver action data, and the corresponding action prediction data.

[0077] In one embodiment, evaluating the driver based on the driver status data, driver action data, and action prediction data may be using a preset driver evaluation model to predict the driving ability of the driver at the current time based on the driver status data, driver action data, and action prediction data at the same time, and evaluating the driving ability of the driver at the current time to obtain driver evaluation data.

[0078] Step S304: Input the flight data and driver evaluation data into the driving authority allocation model to obtain driving authority allocation data.

[0079] Among them, the driving authority allocation model is trained by the above model training method.

[0080] Among them, the driving authority allocation data refers to the data used to indicate whether to allocate the driving authority to the driver or to the autonomous driving system of the flying car.

[0081] In one embodiment, after inputting the flight data and driver evaluation data into the pre-trained driving authority allocation model, the driving authority allocation model determines the current driving ability of the driver and the risk of the current flight environment based on the flight data and driver evaluation data, predicts whether the current driving ability of the driver can cope with the risk of the current flight environment, and obtains corresponding driving authority allocation data according to the prediction result. If it can cope, the driving authority allocation data for allocating the driving authority to the driver is obtained. If it cannot cope, the driving authority allocation data for allocating the driving authority to the autonomous driving system is obtained. The driving authority allocation data may include a driving response plan predicted based on the flight data.

[0082] Step S305: Perform a driving authority allocation action according to the driving authority allocation data.

[0083] The driving right allocation method provided by the embodiments of the present application obtains flight data in real time, predicts driving actions based on the flight data to obtain action prediction data, then evaluates the driver according to the driver status data, driver action data, and action prediction data to obtain driver evaluation data as the basis for evaluating the driving ability of the driver, inputs the flight data and driver sample data into a pre-trained driving right allocation model to obtain driving right allocation data, and finally performs driving right allocation actions according to the driving right allocation data. It can accurately identify the current driving ability of the driver and determine the driving right allocation data. When automatically allocating the driving right according to the driving right allocation data, the flight safety of the flying car can be improved.

[0084] Figure 4 It is a block diagram of an electronic device shown according to an exemplary embodiment.

[0085] The following refers to Figure 4 to describe the electronic device 400 according to this embodiment of the present disclosure. Figure 4 The electronic device 400 shown is only an example and should not impose any restrictions on the functions and usage scope of the embodiments of the present disclosure.

[0086] As Figure 4 shown, the electronic device 400 is presented in the form of a general-purpose computing device. The components of the electronic device 400 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410), a display unit 440, etc.

[0087] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 410, so that the processing unit 410 executes the steps according to various exemplary embodiments of the present disclosure described in the method part of this specification.

[0088] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 4201 and / or a cache storage unit 4202, and may further include a read-only storage unit (ROM) 4203.

[0089] The storage unit 420 may further include a program / utility 4204 having a set (at least one) of program modules 4205. Such program modules 4205 include, but are not limited to: an action system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0090] The bus 430 can represent one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the various bus architectures.

[0091] The electronic device 400 can also communicate with one or more external devices 400' (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 450. Moreover, the electronic device 400 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 460. The network adapter 460 can communicate with other modules of the electronic device 400 through the bus 430. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0092] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0093] The model training method, driving right allocation method, device, and storage medium provided by the embodiment of the present application predict driving actions based on flight sample data to obtain action prediction sample data, and then evaluate the driver according to the driver state sample data, driver action sample data, and action prediction sample data to obtain driver evaluation sample data as the basis for evaluating the driving ability of the driver. The flight sample data and the driver evaluation sample data are input into a preset model to be trained, and the model to be trained is trained using the driving right allocation sample data obtained by the model to be trained, and a driving right allocation model is obtained after the training is completed. Since the driver evaluation sample data is obtained by evaluating the driver according to the driver state sample data, driver action sample data, and action prediction sample data, inputting the flight sample data and the driver evaluation sample data into the model to be trained and training the model to be trained using the driving right allocation sample data can enable the model to be trained to learn the degree of dependence of different driver states and driver operation actions on the driving right allocation result during the training process, so that the trained driving right allocation model can accurately identify the current driving ability of the driver and determine the driving right allocation data. When automatically allocating the driving right according to the driving right allocation data, the flight safety of the flying car can be improved.

[0094] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above methods according to the embodiments of the present disclosure.

[0095] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0096] The computer-readable storage medium may include a data signal propagated in a baseband or as a part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program used by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0097] Those skilled in the art can understand that the above modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are different from the present embodiment only. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0098] The above specifically shows and describes the exemplary embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the detailed structures, setting manners, or implementation methods described herein; on the contrary, the present disclosure covers various modifications and equivalent settings included within the spirit and scope of the appended claims.

Claims

1. A model training method, characterized in that, Including: Obtain flight sample data; the flight sample data includes flight environment sample data, driver state sample data, driver action sample data, and vehicle condition sample data; Perform driving action prediction based on the flight sample data to obtain action prediction sample data; Evaluate the driver based on the driver state sample data, the driver action sample data, and the action prediction sample data to obtain driver evaluation sample data; Input the flight sample data and the driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data; Train the model to be trained according to the driving right allocation sample data to obtain a driving right allocation model.

2. The model training method according to claim 1, wherein The performing driving action prediction based on the flight sample data to obtain action prediction sample data includes: Input the flight sample data into a pre-trained driving action prediction model to obtain the action prediction sample data; the driving action prediction model is trained according to the action prediction loss information of the action prediction sample data, and the action prediction loss information is obtained by fitting the first policy loss information, the first value loss information, and the first entropy loss information. The first policy loss information represents the quality of the policy for generating the action prediction sample data, the first value loss information represents the value deviation between the action prediction sample data and the real action data, and the first entropy loss information represents the uncertainty of the action prediction sample data.

3. The model training method according to claim 1, wherein The evaluating the driver based on the driver state sample data, the driver action sample data, and the action prediction sample data to obtain driver evaluation sample data includes: Generate first evaluation sample data representing the driving state according to the driver state sample data; Generate second evaluation sample data representing the driving skill according to the action prediction sample data and the driver action sample data; Evaluate the driver according to the first evaluation sample data and the second evaluation sample data to obtain the driver evaluation sample data.

4. The model training method according to claim 1, wherein The inputting the flight sample data and the driver evaluation sample data into a preset model to be trained to obtain driving right allocation sample data includes: Predict the driving scenario at a future moment according to the flight sample data and the driver evaluation sample data to obtain multiple future scenario feature sample data; Calculate the scenario reward data corresponding to each of the future scenario feature sample data; Output the corresponding driving right allocation sample data according to the scenario reward data.

5. The model training method according to claim 4, wherein The predicting the driving scenario at a future moment according to the flight sample data and the driver evaluation sample data to obtain multiple future scenario feature sample data includes: Perform feature extraction on the flight sample data and the driver evaluation sample data through continuous convolution to obtain current scenario feature sample data; Predict the driving scenario at a future moment according to the transformation probability of the current scenario feature sample data at the future moment to obtain multiple future scenario feature sample data.

6. The model training method according to claim 4, wherein The described scenario reward data is obtained by fitting the first reward sample data, the second reward sample data, the third reward sample data, the fourth reward sample data, and the fifth reward sample data. The first reward sample data represents the flight efficiency of the driving scenario corresponding to the future scenario feature sample data. The second reward sample data represents the safety level of the driving scenario corresponding to the future scenario feature sample data. The third reward sample data represents the smoothness of the driving scenario corresponding to the future scenario feature sample data. The fourth reward sample data represents the success rate of human-machine collaboration in the driving scenario corresponding to the future scenario feature sample data. The fifth reward sample data represents the quality of the strategy for generating the future scenario feature sample data.

7. The model training method according to claim 1, wherein Training the model to be trained according to the driving right allocation sample data to obtain a driving right allocation model includes: Determining the model loss information corresponding to the driving right allocation sample data; the model loss information is obtained by fitting the second policy loss information, the second value loss information, and the second entropy loss information. The second policy loss information represents the quality of the strategy for generating the driving right allocation sample data. The second value loss information represents the value deviation between the driving right allocation sample data and the true driving right allocation data. The second entropy loss information represents the uncertainty of the driving right allocation sample data; Iteratively adjusting the network parameters in the model to be trained according to the model loss information until the training end condition is met, and obtaining the driving right allocation model.

8. A driving right allocation method, characterized in that, Including: Obtaining flight data; the flight data includes flight environment data, driver status data, driver action data, and vehicle condition data; Performing driving action prediction according to the flight data to obtain action prediction data; Evaluating the driver according to the driver status data, the driver action data, and the action prediction data to obtain driver evaluation data; Inputting the flight data and the driver evaluation data into the driving right allocation model to obtain driving right allocation data; The driving right allocation model is trained by the model training method according to any one of claims 1 to 7; Performing a driving right allocation action according to the driving right allocation data.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Driving right distribution method, device and equipment and readable storage medium

    CN115123301A

  • Target maneuver prediction method based on LSTM, electronic equipment and storage medium

    CN115480582A

  • Driving behavior prediction method and device, electronic equipment and storage medium

    CN115718890A

  • Driving load grading evaluation method based on hidden Markov model

    CN116946149A

  • Driving right interaction decision-making method for hovercar man-machine co-driving

    CN116954202A