Vehicle state control method and system
By constructing a vehicle status model, the problem of unmanned intervention in fully automatic vehicle control and self-evaluation is solved, and accurate control action prediction and effect evaluation is achieved to ensure the automation and accuracy of vehicle control.
Patent Information
- Application Number
- CN202211392759.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-11-08
AI Technical Summary
The prior art cannot achieve end-to-end fully automatic vehicle control without human intervention, and lacks effective control effect evaluation methods.
By building a vehicle status model, including a monitoring data processing module, a map information processing module, a control information input module, a control and prediction first module, a control and prediction second module, a first evaluator group and a second evaluator group, the parameters of the relevant modules are trained and adjusted to achieve independent prediction and control actions and self-evaluation of control effects.
It realizes fully automatic vehicle control without human intervention, and can self-evaluate whether the control actions are in place, ensuring that the predicted control actions are accurate and reasonable, and the control effect is no less than that of manual intervention.
Smart Images

Figure CN115755683B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a vehicle state control method and system. Background Art
[0002] With the development of computer technology and electronically controlled vehicle technology, people can now enjoy more intelligent vehicle control services. Unmanned automatic control is also a major development direction of intelligent vehicle driving. However, the automatic control capabilities of current related technologies are insufficient, and end-to-end fully automatic control without human intervention cannot be achieved. In addition, there is a lack of effective evaluation methods for the control effect of fully automatic control. Summary of the Invention
[0003] In order to solve at least one of the above technical problems, embodiments of the present application provide a vehicle state control method and system.
[0004] On the one hand, an embodiment of the present application provides a vehicle state model training method, wherein the vehicle state model includes a monitoring data processing module, a map information processing module, a manipulation information input module, a first manipulation prediction module, a second manipulation prediction module, a first evaluator group, and a second evaluator group. The monitoring data processing module is connected to the manipulation information input module, the monitoring data processing module and the map information processing module are both connected to the first manipulation prediction module, the monitoring data processing module and the map information processing module are both connected to the second manipulation prediction module, the first manipulation prediction module is connected to the first evaluator group, the second manipulation prediction module is connected to the second evaluator group, and the first evaluator group and the second evaluator group each include at least two evaluators. The method includes:
[0005] Acquiring training information, the training information including a sample manipulation action, map environment information, first monitoring information, second monitoring information, a first motion state, and a second motion state, wherein the first monitoring information is bus data monitored before the sample manipulation action is performed, the second monitoring information is bus data monitored after the sample manipulation action is performed, the first motion state is the motion state of the vehicle when the sample manipulation action is performed, and the second motion state is the motion state of the vehicle at a moment following the first motion state;
[0006] Inputting the sample manipulation action into the manipulation information input module;
[0007] inputting the map environment information into the map information processing module;
[0008] inputting the output of the manipulation information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module;
[0009] inputting the output of the map information processing module, the output of the monitoring data processing module, and the first motion state into the first maneuver prediction module to obtain a first maneuver prediction result;
[0010] inputting the output of the map information processing module and the second motion state into the second maneuver prediction module to obtain a second maneuver prediction result;
[0011] Parameters of related modules are adjusted based on the first manipulation prediction result and the second manipulation prediction result.
[0012] On the other hand, an embodiment of the present application provides a vehicle state control method, the method comprising:
[0013] Monitor the bus data to obtain third monitoring information and fourth monitoring information, wherein the time corresponding to the fourth monitoring information is a time next to the time corresponding to the third monitoring information;
[0014] Acquire a reference motion state, where the reference motion state is the motion state of the vehicle at the moment corresponding to the fourth monitoring information;
[0015] determining a target monitoring difference according to the third monitoring information and the fourth monitoring information;
[0016] Inputting map environment information into a map information processing module to obtain an output of the map information processing module;
[0017] Inputting the output of the map information processing module, the target monitoring difference and the reference motion state into a first control prediction module to obtain a target predicted control action;
[0018] controlling the vehicle according to the target predicted maneuver;
[0019] Wherein, the map information processing module and the first control prediction module both belong to the vehicle state model, and the vehicle state model is trained using the above-mentioned method.
[0020] On the other hand, an embodiment of the present application provides a vehicle state model training device, wherein the vehicle state model includes a monitoring data processing module, a map information processing module, a manipulation information input module, a first manipulation prediction module, a second manipulation prediction module, a first evaluator group, and a second evaluator group. The monitoring data processing module is connected to the manipulation information input module, the monitoring data processing module and the map information processing module are both connected to the first manipulation prediction module, the monitoring data processing module and the map information processing module are both connected to the second manipulation prediction module, the first manipulation prediction module is connected to the first evaluator group, the second manipulation prediction module is connected to the second evaluator group, and the first evaluator group and the second evaluator group each include at least two evaluators. The device includes:
[0021] a training information acquisition module, configured to acquire training information, the training information including a sample manipulation action, map environment information, first monitoring information, second monitoring information, a first motion state, and a second motion state, wherein the first monitoring information is bus data monitored before the sample manipulation action is performed, the second monitoring information is bus data monitored after the sample manipulation action is performed, the first motion state is the motion state of the vehicle when the sample manipulation action is performed, and the second motion state is the motion state of the vehicle at a moment immediately following the first motion state;
[0022] a data processing module configured to input the sample manipulation action into the manipulation information input module; input the map environment information into the map information processing module; input the output of the manipulation information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module; input the output of the map information processing module, the output of the monitoring data processing module, and the first motion state into the first manipulation prediction module to obtain a first manipulation prediction result; and input the output of the map information processing module and the second motion state into the second manipulation prediction module to obtain a second manipulation prediction result;
[0023] A parameter adjustment module is used to adjust the parameters of related modules based on the first control prediction result and the second control prediction result.
[0024] On the other hand, an embodiment of the present application provides a vehicle state control system, the system comprising:
[0025] A monitoring module, configured to monitor bus data to obtain third monitoring information and fourth monitoring information, wherein the time corresponding to the fourth monitoring information is a time next to the time corresponding to the third monitoring information;
[0026] a reference motion state acquisition module, configured to acquire a reference motion state, wherein the reference motion state is the motion state of the vehicle at the moment corresponding to the fourth monitoring information;
[0027] a monitoring difference acquisition module, configured to determine a target monitoring difference according to the third monitoring information and the fourth monitoring information;
[0028] an environment map processing module, configured to input map environment information into the map information processing module and obtain output from the map information processing module;
[0029] a maneuvering information prediction module, configured to input the output of the map information processing module, the target monitoring difference, and the reference motion state into a first maneuvering prediction module to obtain a target predicted maneuvering action;
[0030] a control module, configured to control the vehicle according to the target predicted maneuvering action;
[0031] Wherein, the map information processing module and the first control prediction module both belong to the vehicle state model, and the vehicle state model is trained using the above-mentioned method.
[0032] The embodiment of the present application proposes a vehicle state control method and system. First, a vehicle state model is trained. The vehicle state model has two main functions. One is to set up two modules for predicting control information and their respective evaluator groups, and to adjust the parameters of the modules for predicting control information and their respective evaluator groups so that the modules for predicting control information have the ability to independently predict control actions, and control actions are the information required for vehicle control. Therefore, based on this module, the embodiment of the present application can independently control the vehicle without human intervention. The other is that the vehicle state model can also finally obtain the monitoring difference data corresponding to each single control action. The monitoring difference data represents the inevitable difference between the data on the bus before and after each control action occurs. By actually monitoring the changes in the data in the bus and then comparing the monitoring difference data corresponding to each single control action, it can be known whether the single control action and the combination of control actions are executed in place, thereby achieving the effect of self-evaluation. That is to say, the embodiment of the present application can also independently evaluate whether the vehicle control is in place, and still does not require human intervention. In summary, the embodiment of the present application can independently perform fully automatic vehicle control and determine whether the vehicle control is in place without human intervention.
[0033] Furthermore, the embodiment of the present application requires the use of map information, information monitored in the bus, and the vehicle's motion status to predict control information. That is to say, all available information is effectively utilized to ensure that the predicted control actions are accurate and reasonable, and that correct control decisions are made. The correctness of the decisions makes the implementation effect of the present application no less than the vehicle control implementation effect under manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0035] Figure 1 This is a schematic diagram of a vehicle state model provided in an embodiment of this specification;
[0036] Figure 2 This is a flow chart of a training method provided in an embodiment of the present application;
[0037] Figure 3 It is a flow chart of a vehicle state control method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the embodiments of the present application.
[0039] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0040] In order to make the purpose, technical solutions and advantages disclosed in the embodiments of the present application more clearly understood, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application.
[0041] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0042] This application embodiment provides a vehicle state model training method, please refer to Figure 1 The above-mentioned vehicle state model includes a monitoring data processing module, a map information processing module, a manipulation information input module, a first manipulation prediction module, a second manipulation prediction module, a first evaluator group and a second evaluator group. The above-mentioned monitoring data processing module is connected to the above-mentioned manipulation information input module, the above-mentioned monitoring data processing module and the above-mentioned map information processing module are both connected to the above-mentioned first manipulation prediction module, the above-mentioned monitoring data processing module and the above-mentioned map information processing module are both connected to the above-mentioned second manipulation prediction module, the above-mentioned first manipulation prediction module is connected to the above-mentioned first evaluator group, the above-mentioned second manipulation prediction module is connected to the above-mentioned second evaluator group, and the above-mentioned first evaluator group and the above-mentioned second evaluator group each include at least two evaluators.
[0043] Please refer to Figure 2 , the above method includes:
[0044] S101. Obtain training information, which includes sample control actions, map environment information, first monitoring information, second monitoring information, a first motion state and a second motion state. The first monitoring information is the bus data monitored before the sample control action is implemented, and the second monitoring information is the bus data monitored after the sample control action is implemented. The first motion state is the motion state of the vehicle when the sample control action is implemented, and the second motion state is the motion state of the vehicle at the next moment corresponding to the first motion state.
[0045] This application does not limit the sample control actions, which can be a single control action or a combination of control actions on the vehicle, such as turning left and slowing down, turning right and slowing down, going straight and accelerating, etc. The map environment information is the high-precision map information of the road section where the vehicle is traveling.
[0046] S102. Input the sample manipulation action into the manipulation information input module.
[0047] Specifically, the above-mentioned sample manipulation actions can be subjected to information extraction based on the above-mentioned manipulation information input module to obtain corresponding outputs. Specifically, the sample manipulation actions are extracted into information that can be processed by other related modules. This application does not process the information extraction and can refer to the existing technology.
[0048] S103. Input the above map environment information into the above map information processing module.
[0049] Specifically, the map environment information can be vectorized based on the map information processing module to obtain a corresponding output. The vectorization process is not described in detail in the embodiment of the present application, and reference can be made to the prior art.
[0050] S104. Input the output of the control information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module.
[0051] Monitoring difference information may be determined based on the difference between the first monitoring information and the second monitoring information, and this information may be used as the output of the monitoring data processing module.
[0052] At the same time, the monitoring data processing module can also obtain mapping information based on the output of the manipulation information input module and the monitoring difference information. The mapping information represents the correspondence between the sample operation and the monitoring difference information. For example, the monitoring difference information indicates that bits 1, 3, 5, 8, and 9 of the bus data have undergone corresponding changes. These changes are likely related to the sample operation, thereby constructing the mapping information.
[0053] Of course, there may be individual changes that are unrelated to the sample operation action. For example, only 1, 3, and 5 are related to the sample operation action. That is, as long as the sample operation action is implemented, 1, 3, and 5 will inevitably produce corresponding changes. The changes in bits 8 and 9 are also reflected in the mapping information, which means that the mapping information is noisy, and the noise can be eliminated by performing autocorrelation analysis on a large amount of mapping information. The embodiment of the present application uses noisy monitoring difference information as the output of the monitoring data processing module. This is because this noise is only regarded as "noise" because it is unrelated to the sample operation action, but this noise also reflects the state of the vehicle and is also valid data for predicting control information. Therefore, it is also output by the monitoring data processing module.
[0054] S105. Input the output of the map information processing module, the output of the monitoring data processing module and the first motion state into the first control prediction module to obtain a first control prediction result.
[0055] This application does not limit the structure of the first control prediction module. It only needs to be able to output the first control prediction result based on the output of the above-mentioned map information processing module, the output of the above-mentioned monitoring data processing module, and the above-mentioned first motion state. Any mainstream neural network with prediction capabilities can be selected. Therefore, the structure is not limited. It should be pointed out that the specific structure of each single module in the embodiment of this application can refer to the existing technology, but the organic combination of these modules to form the vehicle state model in this application, as well as the training method of the vehicle state model are all original designs of this application.
[0056] The first manipulation prediction result is a single manipulation action or a combination thereof.
[0057] S106. Input the output of the map information processing module and the second motion state into the second control prediction module to obtain a second control prediction result.
[0058] S107. Adjust parameters of related modules based on the first manipulation prediction result and the second manipulation prediction result.
[0059] Specifically, the second manipulation prediction result can be input into each evaluator in the second evaluator group to obtain the first evaluation prediction information output by each evaluator according to the preset noise signal, and the evaluation prediction information represents the expected value of the reward under the assumption that the second manipulation prediction result is continuously executed for a preset time. The embodiment of the present application does not limit the specific structure of each evaluator used, which can be obtained by a modeling method, and the modeling method can use relevant technologies. Thus, it has the ability to output the expected value of the reward. This application focuses on the specific steps of using the expected value of the reward for parameter adjustment and the model parameter adjustment based on the expected value of the reward. As for the modeling process of the evaluator, no further details will be given.
[0060] A reward prediction target is determined based on the product of the minimum value of each of the first evaluation prediction information and a preset attenuation coefficient. In some embodiments, the reward prediction target can be obtained by adding a preset reward coefficient to the product of the minimum value of each of the first evaluation prediction information and the preset attenuation coefficient. However, both the attenuation coefficient and the preset reward coefficient can be set according to actual circumstances.
[0061] The first manipulation prediction result is input into each evaluator in the first evaluator group to obtain second evaluation prediction information output by each evaluator based on the preset noise signal. The second evaluation prediction information is based on the same inventive concept as the first evaluation prediction information and will not be further described here.
[0062] The evaluation prediction target is obtained based on the weighted square sum of each of the second evaluation prediction information. The weights can be set according to the situation. In some scenarios, they can all be set to 1, that is, the square sum of each of the second evaluation prediction information is directly used as the evaluation prediction target.
[0063] Based on the difference between the reward prediction target and the evaluation prediction target, the parameters of the first manipulation prediction module and the first evaluator group are updated.
[0064] Specifically, the parameters of the first evaluator group are adjusted using gradient descent, and the parameters of the first maneuver prediction module are adjusted using gradient ascent. Furthermore, the parameters of the second maneuver prediction module are updated based on the parameters of the first maneuver prediction module, and the parameters of the second evaluator group are updated based on the parameters of the first evaluator group. Updating one module from another can also be considered parameter update propagation, which can be referenced in existing technologies and is not limited here. Of course, the parameters of other modules not mentioned do not need to be adjusted.
[0065] During the model training process, a large amount of mapping information is generated. Autocorrelation analysis is performed on all of this mapping information to generate a target mapping record table. Each record in this target mapping record table represents the correspondence between a single control action and the corresponding change in the monitored data bit. For example, an acceleration action corresponds to a change in bits 1 and 4 of the bus data, while a steering action corresponds to a change in bits 14 and 63.
[0066] The present application also provides a vehicle state control method, such as Figure 3 As shown, the above method includes:
[0067] S201. Monitor bus data to obtain third monitoring information and fourth monitoring information, where the time corresponding to the fourth monitoring information is the next time after the time corresponding to the third monitoring information.
[0068] The time corresponding to the third monitoring information may be any time during the process of the vehicle state control method being put into use.
[0069] S202. Obtain a reference motion state, where the reference motion state is the motion state of the vehicle at the moment corresponding to the fourth monitoring information.
[0070] S203 . Determine a target monitoring difference according to the third monitoring information and the fourth monitoring information.
[0071] Specifically, the target monitoring difference refers to the difference between the third monitoring information and the fourth monitoring information.
[0072] S204. Input the map environment information into the map information processing module to obtain the output of the map information processing module.
[0073] Please refer to the previous text for the operation of this step.
[0074] S205. Input the output of the map information processing module, the target monitoring difference and the reference motion state into the first control prediction module to obtain the target predicted control action.
[0075] S206. Control the vehicle according to the target predicted maneuver. The various modules used in the vehicle state control method belong to a vehicle state model, which is trained using the method described above.
[0076] Obviously, the vehicle control method can fully automatically affect the vehicle state without manual intervention. In addition, the target prediction control action can be input into the monitoring data processing module in the vehicle state model to obtain target monitoring difference data; based on the monitoring data before the target prediction control action is executed and the monitoring data after the target prediction control action is executed, the measured monitoring difference data is obtained; and the control result is evaluated based on the difference between the target monitoring difference data and the measured monitoring difference data. For example, based on the known target mapping record table and the target prediction control action, the difference in monitoring data when the target prediction control action is fully implemented can be obtained, that is, the target monitoring difference data. The difference between the target monitoring difference data and the measured monitoring difference data can be used to determine whether the target prediction control action is fully implemented. If not, which link is not in place can be directly obtained by analyzing the difference between the target monitoring difference data and the measured monitoring difference data without manual intervention.
[0077] In one embodiment, a vehicle state model training device is further provided. The vehicle state model includes a monitoring data processing module, a map information processing module, a manipulation information input module, a first manipulation prediction module, a second manipulation prediction module, a first evaluator group, and a second evaluator group. The monitoring data processing module is connected to the manipulation information input module. The monitoring data processing module and the map information processing module are both connected to the first manipulation prediction module. The monitoring data processing module and the map information processing module are both connected to the second manipulation prediction module. The first manipulation prediction module is connected to the first evaluator group. The second manipulation prediction module is connected to the second evaluator group. The first evaluator group and the second evaluator group each include at least two evaluators. The device includes:
[0078] a training information acquisition module, configured to acquire training information, the training information including sample manipulation actions, map environment information, first monitoring information, second monitoring information, a first motion state, and a second motion state, wherein the first monitoring information is bus data monitored before the sample manipulation action is performed, the second monitoring information is bus data monitored after the sample manipulation action is performed, the first motion state is the motion state of the vehicle when the sample manipulation action is performed, and the second motion state is the motion state of the vehicle at the next moment corresponding to the first motion state;
[0079] a data processing module configured to input the sample manipulation action into the manipulation information input module; input the map environment information into the map information processing module; input the output of the manipulation information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module; input the output of the map information processing module, the output of the monitoring data processing module, and the first motion state into the first manipulation prediction module to obtain a first manipulation prediction result; and input the output of the map information processing module and the second motion state into the second manipulation prediction module to obtain a second manipulation prediction result;
[0080] A parameter adjustment module is used to adjust the parameters of related modules based on the above-mentioned first control prediction result and the above-mentioned second control prediction result.
[0081] In one embodiment, a vehicle state control system is further provided, the system comprising:
[0082] A monitoring module, configured to monitor bus data to obtain third monitoring information and fourth monitoring information, wherein the time corresponding to the fourth monitoring information is the next time after the time corresponding to the third monitoring information;
[0083] A reference motion state acquisition module, configured to acquire a reference motion state, wherein the reference motion state is the motion state of the vehicle at the moment corresponding to the fourth monitoring information;
[0084] a monitoring difference acquisition module, configured to determine a target monitoring difference based on the third monitoring information and the fourth monitoring information;
[0085] An environmental map processing module, configured to input map environmental information into a map information processing module and obtain an output from the map information processing module;
[0086] a control information prediction module, configured to input the output of the map information processing module, the target monitoring difference, and the reference motion state into a first control prediction module to obtain a target predicted control action;
[0087] A control module, configured to control the vehicle according to the target prediction manipulation action;
[0088] Among them, the above-mentioned map information processing module and the above-mentioned first control prediction module both belong to the vehicle state model, and the above-mentioned vehicle state model is trained using the above-mentioned method.
[0089] The device part and the system part in the embodiments of the present application are based on the same inventive concept as the respective method embodiments, and will not be described in detail here.
[0090] The above is only a preferred embodiment of the embodiment of the present application and is not intended to limit the embodiment of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiment of the present application should be included in the scope of protection of the embodiment of the present application.
Claims
1. A vehicle state model training method, characterized in that: The vehicle state model includes a monitoring data processing module, a map information processing module, a manipulation information input module, a first manipulation prediction module, a second manipulation prediction module, a first evaluator group, and a second evaluator group. The monitoring data processing module is connected to the manipulation information input module. The monitoring data processing module and the map information processing module are both connected to the first manipulation prediction module. The monitoring data processing module and the map information processing module are both connected to the second manipulation prediction module. The first manipulation prediction module is connected to the first evaluator group. The second manipulation prediction module is connected to the second evaluator group. The first evaluator group and the second evaluator group each include at least two evaluators. The method includes: Acquiring training information, the training information including a sample manipulation action, map environment information, first monitoring information, second monitoring information, a first motion state, and a second motion state, wherein the first monitoring information is bus data monitored before the sample manipulation action is performed, the second monitoring information is bus data monitored after the sample manipulation action is performed, the first motion state is the motion state of the vehicle when the sample manipulation action is performed, and the second motion state is the motion state of the vehicle at a moment following the first motion state; Inputting the sample manipulation action into the manipulation information input module; inputting the map environment information into the map information processing module; inputting the output of the manipulation information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module; inputting the output of the map information processing module, the output of the monitoring data processing module, and the first motion state into the first maneuver prediction module to obtain a first maneuver prediction result; inputting the output of the map information processing module and the second motion state into the second maneuver prediction module to obtain a second maneuver prediction result; Adjusting parameters of relevant modules based on the first manipulation prediction result and the second manipulation prediction result, wherein adjusting parameters of relevant modules based on the first manipulation prediction result and the second manipulation prediction result includes: Inputting the second maneuver prediction result into each evaluator in the second evaluator group to obtain first evaluation prediction information output by each evaluator based on a preset noise signal, wherein the evaluation prediction information represents an expected value of a reward under the assumption that the second maneuver prediction result is continuously executed for a preset time; Determining a reward prediction target based on the product of the minimum value of each of the first evaluation prediction information and a preset attenuation coefficient; Inputting the first manipulation prediction result into each evaluator in the first evaluator group to obtain second evaluation prediction information output by each evaluator according to a preset noise signal; Obtaining an evaluation prediction target based on a weighted square sum of each of the second evaluation prediction information; Based on the difference between the reward prediction target and the evaluation prediction target, parameters of the first manipulation prediction module and the first evaluator group are updated.
2. The method according to claim 1, characterized in that The updating of parameters of the first manipulation prediction module and the first evaluator group based on the difference between the reward prediction target and the evaluation prediction target includes: Adjusting the parameters of the first evaluator group by gradient descent; The parameters of the first control prediction module are adjusted in a gradient ascent manner.
3. The method according to claim 2, characterized in that The adjusting parameters of related modules based on the first manipulation prediction result and the second manipulation prediction result further includes: updating parameters of the second maneuvering prediction module based on parameters of the first maneuvering prediction module; The parameters of the second evaluator group are updated based on the parameters of the first evaluator group.
4. The method according to claim 1, wherein The inputting the sample manipulation action into the manipulation information input module includes: extracting information of the sample manipulation action based on the manipulation information input module to obtain a corresponding output; The inputting the map environment information into the map information processing module includes: performing vectorization processing on the map environment information based on the map information processing module to obtain corresponding output; The step of inputting the output of the manipulation information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module includes: Obtaining monitoring difference information based on the first monitoring information and the second monitoring information; Obtaining mapping information according to the output of the manipulation information input module and the monitoring difference information, wherein the mapping information represents a correspondence between the sample manipulation action and the monitoring difference information; The monitoring difference information is used as the output of the monitoring data processing module.
5. The method according to claim 4, characterized in that The method further comprises: Autocorrelation analysis is performed on all mapping information to obtain a target mapping record table, wherein each record in the target mapping record table represents a correspondence between a single manipulation action and a change bit of corresponding monitoring data.
6. A vehicle state control method, characterized in that: The method comprises: Monitor the bus data to obtain third monitoring information and fourth monitoring information, wherein the time corresponding to the fourth monitoring information is a time next to the time corresponding to the third monitoring information; Acquire a reference motion state, where the reference motion state is the motion state of the vehicle at the moment corresponding to the fourth monitoring information; determining a target monitoring difference according to the third monitoring information and the fourth monitoring information; Inputting map environment information into a map information processing module to obtain an output of the map information processing module; Inputting the output of the map information processing module, the target monitoring difference and the reference motion state into a first control prediction module to obtain a target predicted control action; controlling the vehicle according to the target predicted maneuver; Wherein, the map information processing module and the first maneuver prediction module both belong to a vehicle state model, and the vehicle state model is trained using the method according to any one of claims 1 to 5.
7. The method according to claim 6, characterized in that The method further comprises: Inputting the target predicted control action into the monitoring data processing module in the vehicle state model to obtain target monitoring difference data; obtaining measured monitoring difference data based on the monitoring data before the target prediction manipulation action is performed and the monitoring data after the target prediction manipulation action is performed; The control result is evaluated according to the difference between the target monitoring difference data and the measured monitoring difference data.
8. A vehicle state model training device, characterized in that: The vehicle state model includes a monitoring data processing module, a map information processing module, a manipulation information input module, a first manipulation prediction module, a second manipulation prediction module, a first evaluator group, and a second evaluator group. The monitoring data processing module is connected to the manipulation information input module. The monitoring data processing module and the map information processing module are both connected to the first manipulation prediction module. The monitoring data processing module and the map information processing module are both connected to the second manipulation prediction module. The first manipulation prediction module is connected to the first evaluator group. The second manipulation prediction module is connected to the second evaluator group. The first evaluator group and the second evaluator group each include at least two evaluators. The device includes: a training information acquisition module, configured to acquire training information, the training information including a sample manipulation action, map environment information, first monitoring information, second monitoring information, a first motion state, and a second motion state, wherein the first monitoring information is bus data monitored before the sample manipulation action is performed, the second monitoring information is bus data monitored after the sample manipulation action is performed, the first motion state is the motion state of the vehicle when the sample manipulation action is performed, and the second motion state is the motion state of the vehicle at a moment immediately following the first motion state; a data processing module configured to input the sample manipulation action into the manipulation information input module; input the map environment information into the map information processing module; input the output of the manipulation information input module, the first monitoring information, and the second monitoring information into the monitoring data processing module; input the output of the map information processing module, the output of the monitoring data processing module, and the first motion state into the first manipulation prediction module to obtain a first manipulation prediction result; and input the output of the map information processing module and the second motion state into the second manipulation prediction module to obtain a second manipulation prediction result; A parameter adjustment module is used to adjust the parameters of related modules based on the first control prediction result and the second control prediction result.
9. A vehicle state control system, characterized in that: The system comprises: A monitoring module, configured to monitor bus data to obtain third monitoring information and fourth monitoring information, wherein the time corresponding to the fourth monitoring information is a time next to the time corresponding to the third monitoring information; A reference motion state acquisition module, configured to acquire a reference motion state, wherein the reference motion state is the motion state of the vehicle at the moment corresponding to the fourth monitoring information; a monitoring difference acquisition module, configured to determine a target monitoring difference according to the third monitoring information and the fourth monitoring information; an environment map processing module, configured to input map environment information into the map information processing module and obtain output from the map information processing module; a maneuvering information prediction module, configured to input the output of the map information processing module, the target monitoring difference, and the reference motion state into a first maneuvering prediction module to obtain a target predicted maneuvering action; a control module, configured to control the vehicle according to the target predicted maneuvering action; Wherein, the map information processing module and the first maneuver prediction module both belong to a vehicle state model, and the vehicle state model is trained using the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Vehicle control method and device, electronic equipment and storage medium
CN114802233A