Automatic driving vehicle control method, vehicle and storage medium
Patent Information
- Application Number
- CN202611013991.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请实施例提供一种自动驾驶车辆控制方法、车辆及存储介质,以至少解决相关技术中对自动驾驶车辆的控制准确性较差的技术问题
[0019]根据本申请实施例的另一方面,还提供了一种计算机程序产品,包括计算机程序,计算机程序在被处理器执行时实现本申请各个实施例中的方法。
Smart Images

Figure CN122830746A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving, and more specifically, to an autonomous vehicle control method, a vehicle, and a storage medium. Background Technology
[0002] With the deepening application of artificial intelligence technology in the field of intelligent transportation, autonomous driving decision-making based on deep neural networks has become one of the core technologies for achieving autonomous vehicle driving. In complex road environments, autonomous vehicles need to process high-dimensional data from multiple sources of sensors such as LiDAR, cameras, and millimeter-wave radar in real time, and extract features and output decisions through neural network models to complete key tasks such as path planning, speed control, and obstacle avoidance. However, current control strategies have poor accuracy in controlling autonomous vehicles, thus reducing their driving safety.
[0003] There is currently no good solution to the above problems. Summary of the Invention
[0004] This application provides an autonomous vehicle control method, a vehicle, and a storage medium to at least solve the technical problem of poor control accuracy of autonomous vehicles in related technologies.
[0005] According to one aspect of the embodiments of this application, an autonomous vehicle control method is provided, comprising: acquiring environmental perception data and historical decision data of an autonomous vehicle; inputting the environmental perception data and historical decision data into a control decision model; using the environmental perception data and historical decision data to determine the parameter update state of an extraction layer, wherein the parameter update state is used to determine whether to update the extraction parameters of the extraction layer, and the control decision model consists of a decision layer and an extraction layer; based on the parameter update state, controlling the extraction layer and the decision layer to process the environmental perception data and historical decision data to obtain a control strategy for the autonomous vehicle; and controlling the autonomous vehicle to drive based on the control strategy.
[0006] Furthermore, by utilizing environmental perception data and historical decision data, the parameter update status of the extraction layer is determined, including: based on the extraction layer, performing feature extraction on the environmental perception data to obtain decision features; and based on the environmental perception data, historical decision data, and decision features, determining the parameter update status.
[0007] Furthermore, based on environmental perception data, historical decision data, and decision characteristics, the parameter update state is determined, including: determining a first fluctuation state based on environmental perception data, historical decision data, and decision characteristics, wherein the first fluctuation state is used to characterize the environmental fluctuation situation in the current control cycle; determining a second fluctuation state based on historical decision data, wherein the second fluctuation state is used to characterize the environmental fluctuation situation in the historical control cycle; and determining the parameter update state based on the first fluctuation state and the second fluctuation state.
[0008] Further, based on environmental perception data, historical decision data, and decision characteristics, the first fluctuation state is determined, including: acquiring a training dataset for the control decision model, wherein the training dataset includes training perception data and training decision data, wherein the data type of the training perception data is the same as that of the environmental perception data, and the data type of the training decision data is the same as that of the historical decision data; determining a statistical distribution index based on the decision characteristics and the training dataset, wherein the statistical distribution index is used to characterize the similarity between the decision characteristics and the training dataset; and determining the first fluctuation state based on the statistical distribution index, environmental perception data, and historical decision data.
[0009] Furthermore, based on the decision features and the training dataset, statistical distribution indicators are determined, including: determining training features based on training perception data and training decision data; and evaluating the similarity between the training features and the decision features to determine statistical distribution indicators.
[0010] Furthermore, based on statistical distribution indicators, environmental perception data, and historical decision-making data, the first fluctuation state is determined, including: determining decision fluctuation indicators based on historical decision-making data, wherein the decision fluctuation indicators are used to characterize the stability of historical decision-making data; determining the environmental type and environmental change rate based on environmental perception data; and constructing the first fluctuation state based on statistical distribution indicators, decision fluctuation indicators, environmental type, and environmental change rate.
[0011] Further, based on the first fluctuation state and the second fluctuation state, the parameter update state is determined, including: determining a first duration for which the second fluctuation state meets the environmental fluctuation conditions; if the first duration is greater than a preset time and the first fluctuation state meets the environmental fluctuation conditions, the parameter update state is determined to be an active state, wherein the active state is used to characterize the state of updating the extracted parameters; if the first duration is greater than the preset time and the first fluctuation state does not meet the environmental fluctuation conditions, or if the first duration is less than or equal to the preset time, the parameter update state is determined to be a frozen state, wherein the frozen state is used to characterize the state of not updating the extracted parameters.
[0012] Furthermore, the method also includes: determining a rate change threshold based on the environment type; determining that the first fluctuation state meets the environmental fluctuation conditions when the decision fluctuation index is less than or equal to a preset fluctuation index, the environmental change rate is less than or equal to the rate change threshold, and the statistical distribution index is less than or equal to a preset distribution index; and determining that the first fluctuation state does not meet the environmental fluctuation conditions when the decision fluctuation index is greater than the preset fluctuation index, the environmental change rate is greater than the rate change threshold, or the statistical distribution index is greater than the preset distribution index.
[0013] Furthermore, based on the parameter update state, the control extraction layer and the decision layer process the environmental perception data and historical decision data to obtain the control strategy for the autonomous vehicle, including: when the parameter update state is frozen, determining the first update range of the decision parameters of the decision layer based on the preset update strategy; determining the parameters to be updated from the decision parameters based on the first update range; updating the parameters to be updated based on the environmental perception data and historical decision data to obtain the updated decision layer; and processing the decision features based on the updated decision layer to obtain the control strategy.
[0014] Furthermore, based on environmental perception data and historical decision data, the parameters to be updated are updated to obtain the updated decision layer, including: determining the loss value based on historical decision data and the loss function of the decision layer; determining the learning rate of the decision layer based on environmental perception data; and updating the parameters to be updated based on the loss value and the learning rate to obtain the updated decision layer.
[0015] Furthermore, the method also includes: when the parameter update state is in an active state, obtaining a second duration during which the parameter update state is in an active state; updating the decision parameters based on environmental perception data and historical decision data to obtain an updated decision layer; determining a second update range for the extracted parameters based on the second duration; and updating the extracted parameters based on the second update range, environmental perception data, and historical decision data to obtain an updated extraction layer.
[0016] According to another aspect of the embodiments of this application, an autonomous vehicle control device is also provided, comprising: a data acquisition module for acquiring environmental perception data and historical decision data of an autonomous vehicle; a state determination module for inputting the environmental perception data and historical decision data into a control decision model, and using the environmental perception data and historical decision data to determine the parameter update state of the extraction layer, wherein the parameter update state is used to determine whether to update the extraction parameters of the extraction layer, and the control decision model consists of a decision layer and an extraction layer; a strategy construction module for controlling the extraction layer and the decision layer to process the environmental perception data and historical decision data based on the parameter update state to obtain a control strategy for the autonomous vehicle; and a vehicle control module for controlling the autonomous vehicle to drive based on the control strategy.
[0017] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0018] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0019] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0020] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.
[0021] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0022] In this embodiment, the following methods are employed: acquiring environmental perception data and historical decision data of an autonomous vehicle; inputting the environmental perception data and historical decision data into a control decision model; using the environmental perception data and historical decision data to determine the parameter update state of the extraction layer; based on the parameter update state, controlling the extraction layer and the decision layer to process the environmental perception data and historical decision data to obtain a control strategy for the autonomous vehicle; and based on the control strategy, controlling the driving mode of the autonomous vehicle. By dynamically determining the parameter update state of the extraction layer through the environmental perception data and historical decision data, the parameter update logic of the extraction layer and the decision layer can be differentiated based on the parameter update state. This allows for flexible adaptation of the decision layer while ensuring the stability of core feature extraction, ultimately outputting a highly reliable control strategy to accurately control vehicle driving. This achieves the technical effect of improving the control accuracy of autonomous vehicles, thereby solving the technical problem of poor control accuracy of autonomous vehicles in related technologies. Attached Figure Description
[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0024] Figure 1 This is a flowchart of an autonomous vehicle control method according to an embodiment of this application;
[0025] Figure 2 This is a schematic diagram of an optional parameter update state determination process according to an embodiment of this application;
[0026] Figure 3 This is a schematic diagram of an optional process for determining a first fluctuation state according to an embodiment of this application;
[0027] Figure 4 This is a schematic diagram of an optional first fluctuation state and environmental fluctuation condition matching process according to an embodiment of this application;
[0028] Figure 5 This is a schematic diagram of an optional parameter update process according to an embodiment of this application;
[0029] Figure 6 This is a schematic diagram of an autonomous vehicle control device according to an embodiment of this application. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] According to an embodiment of this application, an embodiment of an autonomous vehicle control method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0033] This embodiment provides a method for controlling an autonomous vehicle. Figure 1 This is a flowchart of an autonomous vehicle control method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:
[0034] Step S102: Obtain environmental perception data and historical decision data of the autonomous vehicle.
[0035] The aforementioned autonomous vehicles can refer to intelligent transportation tools that possess environmental perception, path planning, and automatic execution of control operations such as steering, acceleration, or braking, enabling them to drive autonomously without human driver intervention or with minimal monitoring.
[0036] The aforementioned environmental perception data can refer to multi-source heterogeneous data collected and preprocessed by autonomous vehicles through sensors such as LiDAR, cameras, millimeter-wave radar, and inertial measurement units (IMUs). This data may include, but is not limited to, point cloud coordinates, image pixel information, distance, velocity, and azimuth of target objects, as well as the attitude acceleration of the autonomous vehicle itself. It can be used to construct a real-time three-dimensional model and state description of the surrounding environment.
[0037] The aforementioned historical decision data can refer to the control command records generated by autonomous vehicles in the past operating time series and the corresponding environmental state feedback information, which may include, but are not limited to, the action outputs such as steering angle, throttle opening, and braking force, as well as the timestamps, sensor input features, and decision uncertainty indicators associated with these actions. It can be used to evaluate the stability of historical decisions and as a basis for judging the parameter freezing mechanism of artificial intelligence models.
[0038] In one alternative embodiment, considering the acquisition of the external environmental state and internal control history of the autonomous vehicle, a basic input for multi-source data fusion can be constructed, thereby providing the artificial intelligence model with spatiotemporally consistent feature vectors so that the artificial intelligence model can accurately output control actions for the autonomous vehicle.
[0039] Based on this, the autonomous vehicle control system (hereinafter referred to as the control system) can periodically collect LiDAR point clouds, camera images, millimeter-wave radar target lists, and inertial measurement unit data through sensor hardware interfaces, and use timestamp alignment technology to synchronize multimodal data to a unified time reference. Then, the key physical quantities of each sensor of the autonomous vehicle are extracted and stitched together to form the above-mentioned environmental perception data.
[0040] Meanwhile, the control system can read the action output sequence within the historical decision cycle through the vehicle bus interface, store the continuous moment decision values in a sliding window, thereby forming a comprehensive dataset containing current environmental perception data and historical decision data, providing accurate statistical basis for the output of subsequent control actions.
[0041] For example, the control system can acquire multi-source sensor data from LiDAR, millimeter-wave radar, cameras, and inertial measurement units via the CAN (Controller Area Network) bus interface of the autonomous vehicle. After time-stamping synchronization and spatial registration, this multi-source sensor data can generate perception feature vectors containing environmental geometry, dynamic obstacle positions, and motion states, serving as environmental perception data. Simultaneously, the control system can read the control strategy sequences and corresponding state vectors generated over the past N time steps from the historical output buffer of the decision module, serving as historical decision data.
[0042] In another alternative embodiment, the control system can receive standardized environmental state messages from distributed perception nodes via the autonomous vehicle's onboard Ethernet communication protocol. These messages may include a list of predicted trajectories and pose information for the autonomous vehicle after Kalman filtering, serving as environmental perception data. Simultaneously, the control system can record the action execution results, reward function values, and state transition errors of the most recent M control cycles through an internal circular buffer, constructing historical decision sequence data.
[0043] In another alternative embodiment, the control system can also utilize edge computing units to perform feature dimensionality reduction and compression on the original sensor stream, extracting key scene semantic labels and structured features such as relative speed and distance of obstacles as environmental perception data. The control system can also extract historical decision trajectories, control deviation values, and system response times of the autonomous vehicle in the same or similar scenarios as historical decision data by querying event logs stored in non-volatile memory.
[0044] Step S104: Input environmental perception data and historical decision data into the control decision model. Use the environmental perception data and historical decision data to determine the parameter update status of the extraction layer. The parameter update status is used to determine whether to update the extraction parameters of the extraction layer. The control decision model consists of a decision layer and an extraction layer.
[0045] The aforementioned control decision model can refer to an algorithmic architecture based on neural networks used to generate control strategies for autonomous vehicles. The internal logical structure of this model can be clearly divided into two parts: an extraction layer responsible for low-level feature abstraction and a decision layer responsible for high-level policy output. The extraction and decision layers can work together to complete the mapping from environmental perception to control actions.
[0046] The aforementioned parameter update status can refer to the real-time running flag set for the weight parameters in the extraction layer. This parameter update status can be dynamically calculated and determined by the control decision model based on the input environmental perception data and historical decision data. It can be used to indicate whether the extraction parameters of the extraction layer in the current control cycle are in an updatable state that allows gradient backpropagation for fine-tuning, or in a frozen state that locks the weights and prohibits updates.
[0047] The aforementioned extraction parameters can refer to the basic weight matrix and bias term set that constitute the extraction layer of the control decision model. These extraction parameters can be specifically used to perform nonlinear transformations such as convolution, pooling, or fully connected transformations on environmental perception data and historical decision data input from multiple sources of sensors, thereby extracting key high-dimensional feature vectors that can characterize the state of environmental fluctuations.
[0048] The aforementioned decision layer can refer to a neural network submodule in the control decision model that receives the abstract feature vector output from the extraction layer, performs contextual reasoning based on the vector, and finally outputs a specific control strategy. This decision layer can contain fully connected layers and activation functions, and is responsible for mapping the feature space to the action space.
[0049] In one alternative embodiment, it is considered that by determining the parameter update state of the extraction layer, the feature extraction stability and decision adaptability of the control decision model can be dynamically balanced, thereby improving the control decision accuracy and robustness of autonomous vehicles in complex dynamic environments.
[0050] Based on this, the control system can receive and preprocess environmental perception data and historical decision data from multiple sources such as lidar, cameras and millimeter-wave radar, and can input the processed data into the control decision model composed of the extraction layer and the decision layer.
[0051] Subsequently, within the control decision model, environmental perception data and historical decision data can be analyzed to determine the current environmental fluctuations and the rate of environmental change. Next, the control system uses the monitored input data distribution characteristics, historical decision output fluctuations, and the rate of environmental change as indicators, comparing them with preset freeze thresholds. If these indicators exceed the freeze thresholds, the control system determines the parameter update state of the extraction layer to be frozen; if these indicators are below the freeze thresholds, the control system determines the parameter update state of the extraction layer to be active.
[0052] If the parameter update state is frozen, the control system can lock the extraction parameters of the extraction layer and prohibit the extraction layer from performing gradient backpropagation to maintain the stability of the core feature extraction capability. If the parameter update state is active, the control system can allow the extraction parameters to participate in gradient updates to maintain adaptability to environmental changes. Simultaneously, the control system can pass the updated feature vector output by the extraction layer to the decision layer for decision output, thereby realizing the process of dynamically adjusting the behavior of the extraction layer based on the parameter update state and improving the output accuracy of the control decision model.
[0053] For example, the control system can input environmental perception data and historical decision data into the control decision model to determine the statistical distribution variance of the feature vector output by the extraction layer, and the standard deviation of the fluctuation of the historical decision output sequence. If the feature variance exceeds a preset multiple of the training set variance or the standard deviation of the decision fluctuation exceeds a preset threshold, the control system can set the parameter update state of the extraction layer to a frozen state to prohibit parameter updates. Otherwise, the control system can set the parameter update state of the extraction layer to an active state to allow parameter updates.
[0054] In another alternative embodiment, the control system can input environmental perception data and historical decision data into the control decision model, and can obtain the prediction uncertainty estimate of the output features of the extraction layer during multiple forward propagations. Subsequently, the control system can compare this uncertainty estimate with a preset uncertainty threshold. If the preset uncertainty threshold is greater than or equal to the uncertainty estimate, the control system can set the parameter update state of the extraction layer to a frozen state to prohibit the updating of the extracted parameters. If the preset uncertainty threshold is less than the uncertainty estimate, the control system can set the parameter update state of the extraction layer to an active state to allow the extraction parameters to be updated.
[0055] In another optional embodiment, the control system can input environmental perception data and historical decision data into the control decision model to determine the feature distance between the current environmental perception data and the training set data, as well as the rate of change of the historical decision data within a preset time window. If the feature distance is greater than a preset distance threshold, or the rate of change is greater than a preset rate of change threshold, the control system can set the parameter update state of the extraction layer to a frozen state to prohibit parameter updates. If the feature distance is less than or equal to the preset distance threshold, and the rate of change is less than or equal to the preset rate of change threshold, the control system can set the parameter update state of the extraction layer to an active state to allow parameter updates.
[0056] Step S106: Based on the parameter update state, the control extraction layer and the decision layer process the environmental perception data and historical decision data to obtain the control strategy of the autonomous vehicle.
[0057] The aforementioned control strategy can refer to a set of instructions generated by the extraction layer based on the parameter update state, which extracts features from environmental perception data and historical decision data, and then performs calculations based on the output of the extraction layer. This set of instructions guides the autonomous vehicle to perform specific actions such as steering, acceleration, or braking. In complex dynamic environments, this control strategy can balance the stability and real-time adaptability of the decision output through parameter freezing in the extraction layer and dynamic fine-tuning in the decision layer, thereby generating vehicle control instructions that meet both safety and comfort requirements.
[0058] In one alternative embodiment, considering that by dynamically adjusting the extraction layer and decision layer according to the parameter update state, the processing logic of environmental perception data and historical decision data can effectively balance the stability of feature extraction and the adaptability of decision output of the control decision model, thereby achieving precise control of autonomous vehicles in complex dynamic environments.
[0059] Based on this, the control system can monitor the parameter update status of the extraction layer in the control decision model and determine whether the parameter update status is in a frozen state. If the system detects that the input data distribution deviates from the training set or the rate of environmental change exceeds a threshold, the control system can lock the parameter update status of the extraction layer in a frozen state, prohibiting the extraction layer from participating in gradient backpropagation, thereby ensuring the continuity and consistency of environmental feature extraction.
[0060] Subsequently, the control system can input the frozen feature vector from the extraction layer into the decision layer. The decision layer, based on the parameter update status of the extraction layer and the correlation between the extraction and decision layers, determines the update range of its parameters. This decision layer can then combine this feature vector, environmental perception data, and historical decision data, using a pre-defined bi-objective loss function to make minor adjustments to its parameters. This loss function comprehensively considers task execution error and decision stability constraints, enabling the decision layer to smoothly adjust its parameters based solely on the stable features provided by the extraction layer.
[0061] Ultimately, the adjusted decision layer can output a control strategy for autonomous vehicles that meets the requirements of real-time performance and safety based on this feature vector. The entire process combines static feature anchoring at the extraction layer with dynamic fine-tuning at the decision layer, achieving high-precision and robust output of the control strategy in response to sudden environmental disturbances.
[0062] For example, when the parameter update state is frozen, the control system can lock the weight parameters of the extraction layer to fix the environmental feature extraction capability. Simultaneously, the control decision model can receive environmental perception data and input it into the frozen extraction layer to extract feature vectors. These feature vectors, along with historical decision data sequences, can be input into the decision layer. The decision layer can calculate control commands based on a preset constraint fine-tuning strategy and a bi-objective loss function, performing a limited update on the decision layer's parameters after gradient pruning. Finally, the updated decision layer can re-receive the feature vectors output by the extraction layer, thereby outputting a control strategy for the autonomous vehicle that includes steering angle and throttle / brake opening. When the parameter update state is active, the control system can allow the extraction layer and the decision layer to participate in both forward and backward propagation, and can process environmental perception data and historical decision data based on a full-network gradient update mechanism to directly output the control strategy for the autonomous vehicle.
[0063] In another optional embodiment, when the parameter update state is frozen, the extraction layer can output a fixed-dimensional environmental feature embedding vector. The decision layer can input this embedding vector and the historical decision sequence into a temporal processing module with an attention mechanism. This temporal processing module can determine the correlation weight between the current input features and the historical states, perform weighted aggregation of the historical decision data, and then map it through a fully connected layer to obtain the control strategy. At this time, the decision parameters of the decision layer can be locally updated according to the denoising autoencoder loss function to suppress noise interference. When the parameter update state is active, the environmental feature embedding vector output by the extraction layer and the historical decision data can be directly input into the end-to-end control network. This control network can jointly determine the task loss and regularization loss through the backpropagation algorithm, and can simultaneously update the convolutional kernel parameters of the extraction layer and the fully connected layer parameters of the decision layer based on the task loss and regularization loss. Thus, the updated extraction layer and decision layer can be used to generate a control strategy that adapts to the current instantaneous environmental changes.
[0064] Step S108: Based on the control strategy, control the autonomous vehicle to drive.
[0065] In one alternative embodiment, considering dividing the control decision model into an extraction layer for feature extraction and a decision layer for generating action commands, and using a dynamic triggering mechanism based on input data distribution characteristics, historical decision fluctuations and environmental change rates, the decision layer can be used for constraint fine-tuning while the extraction layer parameters are frozen. This can significantly suppress decision oscillations caused by sensor noise or environmental abrupt changes, thereby achieving smooth and adaptive autonomous driving control while ensuring the stability of core perception capabilities.
[0066] Based on this, the control system can construct specific control commands for autonomous vehicles according to the control strategy output by the control decision model, and control the autonomous vehicle to drive according to the control commands, so as to ensure that the autonomous vehicle remains safe and stable during driving.
[0067] In this embodiment, the following methods are employed: acquiring environmental perception data and historical decision data of an autonomous vehicle; inputting the environmental perception data and historical decision data into a control decision model; using the environmental perception data and historical decision data to determine the parameter update state of the extraction layer; based on the parameter update state, controlling the extraction layer and the decision layer to process the environmental perception data and historical decision data to obtain a control strategy for the autonomous vehicle; and based on the control strategy, controlling the driving mode of the autonomous vehicle. By dynamically determining the parameter update state of the extraction layer through the environmental perception data and historical decision data, the parameter update logic of the extraction layer and the decision layer can be differentiated based on the parameter update state. This allows for flexible adaptation of the decision layer while ensuring the stability of core feature extraction, ultimately outputting a highly reliable control strategy to accurately control vehicle driving. This achieves the technical effect of improving the control accuracy of autonomous vehicles, thereby solving the technical problem of poor control accuracy of autonomous vehicles in related technologies.
[0068] Furthermore, by utilizing environmental perception data and historical decision data, the parameter update status of the extraction layer is determined, including: based on the extraction layer, performing feature extraction on the environmental perception data to obtain decision features; and based on the environmental perception data, historical decision data, and decision features, determining the parameter update status.
[0069] The aforementioned decision features can refer to an abstract representation vector output by the extraction layer after preprocessing the input environmental perception data and extracting the core features. This vector carries the key semantic information of the environmental state.
[0070] In one alternative embodiment, by considering the characteristic distribution of environmental perception data and the fluctuation pattern of historical decision outputs, a data-driven state assessment mechanism can be established, which can achieve accurate perception and dynamic control of the operating state of the control decision model.
[0071] Based on this, the control system can extract features from environmental perception data using the extraction layer to obtain decision features. This allows for feature mapping of the multi-source environmental perception data input to the control decision model in both spatial and temporal dimensions, generating decision features with low-dimensional semantic information. The decision feature vector represents the similarity between the feature distribution of the current environmental state and the training dataset, as well as the degree of dynamic change in the current environment.
[0072] Subsequently, the control system can determine the statistical distribution index of the decision characteristics and the standard deviation of historical decision data within a preset time window. This statistical distribution index and standard deviation can then be used as a comprehensive index. The control system can then compare this comprehensive index with a corresponding threshold. If the comprehensive index exceeds the threshold, the control system can determine that the extraction layer enters a parameter freeze state to maintain the stability of feature extraction. If the comprehensive index is within the normal range, the control system can determine that the extraction layer maintains an updatable parameter state to preserve the adaptability of the control decision model to environmental changes, thereby achieving dynamic switching of parameter update states based on real-time environmental conditions and decision stability.
[0073] Furthermore, based on environmental perception data, historical decision data, and decision characteristics, the parameter update state is determined, including: determining a first fluctuation state based on environmental perception data, historical decision data, and decision characteristics, wherein the first fluctuation state is used to characterize the environmental fluctuation situation in the current control cycle; determining a second fluctuation state based on historical decision data, wherein the second fluctuation state is used to characterize the environmental fluctuation situation in the historical control cycle; and determining the parameter update state based on the first fluctuation state and the second fluctuation state.
[0074] The aforementioned first fluctuation state can refer to a characterization index determined by comprehensively considering environmental perception data, historical decision data, and decision characteristics within the current control cycle. This first fluctuation state can be used to quantify and reflect the intensity of fluctuations or uncertainty levels in the environment in which the autonomous vehicle is located at the current moment, and reflects the strength of external disturbances or state changes faced by the control system at the present instant.
[0075] The aforementioned current control cycle can refer to the time period or time step that the control system takes to execute a complete decision-making process. It can be used to define the time boundary of a single decision-making behavior, ensure the real-time nature of decision generation, and serve as a relatively recent observation point in time series analysis.
[0076] The aforementioned environmental fluctuations refer to the degree of dynamic change in the external environment in response to the input of the control system. These fluctuations may include, but are not limited to, the noise level of sensor data, changes in the geometric structure of the physical scene, changes in the motion state of the target object, and unforeseen interference factors. Such environmental fluctuations directly affect the distribution characteristics of the input data and the accuracy of the control system's decisions, and are an important external basis for assessing the stability of the decisions.
[0077] The aforementioned second fluctuation state can refer to a characterization index determined based on the decision output data accumulated within the historical control cycle. This second fluctuation state can reflect the trend of decision stability of the control system in the recent historical period or the continuous impact of environmental disturbances by statistically analyzing the changes in decision results over multiple past time steps, thus reflecting the smoothness of the control system's response to environmental changes over a period of time.
[0078] The aforementioned historical control cycle can refer to one or more consecutive time step sequences that have been completed before the current control cycle. These time step sequences can constitute a sequence of past operating states of the control system, used to store and backtrack previous decision results and corresponding environmental states, thereby providing a data basis for determining the second fluctuation state and helping the control system identify long-term trends or periodic characteristics of environmental changes.
[0079] In one alternative embodiment, considering that by decoupling current environmental disturbances from historical decision-making inertia, the real-time sensitivity and long-term stability of the control decision-making model in a dynamic environment can be quantified, thereby providing an objective mathematical basis for fine-tuning the parameters of the extraction layer and the decision-making layer.
[0080] Based on this, the control system can determine the statistical variance and major eigenvalues of the multi-source sensor feature distribution within the current control cycle based on environmental perception data, and can combine historical decision data to derive the first fluctuation state, so as to characterize the degree of deviation of the current environment from the feature distribution of the training dataset and the rate of dynamic change.
[0081] Subsequently, the control system can determine the standard deviation based on the decision output sequence within the historical control cycle to derive the second fluctuation state, which characterizes the output stationarity of the control decision model within the recent time window.
[0082] Finally, the control system can combine the comparison results of the first fluctuation state with the environmental change threshold, and the comparison results of the second fluctuation state with the decision fluctuation threshold, to determine whether the parameter update state is frozen and limit the parameter update amplitude of the decision layer. If the control system determines that the parameter update state is active, it can perform regular gradient updates on the extraction layer and the decision layer, thereby achieving policy-based adaptive parameter control.
[0083] For ease of understanding, Figure 2 This is a schematic diagram illustrating an optional parameter update state determination process according to an embodiment of this application, such as... Figure 2As shown, the control system can determine a first fluctuation state based on environmental perception data, historical decision data, and decision characteristics to reflect the environmental fluctuations of the current control cycle. Subsequently, the control system can determine a second fluctuation state based on historical decision data to reflect the environmental fluctuations of previous control cycles. Finally, the control system can make a comprehensive judgment based on the first and second fluctuation states and preset state thresholds to determine the parameter update state.
[0084] Further, based on environmental perception data, historical decision data, and decision characteristics, the first fluctuation state is determined, including: acquiring a training dataset for the control decision model, wherein the training dataset includes training perception data and training decision data, wherein the data type of the training perception data is the same as that of the environmental perception data, and the data type of the training decision data is the same as that of the historical decision data; determining a statistical distribution index based on the decision characteristics and the training dataset, wherein the statistical distribution index is used to characterize the similarity between the decision characteristics and the training dataset; and determining the first fluctuation state based on the statistical distribution index, environmental perception data, and historical decision data.
[0085] The aforementioned training dataset can refer to a data set used for pre-training and parameter initialization of the control decision model. It can be used to establish the basic feature extraction capability and decision mapping relationship of the control decision model, and provide an initial weight benchmark for online fine-tuning in the subsequent operation phase.
[0086] The aforementioned training perception data can refer to a set of input samples in the training dataset that are consistent with the environmental perception data in terms of data type, format, and semantics. It can be used to simulate real-world environmental inputs during the training phase of the control decision model in order to extract stable feature representations.
[0087] The aforementioned training decision data can refer to the set of expected output labels in the training dataset that perfectly match the historical decision data in terms of data type, dimension, and physical meaning. It can be used in the training phase of the control decision model to correct the model parameters through supervised learning or reinforcement learning reward mechanisms, so that the output of the control decision model meets the decision criteria for safety and performance requirements.
[0088] The aforementioned statistical distribution index can refer to a quantitative value determined based on decision features and training dataset. It can be used to characterize the distribution of decision features in the current operating environment from a statistical perspective, and the degree of difference or similarity between it and the feature distribution learned during the model training phase, so as to reflect the deviation of environmental data.
[0089] The aforementioned similarity can refer to the degree of closeness between the decision characteristics reflected by the statistical distribution index and the characteristics of historical training data in the distribution space. The higher the value of the similarity, the more likely the current input environment is within the safe operating domain that the model is familiar with. The lower the value of the similarity, the more likely the current environment has an anomaly or extreme situation that is not fully covered by the training data, thereby triggering the subsequent fluctuation state assessment and freezing mechanism.
[0090] In an optional embodiment, considering that by establishing a mapping relationship between the statistical benchmark of the training set and the real-time sensing data, the degree of deviation of environmental characteristics can be quantified, and on this basis, the fluctuation characteristics of historical decisions can be correlated to provide an objective and calculable stability assessment basis for the control decision model.
[0091] Based on this, the control system can obtain a training dataset for the control decision model. This training dataset can include training perception data and training decision data. The data type of the training perception data is the same as that of the environmental perception data, and the data type of the training decision data is the same as that of the historical decision data.
[0092] Subsequently, the control system can calculate the mean and variance of the decision feature vectors output by the extraction layer, and compare these statistics with the statistical distribution of the corresponding features in the training dataset. This allows the determination of numerical indicators that characterize the overlap or deviation between the current environmental features and the training environmental features distribution, which serve as the aforementioned statistical distribution indicators.
[0093] Subsequently, the control system can determine the rate of environmental change in the environmental perception data, and the fluctuation value of historical decisions from recent decision outputs in the historical decision data. Then, the control system can perform weighted fusion or logical judgment on the statistical distribution indicators, the rate of environmental change, and the fluctuation value of historical decisions to generate the first fluctuation state quantity.
[0094] Furthermore, based on the decision features and the training dataset, statistical distribution indicators are determined, including: determining training features based on training perception data and training decision data; and evaluating the similarity between the training features and the decision features to determine statistical distribution indicators.
[0095] The aforementioned training features can refer to the high-dimensional abstract feature vectors output from the extraction layer during the training phase of the control decision model, which represent the environmental state after preprocessing the training perception data and performing forward propagation of the neural network. These feature vectors reflect the inherent distribution patterns and core semantic information of the environmental data in the training set.
[0096] In one alternative embodiment, considering that the difference between the feature spaces of the training phase and the running phase can be used to objectively characterize the degree of environmental disturbance and model confidence, a calculable trigger basis can be provided for the freezing mechanism of parameter update state.
[0097] Based on this, the control system can extract the benchmark feature vector formed in the stable training state as training features through forward propagation of neural networks, based on training perception data and training decision data, so as to establish the theoretical benchmark of feature distribution.
[0098] Subsequently, the control system can acquire the decision features corresponding to the current input during the real-time control process, and evaluate the similarity between the real-time decision features and the benchmark training features. This evaluation process can quantify the degree to which the current input data deviates from the training distribution by determining the statistical distance or distribution overlap between the two in the multidimensional feature space.
[0099] Ultimately, the control system can map the evaluation results into a statistical distribution index, which can specifically reflect the matching degree between the current environmental state and the model pre-training state. This matching degree can be used as a quantitative parameter to determine whether to execute the freeze strategy, thus realizing adaptive control decision based on data distribution characteristics.
[0100] Furthermore, based on statistical distribution indicators, environmental perception data, and historical decision-making data, the first fluctuation state is determined, including: determining decision fluctuation indicators based on historical decision-making data, wherein the decision fluctuation indicators are used to characterize the stability of historical decision-making data; determining the environmental type and environmental change rate based on environmental perception data; and constructing the first fluctuation state based on statistical distribution indicators, decision fluctuation indicators, environmental type, and environmental change rate.
[0101] The aforementioned decision fluctuation index can be determined based on historical decision data and is a quantitative value used to characterize the stability of historical decision data. Specifically, it can be expressed as the standard deviation of the change in decision output over multiple recent decision cycles. The larger the value of this decision fluctuation index, the more drastic the fluctuation of the decision output in the time series, reflecting the stability level of the control system in the recent operation process.
[0102] The aforementioned environment type can refer to the driving or operating scenario categories classified by performing scene clustering analysis on multi-source sensor data, such as stationary scenario, low-speed scenario, medium-speed scenario, high-speed scenario, or extreme emergency scenario. This environment type can be used to distinguish the macroscopic operating environment characteristics of the current control system in order to adapt to different control strategy thresholds.
[0103] The aforementioned rate of environmental change can be a quantitative indicator determined based on the differential sensor data within a continuous time step, which can be used to characterize the drastic degree of dynamic changes in the external environment. For example, it can be reflected by calculating the average movement distance of each point in the lidar point cloud between frames. The larger the value of this rate of environmental change, the higher the frequency or magnitude of changes in the environmental state. Therefore, this rate of environmental change can be directly related to the instability of the sensing input of the control system.
[0104] In one alternative embodiment, considering that multi-dimensional quantitative indicators can comprehensively characterize the operational status of the control decision model in a complex dynamic environment, specifically, the decision fluctuation index can be used to objectively quantify the dispersion of historical decision outputs, and the environmental type and rate of change in environmental perception data can be used to objectively quantify the intensity and frequency of external disturbances, thereby constructing a first fluctuation state that can fully reflect the stability and adaptability requirements of the current control system.
[0105] Based on this, the control system can determine the decision fluctuation index using historical decision data. The control system can quantify the smoothness and consistency of the decision layer's output under current operating conditions by statistically analyzing the variance or standard deviation of the decision layer's output over the last P control cycles. The larger the value of the decision fluctuation index, the lower the stability of the historical decision data and the stronger the volatility.
[0106] Subsequently, the control system can determine the environment type and rate of change based on the environmental sensing data. The control system can use clustering algorithms or rule matching to map the multi-source sensing data currently collected by the sensors to a preset set of environment types, and determine the Euclidean distance or differential amplitude of the feature vectors of the sensing data at adjacent time steps, which serves as the rate of change to quantify the drastic dynamics of the environment.
[0107] Finally, the control system can construct a first fluctuation state based on statistical distribution indicators, decision fluctuation indicators, environmental type, and environmental change rate. It then weights and fuses or logically maps these multiple dimensions of indicators to generate a feature vector representing the overall operating state of the current control decision model. Specifically, the statistical distribution indicator reflects the deviation of the input features from the training distribution, the decision fluctuation indicator reflects the stability of the output response, and the environmental change rate reflects the severity of external disturbances. This forms the foundational state basis for guiding subsequent parameter extraction, freezing, or fine-tuning of the strategy.
[0108] For ease of understanding, Figure 3 This is a schematic diagram illustrating an optional process for determining a first fluctuation state according to an embodiment of this application, as shown below. Figure 3 As shown, the control system can determine training features based on training perception data and training decision data in the training dataset, and can perform similarity evaluation on the training features and decision features to obtain statistical distribution indicators. Subsequently, the control system can determine decision fluctuation indicators based on historical decision data to reflect the stability of historical decision data. The control system can determine the environment type and environmental change rate based on environmental perception data. Finally, the control system can fuse the statistical distribution indicators, decision fluctuation indicators, environment type, and environmental change rate to obtain the first fluctuation state.
[0109] Further, based on the first fluctuation state and the second fluctuation state, the parameter update state is determined, including: determining a first duration for which the second fluctuation state meets the environmental fluctuation conditions; if the first duration is greater than a preset time and the first fluctuation state meets the environmental fluctuation conditions, the parameter update state is determined to be an active state, wherein the active state is used to characterize the state of updating the extracted parameters; if the first duration is greater than the preset time and the first fluctuation state does not meet the environmental fluctuation conditions, or if the first duration is less than or equal to the preset time, the parameter update state is determined to be a frozen state, wherein the frozen state is used to characterize the state of not updating the extracted parameters.
[0110] The aforementioned environmental fluctuation conditions can refer to the judgment criteria used during the operation of the control decision model to determine whether the current input data distribution, historical decision output fluctuations, or environmental change rate exceed the normal safety range. Specifically, it can be manifested as the characteristic distribution variance, decision output standard deviation, or sensor differential calculation value exceeding a preset threshold, thereby indicating that the control system is in a state where output stability needs to be monitored.
[0111] The aforementioned first duration can refer to the length of time during which the second fluctuation state meets the environmental fluctuation conditions, in order to avoid the false triggering of the freeze state caused by instantaneous noise, thereby ensuring the accuracy of the determination of the parameter update state.
[0112] The aforementioned active state can refer to the working state in which the control system determines to update the extraction parameters of the extraction layer using gradients and adjust their weights. In this state, the control system considers the current environmental changes to be within a controllable range or to have an adaptive adjustment requirement. Therefore, the control system can unfreeze the state, allowing the control decision model to fine-tune the parameter weights of the extraction layer based on real-time data through the backpropagation algorithm, thereby enhancing its adaptability to environmental changes.
[0113] The aforementioned frozen state can refer to a working state in which the control system determines not to update the extracted parameters. In this state, the control system believes that the current environment has changed drastically or that there is a high risk of abnormal input. In order to prevent decision oscillation and overfitting, the control system locks the weight matrix and bias term of the extraction layer, prohibits the parameters in the extraction layer from participating in gradient backpropagation, and only allows the decision layer to make limited fine-tuning to maintain the stability and robustness of the decision output.
[0114] In one alternative embodiment, considering that by introducing a state determination mechanism with a duration dimension, it is possible to achieve temporal logic filtering of the coupling relationship between environmental fluctuations and decision fluctuations, thereby effectively distinguishing between transient disturbances and continuous environmental changes, and then accurately and dynamically switching the parameter update state to balance the robustness and adaptability of the control decision model.
[0115] Based on this, the control system can first determine the first duration for which the second fluctuation state continuously meets the environmental fluctuation conditions. Subsequently, the control system can compare the first duration with a preset time threshold. If the first duration is greater than the preset time and the first fluctuation state meets the environmental fluctuation conditions, the control system can determine that the autonomous vehicle has been in a relatively safe driving environment for an extended period. At this point, the control system can determine that the parameter update state is active, and the control decision model allows the parameters of the extraction layer to be updated to adapt to environmental changes.
[0116] If the first duration is greater than a preset time, but the first fluctuation state does not meet the environmental fluctuation conditions, or if the first duration is less than or equal to the preset time, it indicates that the environment in which the autonomous vehicle is currently located is changing drastically. At this time, in order to ensure the output stability of the control decision model and thus guarantee the driving safety of the autonomous vehicle, the control system can determine that the parameter update state is frozen, and the control decision model prohibits updating the parameters of the extraction layer to maintain the stability of feature extraction.
[0117] Furthermore, the method also includes: determining a rate change threshold based on the environment type; determining that the first fluctuation state meets the environmental fluctuation conditions when the decision fluctuation index is less than or equal to a preset fluctuation index, the environmental change rate is less than or equal to the rate change threshold, and the statistical distribution index is less than or equal to a preset distribution index; and determining that the first fluctuation state does not meet the environmental fluctuation conditions when the decision fluctuation index is greater than the preset fluctuation index, the environmental change rate is greater than the rate change threshold, or the statistical distribution index is greater than the preset distribution index.
[0118] The aforementioned rate change threshold can refer to a pre-set environmental dynamics threshold based on the current environment type, used to quantify the degree of change in sensor data. When the calculated rate of environmental change exceeds this threshold, it indicates that the environment in which the autonomous vehicle is located has undergone drastic changes or anomalies, thereby triggering the freezing mechanism of the extraction layer to prevent decision instability.
[0119] The aforementioned preset volatility index can refer to an upper limit reference value for the volatility of decision outputs, derived from historical normal decision data. It can be determined by a fixed multiple of the larger normal decision volatility value or as a percentage of the output amplitude, and is used to measure the drastic changes in recent decision results. When the decision volatility index exceeds this preset volatility index, it indicates that the control system is currently in a high-volatility state, requiring intervention for stabilization control.
[0120] The aforementioned preset distribution index can refer to a statistical threshold used to assess the degree of difference between the current input data feature distribution and the training set feature distribution. It can be set based on a multiple of the variance or standard deviation of the training set features. By comparing the statistical distribution index with the preset distribution index, it can be determined whether the current environment deviates from the normal distribution pre-learned by the control decision model, thereby identifying potential environmental anomalies or distribution shifts.
[0121] In an optional embodiment, considering that determining the rate change threshold based on the environment type can provide differentiated judgment criteria for the environmental dynamic characteristics of different operating scenarios, and combining the condition that the decision fluctuation index, environmental change rate and statistical distribution index are all within the preset threshold range, it is determined that the first fluctuation state meets the environmental fluctuation condition, and that the first fluctuation state does not meet the environmental fluctuation condition when any index exceeds the preset threshold, it is possible to achieve accurate classification and boundary definition of the operating state of the control decision model.
[0122] Based on this, the control system can first match the corresponding rate change threshold according to the current environment type to establish a judgment benchmark adapted to the dynamics of the environment. Subsequently, the control system can compare the decision fluctuation index, the rate of environmental change, and the statistical distribution index with the preset fluctuation index, the rate change threshold, and the preset distribution index one by one.
[0123] When the decision fluctuation index is less than or equal to the preset fluctuation index, the rate of environmental change is less than or equal to the rate change threshold, and the statistical distribution index is less than or equal to the preset distribution index, the control system can determine that the first fluctuation state meets the environmental fluctuation conditions.
[0124] If the decision fluctuation index is greater than the preset fluctuation index, or the rate of environmental change is greater than the rate change threshold, or the statistical distribution index is greater than the preset distribution index, the control system can determine that the first fluctuation state does not meet the environmental fluctuation conditions.
[0125] For ease of understanding, Figure 4 This is a schematic diagram illustrating an optional first fluctuation state and environmental fluctuation condition matching process according to an embodiment of this application, as shown below. Figure 4As shown, the control system can first determine the rate change threshold based on the environment type. Then, it comprehensively considers the decision fluctuation index, the environmental change rate, and the statistical distribution index to determine whether the first fluctuation state meets the environmental fluctuation conditions. If the decision fluctuation index is less than or equal to a preset fluctuation index, the environmental change rate is less than or equal to the rate change threshold, and the statistical distribution index is less than or equal to a preset distribution index, the first fluctuation state is determined to meet the environmental fluctuation conditions. If the decision fluctuation index is greater than the preset fluctuation index, or the environmental change rate is greater than the rate change threshold, or the statistical distribution index is greater than the preset distribution index, the first fluctuation state is determined not to meet the environmental fluctuation conditions. The process for determining whether the second fluctuation state meets the environmental fluctuation conditions is similar to the above process and will not be repeated here.
[0126] Furthermore, based on the parameter update state, the control extraction layer and the decision layer process the environmental perception data and historical decision data to obtain the control strategy for the autonomous vehicle, including: when the parameter update state is frozen, determining the first update range of the decision parameters of the decision layer based on the preset update strategy; determining the parameters to be updated from the decision parameters based on the first update range; updating the parameters to be updated based on the environmental perception data and historical decision data to obtain the updated decision layer; and processing the decision features based on the updated decision layer to obtain the control strategy.
[0127] The aforementioned preset update strategy can refer to a set of rules or constraints used to guide the fine-tuning of decision layer parameters when the extraction layer is frozen. This preset update strategy can limit the magnitude and direction of the adjustment of decision layer parameters while ensuring the output stability of the control decision model, so as to prevent decision oscillation caused by full parameter updates and ensure that the fine-tuning process of decision parameters is smooth and meets safety constraints.
[0128] The aforementioned decision parameters can refer to the complete set of trainable variables that constitute the decision layer in the autonomous vehicle control strategy generation model. These may include, but are not limited to, the weight matrix and bias terms in the decision layer. These decision parameters can map the decision features extracted by the extraction layer to specific vehicle control strategies.
[0129] The aforementioned first update range can refer to the range of parameters that allow the decision parameters to change at the current moment, determined according to a preset update strategy, when the parameter update state is detected to be frozen and a fine-tuning mechanism for the decision parameters is triggered. This limits the step size of a single decision parameter update, ensuring that even if the environment changes drastically, the control output will not deviate from reasonable physical limits or cause abrupt changes in action.
[0130] The aforementioned parameters to be updated can refer to a subset of parameters from all decision parameters in the decision-making layer, which, after initial update range filtering or logical judgment, are determined to actually participate in gradient backpropagation or numerical adjustment within the current time step. These parameters to be updated can be extracted from the decision parameters based on the initial update range, gradient pruning rules, or sensitivity analysis results of a specific layer. Only decision parameters marked as "to be updated" will undergo specific numerical corrections based on environmental perception data and historical decision data, while the remaining parameters remain constant to maintain the output stability of the control decision model.
[0131] In an alternative embodiment, it is considered that by implementing parameter update state differentiation processing on the extraction layer and decision layer in the control decision model, the accuracy and robustness of the control strategy output can be significantly improved, which is conducive to accurately determining the control strategy of autonomous vehicles.
[0132] Based on this, the control system can first monitor the parameter update status of the control decision model. When the parameter update status is frozen, the control system can determine the first update range of the decision parameters of the decision layer according to the preset update strategy. The preset update strategy can be dynamically calibrated based on the historical decision fluctuation threshold or the rate of environmental change to limit the magnitude of the fine adjustment of the decision parameters.
[0133] Subsequently, the control system can filter the parameters to be updated from the decision parameters, ensuring that only sensitive or critical parameters are adjusted locally. Then, the control system can use the current environmental perception data and historical decision data as input to perform gradient descent or incremental update operations on the parameters to be updated, thereby obtaining the updated decision layer.
[0134] This process avoids decision oscillations caused by abrupt changes in decision parameters by limiting the update range of decision parameters in the decision layer. Finally, the control system can input environmental perception data into the frozen extraction layer to obtain stable decision features, and then combine these features with the updated decision layer for processing to output the final control strategy. This allows the system to maintain the stability of core feature extraction under environmental disturbances while adapting to real-time changes through fine-tuning of the decision layer, thereby improving the accuracy and safety of the control strategy.
[0135] Furthermore, based on environmental perception data and historical decision data, the parameters to be updated are updated to obtain the updated decision layer, including: determining the loss value based on historical decision data and the loss function of the decision layer; determining the learning rate of the decision layer based on environmental perception data; and updating the parameters to be updated based on the loss value and the learning rate to obtain the updated decision layer.
[0136] The aforementioned loss function can be a mathematical metric used to quantify the deviation between the decision outputs generated by the decision-making layer in multiple adjacent control cycles prior to the current control cycle and the target desired output. For example, this loss function can be a bi-objective loss function combining task performance and decision stability. This loss function can provide guidance for adjusting decision parameters by quantifying the difference between the model's predicted results and actual control requirements or safety constraints.
[0137] The learning rate mentioned above can be a hyperparameter that controls the step size of the decision layer during backpropagation updates. It can be used to determine the magnitude of the adjustment of the decision parameters to be updated during each gradient descent, so as to prevent decision oscillations caused by excessively large updates of the parameters to be updated in the event of drastic environmental changes or abnormal inputs, or slow adaptation caused by excessively small updates of the parameters to be updated in the event of stable environments.
[0138] In an optional embodiment, considering that updating the parameters to be updated by combining the statistical characteristics of historical decision data with the dynamic characteristics of current environmental perception data, the updating process of the parameters to be updated can have both the goal-oriented nature of task performance optimization and the adaptive adjustment capability to adapt to the rate of environmental change, thereby achieving efficient convergence while ensuring the stability of decision output.
[0139] Based on this, the control system can first use historical decision data to construct and calculate the loss function of the decision layer, so as to quantify the degree of deviation between the decision results generated in the more recent decision cycles and the expected target, and determine the specific loss value.
[0140] Subsequently, the control system can determine the rate of environmental change or uncertainty index based on the environmental perception data, and dynamically determine the learning rate of the decision layer based on the index. This allows a larger learning rate to be used to accelerate parameter adjustment when the environment is relatively stable, and a smaller learning rate to suppress parameter oscillation when the environment changes drastically.
[0141] Finally, the control system can substitute the determined loss value and learning rate into the gradient descent algorithm to perform numerical update operations on the parameters to be updated, thereby obtaining the updated decision layer model.
[0142] In another alternative embodiment, the control system can also utilize the decision-making layer before the update to output a current control strategy based on decision characteristics. Subsequently, the control system can use this current control strategy and the desired value as inputs, and use a loss function to determine the loss value of the current control strategy, facilitating the updating of the parameters to be updated. After the update is completed, the current control strategy can be stored in historical decision data for subsequent parameter adjustments.
[0143] Furthermore, the method also includes: when the parameter update state is in an active state, obtaining a second duration during which the parameter update state is in an active state; updating the decision parameters based on environmental perception data and historical decision data to obtain an updated decision layer; determining a second update range for the extracted parameters based on the second duration; and updating the extracted parameters based on the second update range, environmental perception data, and historical decision data to obtain an updated extraction layer.
[0144] The aforementioned second duration can refer to the length of time that the parameter update state lasts when it is in an active state. This can be used to determine the adjustment range and intensity of the extraction parameters in the extraction layer, so as to avoid large-scale extraction parameter updates when the parameter update state has just switched from a frozen state to an active state, which would cause the control strategy to oscillate.
[0145] The aforementioned second update range may refer to the update range that is allowed to be used to update the extraction parameters, as determined by the aforementioned second duration. This second update range may be positively correlated with or have a fixed functional relationship with the duration of the aforementioned duration. This allows for the gradual relaxation of constraints on the extraction parameters when the active state lasts for a long time and the environment tends to be stable, enabling the extraction layer to be fine-tuned within a controlled range to enhance the long-term adaptability of the model.
[0146] In one alternative embodiment, considering that the activation duration of the parameter update state is dynamically monitored, and the parameters of the decision layer and the extraction layer are updated collaboratively by combining environmental perception data and historical decision data, it is possible to effectively adapt to the control requirements in complex dynamic environments, improve the output accuracy and adaptability of the control decision model in long-term operation, and thus output a better control strategy.
[0147] Based on this, the control system can first obtain the second duration of the parameter update state being in an active state to quantify the time span of current environmental changes or decision adjustments, serving as a benchmark for subsequent parameter updates. Subsequently, the control system can iteratively update the decision parameters in the decision-making layer based on environmental perception data and historical decision data, utilizing the trend information provided by historical decision data to enhance the decision-making layer's responsiveness to environmental changes, thereby obtaining an updated decision-making layer.
[0148] Next, the control system can determine a second update range for the extraction parameters based on the second duration obtained above and the dynamic characteristics of environmental changes. This second update range limits the range of extraction parameter adjustments to prevent the extraction layer from shifting drastically during the update process of the extraction parameters.
[0149] Finally, the control system can perform controlled updates to the extraction parameters in the extraction layer based on the determined second update range, environmental perception data, and historical decision data. While maintaining the stability of core feature extraction, the system can fine-tune the extraction parameters to obtain the updated extraction layer, thereby improving the quality of subsequent feature representation and completing the parameter adaptive adjustment process of the entire control decision model.
[0150] For ease of understanding, Figure 5 This is a schematic diagram of an optional parameter update process according to an embodiment of this application, such as... Figure 5 As shown, when the parameter update state is active, the control system can update both the extracted parameters and the decision parameters simultaneously. However, when the parameter update state is frozen, the control system can prevent the extraction parameters from being updated and only update the decision parameters within a limited range.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0152] According to an embodiment of this application, an embodiment of an autonomous vehicle control device is provided. It should be noted that this device can be used to execute the aforementioned autonomous vehicle control method. The specific implementation process and application scenarios are the same as those in the above embodiment, and will not be repeated here. Figure 6 This is a schematic diagram of an autonomous vehicle control device according to an embodiment of this application, such as... Figure 6 As shown, the device includes:
[0153] The data acquisition module 602 is used to acquire environmental perception data and historical decision data of autonomous vehicles.
[0154] The state determination module 604 is used to input environmental perception data and historical decision data into the control decision model, and use the environmental perception data and historical decision data to determine the parameter update state of the extraction layer. The parameter update state is used to determine whether to update the extraction parameters of the extraction layer. The control decision model consists of a decision layer and an extraction layer.
[0155] The strategy construction module 606 is used to update the state based on parameters, control the extraction layer and the decision layer to process environmental perception data and historical decision data, and obtain the control strategy of autonomous vehicle.
[0156] The vehicle control module 608 is used to control the driving of an autonomous vehicle based on a control strategy.
[0157] Furthermore, the state determination module is also used to: extract features from environmental perception data based on the extraction layer to obtain decision features; and determine the parameter update state based on environmental perception data, historical decision data, and decision features.
[0158] Furthermore, the state determination module is also used to: determine a first fluctuation state based on environmental perception data, historical decision data, and decision characteristics, wherein the first fluctuation state is used to characterize the environmental fluctuation situation in the current control cycle; determine a second fluctuation state based on historical decision data, wherein the second fluctuation state is used to characterize the environmental fluctuation situation in the historical control cycle; and determine the parameter update state based on the first fluctuation state and the second fluctuation state.
[0159] Furthermore, the state determination module is also used to: acquire a training dataset for the control decision model, wherein the training dataset includes training perception data and training decision data, wherein the data type of the training perception data is the same as the data type of the environmental perception data, and the data type of the training decision data is the same as the data type of the historical decision data; determine a statistical distribution index based on the decision features and the training dataset, wherein the statistical distribution index is used to characterize the similarity between the decision features and the training dataset; and determine the first fluctuation state based on the statistical distribution index, the environmental perception data, and the historical decision data.
[0160] Furthermore, the state determination module is also used to: determine training features based on training perception data and training decision data; and to evaluate the similarity between training features and decision features to determine statistical distribution indicators.
[0161] Furthermore, the state determination module is also used to: determine decision fluctuation indicators based on historical decision data, wherein the decision fluctuation indicators are used to characterize the stability of historical decision data; determine the environment type and environmental change rate based on environmental perception data; and construct the first fluctuation state based on statistical distribution indicators, decision fluctuation indicators, environment type, and environmental change rate.
[0162] Furthermore, the state determination module is also used to: determine a first duration for which the second fluctuation state satisfies the environmental fluctuation conditions; if the first duration is greater than a preset time and the first fluctuation state satisfies the environmental fluctuation conditions, determine the parameter update state as an active state, wherein the active state is used to characterize the state of updating the extracted parameters; if the first duration is greater than a preset time and the first fluctuation state does not satisfy the environmental fluctuation conditions, or if the first duration is less than or equal to a preset time, determine the parameter update state as a frozen state, wherein the frozen state is used to characterize the state of not updating the extracted parameters.
[0163] Furthermore, the device also includes: a condition judgment module, used to determine a rate change threshold based on the environment type; when the decision fluctuation index is less than or equal to a preset fluctuation index, and the environmental change rate is less than or equal to the rate change threshold, and the statistical distribution index is less than or equal to a preset distribution index, the first fluctuation state is determined to meet the environmental fluctuation conditions; when the decision fluctuation index is greater than the preset fluctuation index, or the environmental change rate is greater than the rate change threshold, or the statistical distribution index is greater than the preset distribution index, the first fluctuation state is determined to not meet the environmental fluctuation conditions.
[0164] Furthermore, the strategy construction module is also used to: determine the first update range of the decision parameters of the decision layer based on a preset update strategy when the parameter update state is frozen; determine the parameters to be updated from the decision parameters based on the first update range; update the parameters to be updated based on environmental perception data and historical decision data to obtain the updated decision layer; and process the decision features based on the updated decision layer to obtain the control strategy.
[0165] Furthermore, the strategy construction module is also used to: determine the loss value based on historical decision data and the loss function of the decision layer; determine the learning rate of the decision layer based on environmental perception data; and update the parameters to be updated based on the loss value and the learning rate to obtain the updated decision layer.
[0166] Furthermore, the device also includes: an update module, used to obtain a second duration during which the parameter update state is in an active state when the parameter update state is in an active state; update the decision parameters based on environmental perception data and historical decision data to obtain an updated decision layer; determine a second update range for the extracted parameters based on the second duration; and update the extracted parameters based on the second update range, environmental perception data, and historical decision data to obtain an updated extraction layer.
[0167] Embodiments of this application also provide a vehicle, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods described in various embodiments of this application when it runs.
[0168] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0169] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0170] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.
[0171] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.
[0172] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0173] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0174] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0175] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0176] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0177] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for controlling an autonomous vehicle, characterized in that, include: Acquire environmental perception data and historical decision-making data for autonomous vehicles; The environmental perception data and the historical decision data are input into the control decision model. The environmental perception data and the historical decision data are used to determine the parameter update status of the extraction layer. The parameter update status is used to determine whether to update the extraction parameters of the extraction layer. The control decision model consists of a decision layer and the extraction layer. Based on the parameter update status, the extraction layer and the decision layer are controlled to process the environmental perception data and the historical decision data to obtain the control strategy of the autonomous vehicle; Based on the control strategy, the autonomous vehicle is controlled to drive.
2. The method according to claim 1, characterized in that, Using the environmental perception data and the historical decision data, the parameter update status of the extraction layer is determined, including: Based on the extraction layer, feature extraction is performed on the environmental perception data to obtain decision features; Based on the environmental perception data, the historical decision data, and the decision characteristics, the parameter update status is determined.
3. The method according to claim 2, characterized in that, Determining the parameter update status based on the environmental perception data, the historical decision data, and the decision characteristics includes: Based on the environmental perception data, the historical decision data, and the decision characteristics, a first fluctuation state is determined, wherein the first fluctuation state is used to characterize the environmental fluctuation situation of the current control cycle. Based on the historical decision data, a second fluctuation state is determined, wherein the second fluctuation state is used to characterize the environmental fluctuation situation of the historical control cycle; The parameter update state is determined based on the first fluctuation state and the second fluctuation state.
4. The method according to claim 3, characterized in that, Based on the environmental perception data, the historical decision data, and the decision characteristics, the first fluctuation state is determined, including: Obtain the training dataset of the control decision model, wherein the training dataset includes training perception data and training decision data, wherein the data type of the training perception data is the same as the data type of the environmental perception data, and the data type of the training decision data is the same as the data type of the historical decision data; Based on the decision features and the training dataset, a statistical distribution index is determined, wherein the statistical distribution index is used to characterize the similarity between the decision features and the training dataset; The first fluctuation state is determined based on the statistical distribution index, the environmental perception data, and the historical decision data.
5. The method according to claim 4, characterized in that, Based on the decision features and the training dataset, statistical distribution indicators are determined, including: Based on the training perception data and the training decision data, training features are determined; The similarity between the training features and the decision features is evaluated to determine the statistical distribution index.
6. The method according to claim 4, characterized in that, Determining the first fluctuation state based on the statistical distribution index, the environmental perception data, and the historical decision data includes: Based on the historical decision-making data, a decision fluctuation index is determined, wherein the decision fluctuation index is used to characterize the stability of the historical decision-making data; Based on the environmental perception data, the environmental type and rate of environmental change are determined. The first fluctuation state is constructed based on the statistical distribution index, the decision fluctuation index, the environment type, and the environment change rate.
7. The method according to claim 3, characterized in that, Determining the parameter update state based on the first fluctuation state and the second fluctuation state includes: Determine the first duration during which the second fluctuation state satisfies the environmental fluctuation conditions; If the first duration is greater than a preset time and the first fluctuation state meets the environmental fluctuation conditions, the parameter update state is determined to be an active state, wherein the active state is used to characterize the state of updating the extracted parameters; If the first duration is greater than the preset time and the first fluctuation state does not meet the environmental fluctuation conditions, or if the first duration is less than or equal to the preset time, the parameter update state is determined to be a frozen state, wherein the frozen state is used to characterize the state in which the extracted parameters are not updated.
8. The method according to claim 7, characterized in that, The method further includes: Determine the rate change threshold based on the environment type; If the decision fluctuation index is less than or equal to the preset fluctuation index, the rate of environmental change is less than or equal to the rate of change threshold, and the statistical distribution index is less than or equal to the preset distribution index, then the first fluctuation state is determined to meet the environmental fluctuation conditions. If the decision fluctuation index is greater than the preset fluctuation index, or the environmental change rate is greater than the rate change threshold, or the statistical distribution index is greater than the preset distribution index, then the first fluctuation state is determined not to meet the environmental fluctuation conditions.
9. The method according to any one of claims 1 to 8, characterized in that, Based on the parameter update state, the extraction layer and the decision layer are controlled to process the environmental perception data and the historical decision data to obtain the control strategy of the autonomous vehicle, including: When the parameter update status is frozen, the first update range of the decision parameters of the decision layer is determined based on a preset update strategy. Based on the first update range, determine the parameters to be updated from the decision parameters; Based on the environmental perception data and the historical decision data, the parameters to be updated are updated to obtain the updated decision layer; Based on the updated decision layer, the decision features are processed to obtain the control strategy.
10. The method according to claim 9, characterized in that, Based on the environmental perception data and the historical decision data, the parameters to be updated are updated to obtain the updated decision layer, including: The loss value is determined based on the historical decision data and the loss function of the decision layer; Based on the environmental perception data, the learning rate of the decision layer is determined; Based on the loss value and the learning rate, the parameters to be updated are updated to obtain the updated decision layer.
11. The method according to claim 9, characterized in that, The method further includes: When the parameter update state is in an active state, obtain the second duration during which the parameter update state is in the active state; Based on the environmental perception data and the historical decision data, the decision parameters are updated to obtain the updated decision layer; Based on the second duration, a second update range for the extracted parameters is determined; Based on the second update range, the environmental perception data, and the historical decision data, the extraction parameters are updated to obtain the updated extraction layer.
12. A vehicle, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the method according to any one of claims 1 to 11.