Automatic driving behavior decision-making system and method based on visual language model
By introducing a decision-making system based on visual language model in the autonomous driving system, fine-grained comprehensive rewards are generated and strategy training is carried out, the problem that reward design in the existing autonomous driving system depends on manual rules and static reward functions, which significantly improves the intelligence, robustness and safety of the system.
Patent Information
- Application Number
- CN202510615433.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing autonomous driving system, reinforcement learning methods rely on static and preset reward functions, and it is difficult to flexibly deal with complex and changing road environments and dynamic driving scenarios, resulting in significant shortcomings in generalization and adaptability of decision-making and control strategies.
The autonomous driving behavior decision system based on visual language model is adopted to obtain real-time environmental data through the data processing module, extract multiple deep semantic information, generate comparative semantic target rewards, and fuse it with low-dimensional vehicle state data through the reward synthesis module to generate fine-grained comprehensive rewards. At the same time, replay buffering technology and batch processing mechanism are used for training management, and maximum entropy reinforcement learning algorithm is used for strategy training.
It significantly improves the intelligence and adaptability of autonomous driving decisions, can respond more flexibly to complex and changeable driving scenarios, improves the robustness and safety of the system, accelerates the convergence speed of strategy training, and improves the safety and comfort of driving.
Smart Images

Figure CN120123997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of joint control of vehicle subsystems, and particularly to an autonomous driving behavior decision-making system and method based on a vision-language model. Background Art
[0002] With the rapid development of autonomous driving technology, reinforcement learning, as an effective machine learning method, has been widely applied to the training of driving strategies in autonomous driving systems. However, in the prior art, when training driving strategies using reinforcement learning, it usually relies on manually designed reward functions. The design of traditional reward functions usually requires weighted summation by combining multiple driving-related indicators such as vehicle speed, lane departure, and collision status to guide the agent to learn strategies. But this method has many problems. On the one hand, the debugging process of manually designing reward functions is usually time-consuming and laborious, and due to the complex and changeable driving scenarios, the designed reward functions often have insufficient generalization ability and are difficult to adapt to multiple scenarios. On the other hand, a single manually designed reward function is difficult to fully capture the implicit semantic information in driving tasks and is difficult to achieve a balance between multi-task objectives, resulting in limited performance of driving strategies.
[0003] In recent years, with the development of deep learning technology, pre-trained vision-language models (such as CLIP) have demonstrated excellent capabilities in multi-modal semantic understanding of images and natural languages; this makes it possible to generate flexible and accurate reward signals through semantic understanding. However, in autonomous driving scenarios, how to effectively apply vision-language models to real-time generate reward signals with high recognition and robustness still faces many technical challenges; this limits the practical application of reward functions based on semantic understanding in autonomous driving.
[0004] In order to improve the performance of autonomous driving control systems, the following existing similar patent technologies are as follows: (1) Chinese Patent Publication No. CN119329519A discloses an optimized method for vehicle adaptive cruise control based on SAC for the problem of adaptive cruise control under urban conditions. This patent has carried out fine optimization in the field of cruise control and mainly focuses on low-level control tasks such as following vehicles, accelerating, and decelerating. However, its reward function still relies on dynamic adjustment based on physical quantities and empirical formulas (such as safety, efficiency, and comfort rewards), and manual parameter adjustment is required for specific application scenarios, and the applicable range is relatively limited. In addition, its technical design is mainly oriented to specific task scenarios and is difficult to extend to more complex multi-task driving environments.
[0005] (2) Chinese Patent CN119283855A discloses an autonomous driving lane - keeping decision - making method based on deep reinforcement learning. Through techniques such as state fusion, composite reward function, and prioritized experience replay, this method realizes deep reinforcement learning control in the lane - keeping task and verifies its safety and efficiency in a simulation environment. However, the reward function designed in this patent mainly relies on fixed manually - designed reward and punishment items, and its structure is relatively static and fixed, making it difficult to flexibly capture the semantic information hidden in complex and changeable driving scenarios. In addition, in data processing, this patent mainly depends on the traditional fusion method of images and low - dimensional motion data, ignoring the potential of large - scale pre - trained models in extracting deep - level semantic information from images.
[0006] In summary, the reinforcement learning methods in current autonomous driving systems mainly rely on static and preset reward functions and experience rules. This approach is difficult to flexibly handle complex and changeable road environments and dynamic driving scenarios, resulting in significant deficiencies in the generalization and adaptability of autonomous driving decision - making and control strategies, and it still cannot fully meet the requirements of autonomous driving technology in practical applications. Therefore, it is urgent to develop new technical methods to address these challenges. Summary of the Invention
[0007] The purpose of the present invention is to provide an autonomous driving behavior decision - making system and method based on a vision - language model, so as to solve all or one of the above problems existing in the prior art.
[0008] To solve the above - mentioned technical problems, the specific technical solutions of the present invention are as follows: On the one hand, the present invention provides an autonomous driving behavior decision - making system based on a vision - language model, including: A data processing module, which is used to: obtain real - time environmental data about the vehicle, pre - process the real - time environmental data, and extract bidirectional language objectives and multi - dimensional deep semantic information of the pre - processed data based on the vision - language model. The multi - dimensional deep semantic information is semantic information about visual features and text features; A reward generation module, which is used to: define the bidirectional language objectives, calculate the similarity between visual features and text features based on the multi - dimensional deep semantic information, and generate a contrast semantic objective reward based on the similarity calculation; A reward synthesis module, which is used to: perform normalization processing on the contrast semantic objective reward, fuse and calculate the normalized reward with low - dimensional vehicle state data to obtain a fine - grained comprehensive reward; A training management module, which is used to: store the real - time state data during training by using the replay buffer technology, perform unified calculation of the fine - grained comprehensive reward by using a batch - processing mechanism, and perform autonomous driving policy training based on the maximum entropy reinforcement learning algorithm after the unified calculation; A decision control module, configured to: deploy the trained autonomous driving policy network to a vehicle, input the real-time state of the vehicle into the autonomous driving policy network, and control the vehicle according to the optimal action output by the autonomous driving policy network.
[0009] As an improved solution, the data processing module includes: a data preprocessing unit; The data preprocessing unit is configured to: perform size and image enhancement processing on the front-view camera data in the real-time environmental data; the image enhancement processing includes denoising processing, histogram equalization processing, and color normalization processing.
[0010] As an improved solution, the data processing module further includes: an inference unit; The inference unit is configured to: perform feature extraction and semantic projection processing on the preprocessed data and the bidirectional language target using the vision-language model; The feature extraction operation performed by the inference unit includes: the inference unit uses the vision encoder of the vision-language model to extract the image features of the preprocessed image; The semantic projection processing performed by the inference unit includes: the inference unit uses the vision encoder of the vision-language model to project the extracted image features into a shared semantic space; the inference unit uses the language encoder of the vision-language model to encode the bidirectional language target and project it into the shared semantic space.
[0011] As an improved solution, the bidirectional language target includes: A positive language target used to reflect the consistency between the current state of the vehicle and the ideal driving state; And, A negative language target used to reflect the degree of proximity between the current state of the vehicle and the non-ideal driving state.
[0012] As an improved solution, the reward generation module includes: a semantic reward calculation unit; The semantic reward calculation unit is configured to: use the vision encoder of the vision-language model to obtain a state embedding vector regarding visual features, calculate a first similarity between the state embedding vector and the positive language target, calculate a second similarity between the state embedding vector and the negative language target, and use the weighted difference result of the first similarity and the second similarity as the contrast semantic target reward.
[0013] As an improved solution, the reward synthesis module includes: a reward normalization unit, a state information fusion unit, and a comprehensive reward calculation unit; The reward normalization unit is used to normalize the contrast semantic target reward to obtain the normalized reward; The state information fusion unit is used to obtain the low-dimensional vehicle state data of the vehicle and calculate the state factor corresponding to the low-dimensional vehicle state data; The comprehensive reward calculation unit is used to fuse the state factors by product calculation to obtain a preliminary comprehensive reward; the comprehensive reward calculation unit determines the fine-grained comprehensive reward according to the sparse reward and the preliminary comprehensive reward.
[0014] As an improved scheme, the low-dimensional vehicle state data includes: vehicle speed, lane center deviation, matching degree between the vehicle and the road direction angle, and driving stability index; The state factors include: vehicle speed factor, lane deviation factor, direction consistency factor, and driving stability factor.
[0015] As an improved scheme, the training management module includes: an experience storage unit and a batch reward calculation unit; The experience storage unit is used to store the real-time state data generated by the vehicle automatic driving system based on the automatic driving policy network in the replay buffer in the form of transition tuples. The real-time state data includes: state, action, reward, and the state at the next moment; The batch reward calculation unit is used to randomly sample part of the state data from the replay buffer at a fixed time interval and uniformly calculate the contrast semantic target reward and the preliminary comprehensive reward corresponding to the part of the state data through the vision-language model.
[0016] As an improved scheme, the training management module further includes: a SAC training unit; The SAC training unit is used to perform automatic driving policy training using the SAC algorithm; The automatic driving policy training executed by the SAC training unit includes: Critic network update and Actor network update; The SAC training unit is further used to update the Critic network parameters by minimizing the mean square Bellman residual during the update process of the Critic network; The SAC training unit is further used to optimize the Actor network by maximizing the weighted sum of the reward and the entropy during the update process of the Actor network, and optimize the Actor network using the reparameterization technique.
[0017] On the other hand, the present invention also provides an automatic driving behavior decision method based on a vision-language model, including the following steps: Data processing steps based on a vision - language model: Obtain real - time environmental data about the vehicle, pre - process the real - time environmental data, and extract bidirectional language targets and multi - dimensional deep semantic information of the pre - processed data based on the vision - language model. The multi - dimensional deep semantic information is semantic information about visual features and text features; Reward generation steps based on contrastive language targets: Define the bidirectional language targets, calculate the similarity between visual features and text features based on the multi - dimensional deep semantic information, and generate a contrastive semantic target reward based on the similarity calculation; Hierarchical reward synthesis steps: Normalize the contrastive semantic target reward, fuse and calculate the normalized reward with low - dimensional vehicle state data to obtain a fine - grained comprehensive reward; Reinforcement learning training management steps: Adopt the replay buffer technology to store the real - time state data during training, adopt a batch - processing mechanism to uniformly calculate the fine - grained comprehensive reward, and perform autonomous driving policy training based on the maximum entropy reinforcement learning algorithm after the unified calculation; Real - time decision - making and vehicle control steps: Deploy the trained autonomous driving policy network to the vehicle, input the real - time state of the vehicle into the autonomous driving policy network, and control the vehicle according to the optimal action output by the autonomous driving policy network.
[0018] In view of the problems in the existing autonomous driving technology, such as the reward design relying on manual rules, single state representation, large computational resource overhead, insufficient real - time response ability, and non - intelligent multi - modal information fusion, the present invention proposes an innovative solution, which has the following remarkable beneficial effects: 1. By introducing a pre - trained vision - language model, the present invention automatically generates semantic reward signals, getting rid of the complex process of manually designing reward functions in traditional methods; it can automatically extract deep semantic information from images and adjust the reward signal in real time by combining positive and negative language targets, thereby dynamically reflecting the changes in the driving state, and then significantly improving the intelligence and adaptability of autonomous driving decisions, enabling the system to more flexibly cope with complex and changeable driving scenarios.
[0019] 2. By fusing high - dimensional visual information and low - dimensional vehicle state data, the present invention realizes a comprehensive description of the driving environment; this multi - modal data fusion method provides richer context information for decision - making, can more accurately capture environmental changes, and significantly improves the robustness and safety of the autonomous driving system in complex scenarios.
[0020] 3. The present invention proposes a hierarchical reward synthesis mechanism. By normalizing the automatically generated semantic rewards and fusing them with real-time vehicle state data, a fine-grained comprehensive reward signal is generated. This dynamic reward design can flexibly adjust the weights of various rewards, accurately reflect the vehicle's performance in different driving scenarios, thereby accelerating the convergence speed of policy training and enhancing the safety and comfort of driving.
[0021] 4. The present invention introduces a batch reward calculation mechanism. By batch sampling state data from the replay buffer and uniformly calculating semantic rewards, and using an asynchronous update strategy to feedback the calculation results into the experience tuple. This mechanism significantly reduces the real-time calculation burden, improves the training efficiency and the system's response speed, enabling the autonomous driving system to still maintain low-latency operation in a high-load environment, and further enhancing the practicality of the system.
[0022] 5. By closely combining visual semantics with reinforcement learning, the present invention realizes an end-to-end autonomous driving control system, which can respond in real time and make safe decisions in complex driving environments. Compared with traditional solutions, the present invention significantly improves the safety, robustness, and generalization ability of the autonomous driving system, providing solid technical support for the popularization of autonomous driving technology in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic diagram of the logical architecture of the autonomous driving behavior decision-making system based on the vision-language model described in Embodiment 1 of the present invention; Figure 2 It is a schematic diagram of the flow of the autonomous driving behavior decision-making method based on the vision-language model described in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following elaborates on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0026] In the description of the present invention, it should be noted that the embodiments described are some embodiments of the present invention, rather than all embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0027] In the description, claims and above-mentioned drawings of this article, terms such as "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this article described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0028] Embodiment 1. This embodiment provides an autonomous driving behavior decision-making system based on a vision language model, as Figure 1 shown, including: (1) A data processing module based on a vision language model, including: (1.1) A data acquisition unit, which is used to acquire real-time environmental data under the call of the vehicle; Among them, the data acquisition unit includes: a front camera, a lidar, a GPS, and an IMU; Among them, the acquired real-time environmental data includes: the front road scene, traffic signs, lane line information, the situation of surrounding pedestrians, and the situation of surrounding vehicles; Among them, the input data for the vision language model (VLM) is RGB camera data.
[0029] (1.2) A data preprocessing unit, which is used to adjust the size of the original image collected by the front camera and perform enhancement processing on the image; Among them, the size adjustment includes, but is not limited to: adjusting to 224×224 pixels; Among them, the enhancement processing includes, but is not limited to: denoising and histogram equalization processing for optimizing visual feature extraction, and color normalization processing for matching the input format of the pre-trained VLM.
[0030] (1.3) An inference unit, which is used to perform feature extraction and semantic-based projection processing using a mature vision language model (including but not limited to CLIP); Among them, the feature extraction includes, but is not limited to, using the visual encoder VLM of CLIP I to extract the image features of the preprocessed image; Among them, the semantic-based projection processing includes, but is not limited to, using the visual encoder VLM of CLIP IProject the extracted image features into the shared semantic space and use the language encoder VLM of CLIP L Encode the predefined text features (i.e., positive language targets such as "smooth driving" and negative language targets such as "collision occurred") respectively and project them into the same semantic space (i.e., the shared semantic space); Among them, the extracted visual features and the predefined text features are used to calculate the similarity in the same embedding space in subsequent operations, providing input for subsequent reward generation.
[0031] (2) The reward generation module based on CLG, including: (2.1) Language target preset unit, used to set a group of positive language targets during system initialization and negative language targets ; Among them, the positive language target , used to reflect the consistency between the current state and the ideal driving state (such as "the vehicle stays in the lane and drives smoothly"); Among them, the negative language target , used to reflect the degree of proximity between the current state and the non-ideal driving state (such as "the vehicle has collided" or "departed from the lane"); Among them, to ensure the consistency of reward calculation, and remain unchanged throughout the training process.
[0032] (2.2) Semantic reward calculation unit, used to calculate the similarity between the visual features of the current vehicle state s and the positive language target and the negative language target respectively, and determine the semantic reward output based on the similarity calculation results; Among them, the semantic reward calculation unit uses the visual encoder VLM I to obtain the state embedding vector ; Among them, the semantic reward calculation unit calculates and the first similarity ; Among them, the semantic reward calculation unit calculates and the second similarity ; Among them, the semantic reward calculation unit weights and subtracts and and takes the obtained difference as the preliminary semantic reward output , based on It can help the system intuitively reflect whether the current driving state is more inclined to the ideal state or the bad state, and its calculation formula is as follows: ; Among them, represents the cosine similarity, and α and β are weight factors respectively. Usually, α = β = 0.5.
[0033] (3) Hierarchical reward synthesis module, including: (3.1) Reward normalization unit, which is used to normalize the preliminary semantic reward output obtained from the above operations to ensure that the reward value is within a stable and reasonable range, and avoid the problem that the numerical range of the CLG reward depends on the cosine similarity calculated by the VLM. Its normalization formula is as follows: ; Among them, the function is used to limit within the interval to prevent outliers from interfering. And "x" in this function is " " in the above normalization formula
[0034] (3.2) State information fusion unit, which is used to simultaneously obtain several state information of the vehicle from on-vehicle sensors and calculate the corresponding factors respectively; Among them, several state information includes, but is not limited to: vehicle speed, lane center deviation, matching degree between the vehicle and the road direction angle, and driving stability index; Among them, the corresponding factors are as follows: Vehicle speed factor, , and , ; The formula can show that the CLG semantic reward determines the target speed, and this vehicle speed factor reflects the matching degree between the current vehicle speed and the target vehicle speed; Lane deviation factor, , which is used to reflect the deviation between the vehicle and the lane center; Direction consistency factor, , which is used to measure the consistency between the vehicle driving direction and the road direction; Driving stability factor, , which is used to evaluate the stability of the vehicle during driving.
[0035] (3.3) Comprehensive reward calculation unit, which is used to fuse the above factors in the form of a product to obtain a preliminary comprehensive reward , and finally fuse the sparse reward of the task itself and the preliminary comprehensive reward in proportion to obtain the final reward (i.e., the fine-grained comprehensive reward); Among them, the calculation formula for the preliminary comprehensive reward is as follows: ; Among them, the calculation formula for the final reward is as follows: ; Among them, is a weight parameter, and > 0. This weight parameter is used to balance the relative importance between the comprehensive reward and the sparse task reward. Additionally, it should be noted that for , in autonomous driving or other reinforcement learning tasks, positive rewards are only given when specific key events (such as task success, reaching the destination, or completing the goal) occur, and most of the time it belongs to a reward signal with a reward value of 0.
[0036] (4) Training management module, including: (4.1) Experience storage unit, which is used to store the state data generated by the autonomous driving system based on the policy at each time t using a replay buffer; Among them, the state data includes: the vehicle state at time t , the vehicle action at time t , the final reward and the predicted vehicle state at time t + 1 ; Among them, the above state data is stored in the form of a transition tuple ( ) and is used in subsequent training; It should be noted that the Actor network is a neural network structure, where π represents the policy, which is a function that describes the rules for what actions an agent should take given a state; a represents the action, that is, the specific behavior that an agent can perform in the environment. For example, in a vehicle control task, the action may be the steering angle of the vehicle; s represents the state, which is a description of the environment at a certain moment and contains all the information that the agent can observe. For example, in an autonomous driving scenario, the state may include the vehicle's speed, position, distance to surrounding obstacles, etc. Generally speaking, in the Actor network, its input is the state s of the environment, and the output is the policy π(a|s) corresponding to this state, that is, the probability distribution of each possible action given the state s.
[0037] (4.2) Batch reward calculation unit, which is used to calculate the semantic reward and the comprehensive reward using a batch processing mechanism; Among them, the batch processing mechanism is as follows: within a fixed time interval, a batch of state data is randomly sampled from the replay buffer, and the semantic reward corresponding to the current state is uniformly calculated through the vision - language model and comprehensive rewards , and update the corresponding value in the buffer, thereby alleviating the high computational load brought by using the pre-trained vision-language model in real-time computing, reducing resource occupancy, and achieving "asynchronous" calculation of rewards without disturbing the main training loop.
[0038] (4.3) SAC training unit, which is used to perform policy training using the Soft Actor-Critic (SAC) algorithm in response to the completion of batch reward calculation; Among them, during the SAC training process, it mainly includes: updating the Critic network and updating the Actor network; specifically, the Critic network is used to estimate the Q value, that is, the long-term return expectation value of a certain state-action, and is used to evaluate the quality of the action selected by the Actor network; the Actor network is used to optimize the policy , that is, the probability distribution of selecting the optimal action under a given state, and maximizing the long-term return; Among them, during the SAC training process, the main optimization objectives are as follows: (i) Critic network update objective: Update the Critic network parameters by minimizing the mean square Bellman residual (Soft Bellman Residual, SBR), and its loss function is specifically as follows: ; Among them, is the final reward after batch reward calculation, is the Q value estimated by the Critic network, is the target Q value; Among them, The calculation formula of is: Among them, is the discount factor, which is used to control the influence degree of future rewards; Adopt a double Q network to further reduce the problem of overestimation of Q values; is the entropy regularization term, which is used to encourage policy exploration.
[0039] (ii) Actor network update objective: 1) Optimize the Actor network by maximizing the weighted sum of rewards and entropy, and its loss function is specifically as follows: ; Among them, represents that the policy tends to select the action with the largest Q value; is the entropy reward term, which is used to control the exploration of the policy and prevent the policy from converging to the local optimum prematurely.
[0040] 2) The Actor network is optimized using the Reparameterization Trick, and the specific formula is as follows: ; Among them, is the action mean output by the policy network; is the standard deviation output by the policy network; is the standard normal noise, which is used to introduce randomness.
[0041] (5) The decision control module is used to deploy the reinforcement learning model to the vehicle's autonomous driving system after training: After the model is deployed, the current state of the vehicle is input into the model, and the optimal action is output by the policy network of the model (including but not limited to steering wheel angle, acceleration, braking force, etc.); then the vehicle calls the PID controller or MPC (Model Predictive Control) to adjust the action according to to ensure the smooth execution of the control signal; finally, the vehicle executes the action , completes the execution of the autonomous driving decision at the current time t, and enters the next state , realizing the closed-loop of real-time perception → decision → control.
[0042] In summary, this solution uses a pre-trained vision-language model to automatically extract deep semantic information from vehicle camera images, combines positive and negative language objectives to generate contrastive language objective rewards, breaks away from the limitations of traditional manually designed reward functions, and improves the dynamic adaptability of the reward signal; by normalizing the semantic reward and fusing it with low-dimensional vehicle state data, a fine-grained comprehensive reward signal is constructed to achieve multi-index collaborative feedback in complex scenarios; at the same time, a batch processing strategy is adopted to batch sample state data from the replay buffer to uniformly calculate the semantic reward and asynchronously update the experience tuple, improving the sample utilization rate, accelerating training and reducing the real-time calculation burden; finally, this solution forms a new reward design method and training framework (UVL-RLAD: Unified Visual-Language-driven Reinforcement Learning for Autonomous Driving), which improves the intelligence, stability and generalization ability of the reinforcement learning strategy in the autonomous driving system.
[0043] It should be noted that the above examples are only for explaining the present invention and should not limit the protection scope of the present invention accordingly.
[0044] Embodiment 2. This embodiment is based on the same inventive concept as the autonomous driving behavior decision-making system based on a vision-language model described in Embodiment 1, and provides an autonomous driving behavior decision-making method based on a vision-language model, as Figure 2 shown, which includes the following steps: S100. Data processing step based on a vision-language model, specifically including: Obtain real-time environmental data about the vehicle, preprocess the real-time environmental data, and extract bidirectional language targets and multi-dimensional deep semantic information of the preprocessed data based on the vision-language model. The multi-dimensional deep semantic information is semantic information about visual features and text features; S200. Reward generation step based on contrastive language targets, specifically including: Define the bidirectional language targets, calculate the similarity between visual features and text features based on the multi-dimensional deep semantic information, and generate a contrastive semantic target reward based on the similarity calculation; S300. Hierarchical reward synthesis step, specifically including: Normalize the contrastive semantic target reward, and perform fusion calculation on the normalized reward and low-dimensional vehicle state data to obtain a fine-grained comprehensive reward; S400. Reinforcement learning training management step, specifically including: Adopt a replay buffer technique to store real-time state data during training, adopt a batch processing mechanism to uniformly calculate the fine-grained comprehensive reward, and perform autonomous driving policy training based on the maximum entropy reinforcement learning algorithm after the unified calculation; S500. Real-time decision-making and vehicle control step, specifically including: Deploy the trained autonomous driving policy network to the vehicle, input the real-time state of the vehicle into the autonomous driving policy network, and control the vehicle according to the optimal action output by the autonomous driving policy network.
[0045] Different from the prior art, by using the autonomous driving behavior decision-making system and method based on a vision-language model of the present application, a semantic reward signal can be automatically generated by pre-training the vision-language model, hierarchical reward synthesis can be performed in combination with vehicle state information to achieve dynamic and fine-grained reward design; at the same time, a batch processing mechanism is introduced to optimize the calculation process, reduce the real-time calculation burden, and improve the training efficiency; by fusing high-dimensional vision and low-dimensional vehicle state data, the multi-modal information expression is enriched, and the safety, robustness, and generalization ability of the autonomous driving system are significantly improved, realizing end-to-end efficient control of autonomous driving.
[0046] It should be understood that in various embodiments herein, the sequence numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments herein.
[0047] It should also be understood that in the embodiments herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0048] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.
[0049] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0050] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be in the form of electrical, mechanical, or other connections.
[0051] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments herein.
[0052] In addition, each functional unit in the various embodiments of this document may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0053] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution in this document, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this document. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0054] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present invention.
Claims
1. An autonomous driving behavior decision system based on a visual language model, characterized in that: include: A data processing module, used to: obtain real-time environmental data about the vehicle, preprocess the real-time environmental data, extract bidirectional language targets and multi-dimensional deep semantic information of the preprocessed data based on a visual language model, wherein the multi-dimensional deep semantic information is semantic information about visual features and text features; A reward generation module, used to: define the bidirectional language target, and calculate the similarity between visual features and text features based on the multi-dimensional deep semantic information, and generate a comparison semantic target reward based on the similarity calculation; A reward synthesis module is used to: normalize the contrast semantic target reward, and fuse the normalized reward with the low-dimensional vehicle state data to obtain a fine-grained comprehensive reward; A training management module, for: using a replay buffering technique to store real-time status data in training, using a batch processing mechanism to perform unified calculation of the fine-grained comprehensive rewards, and performing autonomous driving strategy training based on a maximum entropy reinforcement learning algorithm after the unified calculation; The decision control module is used to: deploy the trained autonomous driving strategy network on the vehicle, input the real-time status of the vehicle into the autonomous driving strategy network, and control the vehicle according to the optimal action output by the autonomous driving strategy network.
2. The autonomous driving behavior decision system based on the visual language model according to claim 1, characterized in that: The data processing module includes: a data preprocessing unit; The data preprocessing unit is used to: resize and perform image enhancement processing on the forward-looking camera data in the real-time environmental data; the image enhancement processing includes denoising processing, histogram equalization processing and color normalization processing.
3. The autonomous driving behavior decision system based on the visual language model according to claim 2, characterized in that: The data processing module further includes: an inference unit; The reasoning unit is used to: use the visual language model to perform feature extraction and semantic projection processing on the preprocessed data and the bidirectional language target; The feature extraction operation performed by the inference unit includes: the inference unit extracting image features of the preprocessed image using a visual encoder of the visual language model; The semantic projection processing performed by the inference unit includes: the inference unit uses the visual encoder of the visual language model to project the extracted image features into a shared semantic space; the inference unit uses the language encoder of the visual language model to encode the bidirectional language target and project it into the shared semantic space.
4. The autonomous driving behavior decision system based on a visual language model according to claim 1, characterized in that: The two-way language objectives include: Positive language goals used to reflect the consistency between the vehicle's current state and the ideal driving state; as well as, A negative language target used to reflect the degree of proximity between the vehicle's current state and a non-ideal driving state.
5. The autonomous driving behavior decision system based on a visual language model according to claim 4, characterized in that: The reward generation module includes: a semantic reward calculation unit; The semantic reward calculation unit is used to: use the visual encoder of the visual language model to obtain a state embedding vector about the visual feature, calculate a first similarity between the state embedding vector and the positive language target, calculate a second similarity between the state embedding vector and the negative language target, and use a weighted difference result of the first similarity and the second similarity as the comparison semantic target reward.
6. The autonomous driving behavior decision system based on a visual language model according to claim 1, characterized in that: The reward synthesis module includes: a reward normalization unit, a state information fusion unit and a comprehensive reward calculation unit; The reward normalization unit is used to normalize the contrast semantic target reward to obtain a normalized reward; The state information fusion unit is used to obtain the low-dimensional vehicle state data of the vehicle and calculate the state factor corresponding to the low-dimensional vehicle state data; The comprehensive reward calculation unit is used to merge the state factors in a multiplication calculation manner to obtain a preliminary comprehensive reward; the comprehensive reward calculation unit determines the fine-grained comprehensive reward based on the sparse reward and the preliminary comprehensive reward.
7. The autonomous driving behavior decision system based on a visual language model according to claim 6, characterized in that: The low-dimensional vehicle state data includes: vehicle speed, lane center deviation, matching degree between vehicle and road direction angle, and driving stability index; The state factors include: vehicle speed factor, lane departure factor, direction consistency factor and driving stability factor.
8. The autonomous driving behavior decision system based on a visual language model according to claim 1, characterized in that: The training management module includes: an experience storage unit and a batch reward calculation unit; The experience storage unit is used to store the real-time state data generated by the vehicle automatic driving system based on the automatic driving strategy network in the form of conversion tuples using a replay buffer, wherein the real-time state data includes: state, action, reward, and next moment state; The batch reward calculation unit is used to randomly sample part of the state data from the replay buffer within a fixed time interval, and uniformly calculate the comparative semantic target rewards and preliminary comprehensive rewards corresponding to the part of the state data respectively through the visual language model.
9. The autonomous driving behavior decision system based on a visual language model according to claim 1, characterized in that: The training management module further includes: a SAC training unit; The SAC training unit is used to perform autonomous driving strategy training using the SAC algorithm; The autonomous driving strategy training process performed by the SAC training unit includes: updating the Critic network and updating the Actor network; The SAC training unit is further used to update the critic network parameters by minimizing the mean square Bellman residual during the critic network update process; The SAC training unit is also used to optimize the Actor network by maximizing the weighted sum of reward and entropy during the Actor network update process, and to optimize the Actor network using a reparameterization technique.
10. A method for autonomous driving behavior decision-making based on a visual language model, characterized in that: The following steps are involved: Data processing steps based on visual language model: Acquire real-time environmental data about the vehicle, preprocess the real-time environmental data, extract bidirectional language targets and multi-dimensional deep semantic information of the preprocessed data based on a visual language model, wherein the multi-dimensional deep semantic information is semantic information about visual features and text features; Steps for reward generation based on contrastive language objectives: defining the bidirectional language target, and performing similarity calculation between visual features and text features based on the multi-dimensional deep semantic information, and generating a comparative semantic target reward based on the similarity calculation; Steps for synthesizing hierarchical rewards: Normalizing the contrast semantic target reward, and fusing the normalized reward with the low-dimensional vehicle state data to obtain a fine-grained comprehensive reward; Reinforcement learning training management steps: A replay buffer technology is used to store the real-time state data in training, a batch processing mechanism is used to uniformly calculate the fine-grained comprehensive rewards, and after the uniform calculation, the autonomous driving strategy training is performed based on the maximum entropy reinforcement learning algorithm; Real-time decision making and vehicle control steps: The trained autonomous driving strategy network is deployed on the vehicle, the real-time state of the vehicle is input into the autonomous driving strategy network, and the vehicle is controlled according to the optimal action output by the autonomous driving strategy network.
Citation Information
Patent Citations
Automatic driving lane keeping decision-making method based on deep reinforcement learning
CN119283855A
SAC-based automobile adaptive cruise control optimization method
CN119329519A
Image processing method and device, computer equipment and image processing system
CN117727086A
Automatic driving decision-making system based on deep reinforcement learning
CN117962926A
Cited By
Safe automatic driving method based on reward mechanism
CN120526407A
Unmanned vehicle control system and method based on multi-agent cooperation and computer equipment
CN121325882A
Multi-agent cooperative unmanned vehicle control system, method, and computer equipment
CN121325882B