Automatic driving continuous optimization method and system based on latent state consistency modeling and storage medium

By using a latent state consistency modeling approach, autonomous driving systems can extract environmental evolution patterns and construct stable temporal structure features in dynamic traffic environments, enabling cross-driving task knowledge transfer and optimization. This solves the adaptability and robustness issues of existing systems in complex environments and enhances the decision-making and control capabilities of autonomous driving systems.

CN121543695BActive Publication Date: 2026-04-07TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing autonomous driving systems lack the ability to uniformly characterize the changing patterns of the environment over time in dynamic traffic environments. This makes it difficult for models to form a stable and reliable temporal structure, leading to catastrophic forgetting phenomena and affecting the continuity and robustness of vehicle behavior decision-making and control.

Method used

A latent state consistency modeling approach is adopted, which extracts stable and dynamic knowledge domains through feature encoding networks, constructs a latent state transition network to capture the evolution of the environment, and realizes cross-task knowledge transfer and optimization through a policy decision network, thereby improving the system's adaptability and robustness in complex traffic environments.

Benefits of technology

Without relying on full retraining, the autonomous driving system can efficiently adapt to new tasks, maintain stable cognition and knowledge reuse of historical tasks, and significantly improve its adaptability, scalability and decision-making and control robustness in complex traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543695B_ABST
    Figure CN121543695B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of intelligent driving decision control, specifically a method, system, and storage medium for continuous optimization of autonomous driving based on latent state consistency modeling. It addresses issues such as insufficient adaptability of existing autonomous driving systems in complex, multi-scenario, and dynamic traffic environments, decreased decision control stability during long-term operation, and inadequate characterization of environmental changes. This solution learns by dividing the dynamic traffic environment state information into knowledge domains based on structural stability. This allows the latent space to extract stable, predictable, and task-independent environmental state features, enabling cross-driving task knowledge transfer and optimization, and improving the robustness and long-term reliability of autonomous driving systems in real dynamic traffic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent driving decision control, and in particular to a method, system and storage medium for continuous optimization of autonomous driving based on potential state consistency modeling. It is used to realize cross-driving task knowledge transfer and optimization in dynamic traffic environments, and improve the adaptability and robustness of autonomous driving decision control system in complex long-tail scenarios. Background Technology

[0002] Autonomous driving, as an important direction for the deep integration of artificial intelligence and transportation, has made significant progress in environmental perception, decision-making and control, and path planning in recent years. With the development of large-scale driving data and deep learning methods, autonomous driving systems are gradually gaining the ability to perform complex driving behaviors in structured scenarios, demonstrating higher flexibility and efficiency compared to traditional rule-based methods.

[0003] However, in real-world road environments, the behavior of traffic participants, traffic flow structure, road conditions, and weather factors all change continuously over time. Autonomous driving systems need to constantly adapt to these dynamically evolving environments during long-term operation. Most existing autonomous driving decision-making and control methods rely primarily on single-moment state inputs, lacking a unified ability to characterize the temporal changes in the environment. This makes it difficult for models to form stable and reliable temporal structure perceptions. In complex or rare scenarios, models often fail to accurately predict the dynamic trends of the environment, thus affecting the continuity and robustness of vehicle behavior decision-making and control.

[0004] Furthermore, deep learning models are prone to "catastrophic forgetting" when faced with constantly changing driving environments: dynamic information in new scenarios can overwrite effective knowledge accumulated in existing scenarios, causing the system to continuously lose its understanding of long-term environmental structures. The knowledge formed by the system at different operational stages is difficult to maintain consistency, often resulting in unstable, unpredictable, or poorly generalized decision-making and control behavior under complex traffic conditions.

[0005] To alleviate these problems, existing research attempts to mine temporal patterns in the environment from continuous state data through methods such as temporal prediction, contrastive learning, and self-supervised representation learning, in order to improve the model's ability to understand dynamic traffic environments. However, such methods typically rely on long historical sequences or specific training settings, resulting in complex model structures, high computational costs, and difficulty in maintaining consistent temporal modeling performance across different scenarios. Furthermore, existing methods lack effective extraction and structured representation mechanisms for key changes in the environment over time, making it difficult to stably capture predictable temporal relationships during state changes, thus limiting the system's adaptability and long-term decision-control stability in continuous task operation. Summary of the Invention

[0006] In order to more effectively extract the environmental evolution law from the process of state change over time, construct stable time structure features, and realize the technical solution of continuous optimization of driving decision control, so as to improve the robustness and long-term reliability of autonomous driving system in real dynamic traffic environment, the purpose of this disclosure is to propose an autonomous driving continuous optimization method, system and storage medium based on potential state consistency modeling, which is used to realize cross-driving task knowledge transfer and optimization in dynamic traffic environment, and improve the adaptability and robustness of autonomous driving decision control system in complex long-tail scenarios.

[0007] To achieve the aforementioned technical objectives, a continuous optimization method for autonomous driving based on latent state consistency modeling is proposed. The steps include: during training, acquiring a sequence of environmental state information from n consecutive time steps within the same real-world trajectory of the controlled object for the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; based on the structural stability index and preset stability threshold for each basic state characteristic... The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain. After mapping the stable knowledge domain and the dynamic knowledge domain to a unified dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t , , , , Weights; based on environmental state characteristics and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragment The latent state representation is obtained. ; to provide environmental status information and potential state representation Input a policy decision network and output a predicted action; calculate the total loss during training, which consists of an action loss and a contrastive loss, wherein the contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features, thereby optimizing the robustness of the system controller. During inference, the current environmental state information and the latent state representation at the end of training are input into the trained policy decision network to output the predicted action.

[0008] Based on the autonomous driving continuous optimization method based on potential state consistency modeling proposed in this disclosure, a computer-readable storage medium can be obtained, which stores a computer program that can be loaded by a processor and execute any of the methods described in this disclosure.

[0009] Based on the autonomous driving continuous optimization method based on latent state consistency modeling proposed in this disclosure, an autonomous driving continuous optimization system based on latent state consistency modeling can be obtained. The system includes a feature encoding network module, a latent state transition network module, a policy decision network module, and a training loss calculation module. The feature encoding network module is configured to be used during training to acquire a sequence of environmental state information from n consecutive time steps within the same real trajectory of the controlled object in the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; based on the structural stability index and preset stability threshold for each basic state characteristic... The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain. After mapping the stable knowledge domain and the dynamic knowledge domain to a unified dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t , , , , As weight, To stabilize knowledge-related features, The latent state transition network module is configured to be used during training, based on dynamic knowledge-related features, and is based on environmental state features. and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragment The latent state representation is obtained. The policy decision network module is configured to incorporate environmental state information during training. and potential state representation The input policy decision network outputs predicted actions. During inference, the current environment state information and the latent state representation at the end of training are input into the trained policy decision network, which outputs predicted actions. The training loss calculation module is configured to calculate the total loss, which consists of action loss and contrastive loss. The contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features, thereby optimizing the robustness of the system controller.

[0010] The beneficial technical effects of this invention are: This solution can improve the adaptability, scalability and decision control robustness of the autonomous driving system in complex traffic environments, thereby improving the accuracy of autonomous driving decision control and enhancing the safety of autonomous driving. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of a continuous optimization method for autonomous driving based on potential state consistency modeling. Detailed Implementation

[0013] To address the shortcomings of existing autonomous driving systems, such as insufficient adaptability in complex multi-scenario and dynamic traffic environments, decreased decision-control stability during long-term operation, and inadequate characterization of environmental changes, this invention provides a continuous optimization method for autonomous driving based on latent state consistency modeling. The method includes the following steps: during training, acquiring a sequence of environmental state information from n consecutive time steps within the same real-world trajectory of the controlled object for the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; based on the structural stability index and preset stability threshold for each basic state characteristic... The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain, and the environmental state features at time t are extracted. , , , , As weight, To stabilize knowledge-related features, For dynamic knowledge-related features; based on environmental state features and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragment The latent state representation is obtained. ; to display current environmental status information and potential state representation Input a policy decision network and output a predicted action; calculate the total loss during training, which consists of an action loss and a contrastive loss, wherein the contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features; during inference, the current environmental state information and the latent state representation at the end of training are input into the trained policy decision network to output the predicted action.

[0014] Compared with traditional methods that rely on a single scenario or static model, the method of this invention can efficiently adapt to new tasks without relying on full retraining, while maintaining the system's stable cognition and knowledge reuse of historical tasks, thereby significantly improving the adaptability, scalability and decision control robustness of autonomous driving systems in complex traffic environments.

[0015] The following provides a clear and complete description of how the technical solution of this case is implemented. Obviously, the described implementation methods are only a part of the implementation methods of this case, and not all of them. Based on the implementation methods in this case, all other implementation methods obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0016] In one embodiment, the model used in the method of the present invention includes a feature encoding network, a latent state transition network, and a policy decision network. The feature encoding network acquires environmental state features considering structural stability from multi-source environmental state information in an autonomous driving scenario. The latent state transition network constructs state transition segments based on the environmental state features, thereby obtaining a latent state representation. The policy decision network outputs a predicted action based on the environmental state information and the latent state representation. These models can be configured in the same or different modules.

[0017] (a) Feature coding network

[0018] The environmental state information in this method includes camera images, LiDAR point clouds, millimeter-wave radar detection signals, ultrasonic distance measurements, GPS / IMU fusion positioning data, and multimodal data such as vehicle speed, acceleration, and attitude angles. The various sensor data within the environmental state information simultaneously contain information related to stable structures and information related to dynamic scenes. To enhance cross-task generalization capabilities, this invention designs a method to differentiate environmental state information based on structural stability.

[0019] In one implementation, environmental state information is recorded as , For camera images, For lidar point clouds, For millimeter-wave radar to detect signals, This refers to the ultrasonic distance. For GPS / IMU fusion positioning data, The vehicle state data consists of the vehicle's own speed, acceleration, and attitude angles.

[0020] Since environmental state information is multimodal, a multimodal feature encoding network is constructed using a convolutional neural network and a Transformer encoder. This network encodes the environmental state information of the controlled object for the current driving task over n consecutive time steps. Input the feature encoding network to extract the basic state feature sequence over n consecutive time steps. The basic state features are used to characterize the original evolutionary information of the environment in the time dimension.

[0021] Structural stability is assessed based on the magnitude of changes in the basic states over time, dividing the basic state feature sequence into a stable knowledge domain and a dynamic knowledge domain. Specifically, a structural stability index is calculated: , The basic state characteristics at time step i. When the structural stability index... Less than the preset stability threshold When the structural stability index is greater than or equal to the stability threshold, the corresponding basic state features are divided into a stable knowledge domain; when the structural stability index is greater than or equal to the stability threshold, the corresponding basic state features are divided into a dynamic knowledge domain.

[0022] Through the above implementation method, based on structural stability, this application divides the environmental state information of each controlled object in driving tasks at n time steps into a stable knowledge domain and a dynamic knowledge domain, such as... Figure 1 As shown.

[0023] Next, based on the structural stability index and preset stability threshold of each basic state characteristic, The basic state features are divided into stable knowledge domains or dynamic knowledge domains.

[0024] The stable knowledge domain and the dynamic knowledge domain can be viewed as two complementary components in the same high-dimensional feature space, both having the same vector dimension. The stable knowledge-related features characterize the environmental structure information that remains relatively unchanged across different time steps and driving tasks. The dynamic knowledge-related features characterize the environmental information that dynamically updates with changes in specific scenarios and tasks.

[0025] Task-independent features include environmental structural information such as road geometry, lane topology, traffic rules, and vehicle dynamics constraints. Traffic rules, such as speed limits and stop signs, can be represented using rule settings or decision trees, and the control decisions can be dynamically adjusted based on real-time input. Vehicle dynamics constraints, such as maximum acceleration and maximum steering angle, are represented by numerical constraints to ensure that the control strategy does not exceed physical limits. Dynamically updated environmental information includes traffic density, neighboring vehicle behavior, and weather conditions. Neighboring vehicle behavior, such as acceleration, deceleration, and lane changes, can be modeled using behavior prediction models and used as input for real-time decision control. This data can be processed using recurrent neural networks to predict future vehicle behavior. Weather conditions, such as rain, snow, and fog, are typically obtained through cameras and can affect vehicle speed and driving patterns.

[0026] Next, after mapping the stable knowledge domain and the dynamic knowledge domain to the same dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t : .in, , , where is the weight, used to adjust the contribution ratio of the two types of features in the construction of environmental state features. Due to stable knowledge-related features... It plays a dominant role in the construction of environmental state features and can also enhance the stability of temporal structure information in subsequent potential state consistency modeling. For example, it will... The value range is limited to 0.6 to 0.9. The value range is limited to 0.1 to 0.4.

[0027] The aforementioned weights are used in the feature fusion stage to distinguish the contributions of two types of features to model learning: stable knowledge-related features. It is given a higher weight during fusion, allowing it to have a greater impact on model parameter updates during backpropagation, thus reinforcing the structural constraints shared across tasks; dynamic knowledge-related features. They are assigned relatively low or adjustable weights to characterize the changing factors in specific tasks and scenarios, and to support the adaptive updates of the model to new tasks.

[0028] This weighted fusion method allows for the gradual absorption of dynamic information from new tasks while maintaining the stability of general knowledge, thereby improving the stability of continuous learning and cross-task transfer. The output forms the basic input for subsequent potential state transfer modeling. Using the same method, the future state relative to the current time step t is extracted from the same real trajectory. Step environmental state characteristics For example, the value of k is selected from the range 3-5.

[0029] In one implementation, a feature encoding network module is implemented, configured to sequence environmental state information over n consecutive time steps along the same real-world trajectory of the controlled object in the same driving task. As input, output the environmental state features at time t. The environmental state characteristics satisfy , , , As weight, To stabilize knowledge-related features, This refers to dynamic knowledge-related features. The feature encoding network is a multimodal feature encoding network, which consists of a convolutional neural network, a first Transformer encoder, an index calculation and classification module, and a second Transformer encoder connected in sequence. The environmental state information sequence is then used to... The input is a convolutional neural network, and the first Transformer encoder obtains the basic state feature sequence. The basic state feature sequence The input index calculation and classification module classifies stable and dynamic knowledge domains based on the structural stability index of each basic state feature. It then maps the stable and dynamic knowledge domains to the same dimension and outputs concatenated vectors of the stable and dynamic knowledge domains. These vectors are input to the second Transformer encoder, which extracts stable and dynamic knowledge-related features. The second Transformer encoder then performs weighted fusion at the output layer to output the environmental state features at time t. .

[0030] (ii) Latent State Transition Network

[0031] Autonomous driving systems are dynamic systems that continuously evolve over time, and the environmental state features at a single moment are insufficient to fully express the changing trends of the environment. To obtain reusable temporal evolution patterns across tasks, this invention constructs a latent state transition network based on the environmental state features output by the feature encoding network. The latent state transition network models the recursive relationship of environmental state evolution over time in the latent space. It generates the current latent state representation by combining the previous latent state with the current unified environmental state features, and constructs a contrastive learning loss by incorporating future environmental state features to enhance the temporal consistency and predictability of the latent representation. During training, this network employs a contrastive learning objective based on maximizing mutual information, with all network parameters jointly minimizing the same loss function.

[0032] Specifically, the input to the latent state transition network is not a long-term sequence, but rather a minimal time segment composed of environmental state features at the current and future times. The latent state transition network processes the state transition segment through a predictor and discriminator network. Perform contrastive modeling to maximize the latent representation The mutual information between the characteristics of the future environmental state and the target environment allows for the acquisition of stable and transferable latent state representations. This structure effectively captures the shareable patterns of environmental evolution over time and provides a unified latent temporal input for policy decision networks.

[0033] Environmental state features obtained from feature coding networks And extract the future first from the same real trajectory Step environmental state characteristics Together, they constitute a state transition segment: The state transition segments are used to characterize the changing trends of vehicle behavior, traffic flow, and road scene evolution over time.

[0034] Next, using latent state transition networks Mapping the state transition fragments to the latent state consistency space yields the latent state representation: The output of the latent state transition network is the latent state representation. This represents the stable implicit evolutionary structure of the environment state over time, and serves as the unified temporal input for subsequent policy learning.

[0035] Then, a contrastive predictive coding (CPC) objective based on contrastive learning is adopted to maximize the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features, ensuring that the latent state representation can capture the shared, consistent structure of different tasks.

[0036] Specifically, Figure 1 This illustrates Task 1, Task 2, Task 3, ..., Task N. During the process of constructing the potential space for these N tasks, the system represents the potential state obtained at the current moment. Environmental state features obtained from feature encoding networks at future times Form positive sample pairs, and simultaneously select environmental state features from other time steps of the same trajectory or from any time step of other trajectories from the same training batch. As negative samples, they form a contrastive set for time prediction. Based on this, a contrastive loss function based on vector similarity is adopted. By comparing positive and negative samples, the model strengthens the state association under real time progression and suppresses the confusion between irrelevant states. This makes the latent state representation have a higher degree of matching with the "real future state" while maintaining a lower response to irrelevant states, as shown below.

[0037]

[0038] in, This represents an exponential function with base e. This indicates the similarity between the characteristics of a potential state and its environmental state. is a trainable linear mapping matrix used to adjust the alignment between the latent space and the environmental state feature space, where T denotes transpose and N is the total number of negative samples. These are negative samples. By minimizing this loss, the system can enhance the sensitivity of the latent state representation to real-time progression, suppress responses to erroneous future candidates, and gradually form a stable and transferable temporal evolution structure in the latent space. This latent state representation not only fully preserves the changing patterns of the environment over time but also provides a unified temporal structure to support the policy decision network, enabling the system to naturally accumulate new experiences and maintain existing capabilities in continuous task sequences, achieving true cross-task transfer and continuous learning.

[0039] The size of the negative sample set is not fixed and can be dynamically determined based on the training batch size. Typically, all environmental state features other than positive samples in the same batch are treated as negative samples to enhance the discriminative power of the latent state representation. Those skilled in the art can adjust the number of negative samples appropriately based on computational resources, model size, and convergence speed requirements.

[0040] One embodiment of the potential state transition network module is configured to be based on environmental state features and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragment The latent state representation is obtained. The latent state representation output by the latent state transition network at the end of training is used as a consistent structural representation shared by the different tasks captured in this application.

[0041] (III) Strategy Decision Network

[0042] Represent the generated latent states As one of the unified inputs to the strategy decision network, it enables the strategy decision network to share time structure knowledge across different tasks and scenarios, thereby achieving continuous learning and cross-task transfer capabilities.

[0043] During the training phase, the policy decision network of this invention, at time step t, will process environmental state information. The latent state representation is output by the latent state transition modeling network. Furthermore, a multi-layer neural network is used to fuse and encode the state and temporal structure. The policy decision network calculates the action distribution based on the fused representation. The system outputs a predicted action. During the inference phase, only the trained policy decision network is applied. The current real-time environment state information and the latent state representation at the end of training are input into the trained policy decision network, which then outputs a predicted action.

[0044] The predicted actions include signals such as vehicle acceleration, steering angle, or higher-level driving intentions. The network is continuously updated during training, enabling it to fully utilize latent state representations in different tasks and scenarios, thereby improving the adaptability and robustness of the autonomous driving system.

[0045] Because the policy decision network directly utilizes the latent state representation By modeling the stable structural characteristics of the environment as it evolves over time, this invention eliminates the need for model switching, expert division, or parameter reset during continuous operation. Instead, it achieves gradual learning of environmental change patterns and continuous optimization of decision-making capabilities through continuous updates of the strategy decision network parameters.

[0046] In addition to the contrastive loss mentioned above, the feature encoding network, latent state transition network, and policy decision network of this invention also calculate the action loss during training. The total loss is the weighted sum of the action loss and the contrast loss. L= , These are preset weight parameters. A separate training loss calculation module can be configured to calculate the total loss.

[0047] In summary, during the training phase, this invention constructs a continuous processing link from environmental state information to potential structure and then to decision-making control strategy through the collaborative work of a feature encoding network, a latent state transition network, and a policy decision network, thereby achieving continuous learning capability for autonomous driving tasks in dynamic multi-scenario environments. Specifically, the feature encoding network first maps multi-source sensor inputs into environmental state features that include both stable knowledge-related features and dynamic knowledge-related features. This provides a complete environmental representation for subsequent time structure modeling; the latent state transition network constructs state transition segments based on environmental state features at intervals of k steps, and extracts latent state representations describing the environmental evolution law. This enables the system to capture common temporal patterns across tasks, laying a unified latent space for long-term knowledge retention. Based on this, the policy decision network generates control actions by fusing environmental state information with the long-term structure of the latent space, achieving real-time adjustment and policy updates for driving behavior. The three networks operate interconnected through parameter sharing and end-to-end training, with latent state representation... By continuously accumulating cross-scenario evolutionary features in the task sequence, the policy decision network can continuously absorb state changes in new tasks without resetting parameters or switching models. This achieves gradual knowledge expansion and long-term retention of old knowledge, forming a continuous learning mechanism based on latent structure, i.e., correlation-driven learning. Applying the latent state representation at the end of training to the inference stage enables the autonomous driving system to maintain stable, coherent decision-making and control performance with cross-task generalization capabilities in continuously changing task environments.

[0048] This invention can be a system, method, and / or computer program product.

[0049] For example, an autonomous driving continuous optimization method based on latent state consistency modeling includes the following steps: during training, acquiring a sequence of environmental state information for n consecutive time steps on the same real-world trajectory of the controlled object for the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; based on the structural stability index and preset stability threshold for each basic state characteristic... The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain. After mapping the stable knowledge domain and the dynamic knowledge domain to a unified dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t , , , , Weights; based on environmental state characteristics and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragment The latent state representation is obtained. ; to provide environmental status information and potential state representation Input a policy decision network and output a predicted action; calculate the total loss during training, which consists of an action loss and a contrastive loss, wherein the contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features; during inference, the current environmental state information and the latent state representation at the end of training are input into the trained policy decision network to output the predicted action.

[0050] For example, an autonomous driving continuous optimization system based on latent state consistency modeling includes a feature encoding network module, a latent state transition network module, a policy decision network module, and a training loss calculation module. The feature encoding network module is configured to be used during training to acquire a sequence of environmental state information from n consecutive time steps along the same real trajectory of the controlled object in the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; based on the structural stability index and preset stability threshold for each basic state characteristic... The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain. After mapping the stable knowledge domain and the dynamic knowledge domain to a unified dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t , , , , The weights are used as parameters; the latent state transition network module is configured to be used during training, based on features derived from the environment state. and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragment The latent state representation is obtained. The policy decision network module is configured to, during training, incorporate the current environment state information. and potential state representation The input policy decision network outputs predicted actions. During inference, the current environment state information and the latent state representation at the end of training are input into the trained policy decision network, which outputs predicted actions. The training loss calculation module is configured to calculate the total loss, which consists of action loss and contrastive loss. The contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features.

[0051] Computer program products may include computer-readable storage media on which computer-readable program instructions are loaded to enable a processor to implement various aspects of the present invention. A computer-readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combinations thereof. The computer-readable storage medium as used herein is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0052] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0053] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Python, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0054] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0055] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0056] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0057] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.

[0058] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.

Claims

1. A continuous optimization method for autonomous driving based on latent state consistency modeling, characterized by the following steps: include: During training, a sequence of environmental state information for n consecutive time steps is obtained from the same real-world trajectory of the controlled object in the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; Structural stability index and preset stability threshold based on each basic state characteristic The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain. After mapping the stable knowledge domain and the dynamic knowledge domain to a unified dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t , , , , As weight; Based on environmental state characteristics and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragments constituted The latent state representation is obtained. ; Environmental status information and potential state representation Input the policy decision network, output the predicted action; During training, the total loss is calculated, consisting of action loss and contrastive loss, where the contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features. During inference, the current environment state information and the potential state representation at the end of training are input into the trained policy decision network, which then outputs the predicted action.

2. The method according to claim 1, characterized in that, The environmental status information includes camera images, lidar point clouds, millimeter-wave radar detection signals, ultrasonic distance measurements, GPS / IMU fusion positioning data, as well as the vehicle's own speed, acceleration, and attitude angle.

3. The method according to claim 1, characterized in that, The structural stability index is calculated using the following formula. , The basic state features at time step i.

4. The method according to claim 1, characterized in that, By maximizing the latent state representation Characteristics of future environmental conditions The mutual information is achieved through the following steps: Representing the latent state Characteristics of environmental conditions To form positive sample pairs, environmental state features from other time steps of the same trajectory or any time step of other trajectories are selected from the same training batch. If the sample is negative, the contrastive loss is calculated as follows: In the formula: This indicates the similarity between the characteristics of a potential state and its environmental state. is a trainable linear mapping matrix used to adjust the alignment between the latent space and the environmental state feature space, where T denotes transpose and N is the total number of negative samples.

5. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed according to any one of claims 1 to 4.

6. An autonomous driving continuous optimization system based on latent state consistency modeling, characterized in that, The system includes a feature encoding network module, a latent state transition network module, a policy decision network module, and a training loss calculation module; The feature encoding network module is configured to be used during training to obtain a sequence of environmental state information from n consecutive time steps in the same real trajectory of the controlled object for the same driving task. Based on environmental state information sequence Extracting basic state feature sequences For each basic state characteristic, calculate the structural stability index; based on the structural stability index and preset stability threshold for each basic state characteristic... The basic state feature sequence is divided into a stable knowledge domain and a dynamic knowledge domain. After mapping the stable knowledge domain and the dynamic knowledge domain to a unified dimension, the corresponding stable knowledge-related features are extracted. Features related to dynamic knowledge Construct the environmental state features at time t , , , , As weight; The latent state transition network module is configured to be used during training, based on features derived from the environment state. and the future extracted from the same real trajectory Step environmental state characteristics The state transition fragments constituted The latent state representation is obtained. ; The policy decision network module is configured to incorporate environmental state information during training. and potential state representation Input the policy decision network and output the predicted action; during inference, input the current environment state information and the potential state representation at the end of training into the trained policy decision network and output the predicted action. The training loss calculation module is configured to calculate the total loss, which consists of action loss and contrastive loss. The contrastive loss is calculated by maximizing the latent state representation. Characteristics of future environmental conditions The mutual information enables the latent space to extract stable, predictable, and task-independent environmental evolution features.

7. The system according to claim 6, characterized in that, The environmental status information includes camera images, lidar point clouds, millimeter-wave radar detection signals, ultrasonic distance measurements, GPS / IMU fusion positioning data, as well as the vehicle's own speed, acceleration, and attitude angle.

8. The system according to claim 6, characterized in that, The structural stability index is calculated using the following formula. , The basic state features at time step i.

9. The system according to claim 6, characterized in that, By maximizing the latent state representation Characteristics of future environmental conditions The mutual information is achieved through the following steps: Representing the latent state Characteristics of environmental conditions Form positive sample pairs, and simultaneously select environmental state features from other time points or scenarios from the same training batch. If the sample is negative, the contrastive loss is calculated as follows: In the formula: This indicates the similarity between the characteristics of a potential state and its environmental state. is a trainable linear mapping matrix used to adjust the alignment between the latent space and the environmental state feature space, where T denotes transpose and N is the total number of negative samples.

Citation Information

Patent Citations

  • Safety key scene generation system and method for automatic driving decision algorithm

    CN118709520A

  • Method for training end-to-end autonomous driving strategy

    WO2023102962A1