Target behavior prediction method and related equipment

By constructing a dynamic causal graph and updating the node weights in real time, the problem of large prediction deviations in the existing technology in complex scenarios is solved, and high-accurate target behavior prediction and environmental adaptability are achieved.

CN120145203AActive Publication Date: 2025-06-13BYZORO NETWORK LTD +1

Patent Information

Application Number
CN202510629493.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing target behavior prediction technology has large prediction deviations in complex scenarios and lacks the ability to deeply model the causal relationship between environmental changes and target behavior.

Method used

By obtaining the multimodal historical behavior data and real-time environmental semantic information of the target object, key causal event nodes are determined based on the spatiotemporal causal model, a dynamic causal graph containing environmental interaction relationships is constructed, node weights are updated in real time, and predicted behavior trajectories of the target object are generated.

Benefits of technology

It significantly improves the accuracy and environmental adaptability of target behavior prediction, overcomes the prediction deviation problem of traditional methods in complex scenarios, and achieves the robustness of dynamic avoidance and behavior prediction of sudden obstacle scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145203A_ABST
    Figure CN120145203A_ABST
Patent Text Reader

Abstract

The invention discloses a target behavior prediction method and related equipment, and relates to the technical field of target prediction, and the method comprises the steps: obtaining multi-modal historical behavior data and real-time environment semantic information of a target object; determining key causal event nodes based on the multi-modal historical behavior data and a space-time causal model; based on the key causal event nodes and the real-time environment semantic information, determining a dynamic causal graph; determining a future potential event type and a corresponding event occurrence probability based on the dynamic cause and effect graph; determining an event triggering constraint condition based on the event occurrence probability; determining node weight updating parameters of the dynamic causal graph based on an incremental learning framework and multi-modal historical behavior data input in real time; and determining a predicted behavior track of the target object based on the event triggering constraint condition and the updated dynamic causal graph. According to the invention, by fusing the multi-modal data and space-time causal reasoning, the accuracy and environmental adaptability of target behavior prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of target prediction, and in particular, to a target behavior prediction method and related devices. Background Art

[0002] In the field of target behavior prediction, existing technologies mostly rely on static environment modeling and simple time-series analysis of historical behavior data. Traditional methods usually adopt prediction frameworks based on rules or statistical models. For example, in the field of autonomous driving, sudden behaviors of pedestrians or vehicles often lead to prediction deviations, and existing models lack the ability to deeply model the causal relationship between environmental changes and target behaviors. Therefore, there is an urgent need for a target behavior prediction method to solve the above-mentioned technical problems. Summary of the Invention

[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further elaborated in detail in the Detailed Description section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0004] In a first aspect, this application provides a target behavior prediction method, the method comprising: Obtaining multi-modal historical behavior data of a target object and real-time environmental semantic information; Determining key causal event nodes based on the multi-modal historical behavior data and a spatio-temporal causal model; Determining a dynamic causal graph including environmental interaction relationships based on the key causal event nodes and the real-time environmental semantic information; Determining future potential event types and corresponding event occurrence probabilities based on the dynamic causal graph; Determining event trigger constraint conditions based on the event occurrence probabilities; Determining node weight update parameters of the dynamic causal graph based on an incremental learning framework and real-time input multi-modal historical behavior data; Determining a predicted behavior trajectory of the target object based on the event trigger constraint conditions and the updated dynamic causal graph.

[0005] In some embodiments, obtaining multi-modal historical behavior data of a target object and real-time environmental semantic information includes: Determining the motion speed and direction angle data of the target object based on the detection signal of a millimeter-wave radar; Determining the contour features and local environmental images of the target object based on the image acquisition results of a vision sensor; Determining a three-dimensional environmental topology map based on the point cloud data of a lidar, wherein the three-dimensional environmental topology map includes motion direction annotation information of dynamic obstacles. Based on a preset timestamp synchronization rule, align the motion speed, direction angle data, contour features, local environment images, and three-dimensional environment topological maps to determine multi-modal historical behavior data; Based on the motion direction annotation information, determine the real-time environment semantic information.

[0006] In some embodiments, based on the multi-modal historical behavior data and the spatio-temporal causal model, determine key causal event nodes, including: Based on the time series characteristics of the multi-modal historical behavior data, determine multiple historical time segments; Based on the event detection rules in the spatio-temporal causal model, extract features from each historical time segment to determine speed mutation points, direction turning points, and environment interaction events; Based on a preset causal association threshold, mark the speed mutation points, direction turning points, and environment interaction events to determine a set of candidate causal event nodes; Based on the Granger causality test algorithm in the spatio-temporal causal model, perform association elimination on the set of candidate causal event nodes to determine key causal event nodes.

[0007] In some embodiments, based on the key causal event nodes and the real-time environment semantic information, determine a dynamic causal graph including environment interaction relationships, including: Based on the key causal event nodes, determine a set of event nodes; Based on the real-time environment semantic information, determine the motion direction and position of dynamic obstacles; Based on the set of event nodes, motion direction, and position, determine the environment interaction weight, where the environment interaction weight is used to characterize the association strength between event nodes and dynamic obstacles; Based on the set of event nodes and the environment interaction weight, determine an initial causal graph with weights; Based on a preset causal redundancy threshold, eliminate the connection edges in the initial causal graph with weights less than the causal redundancy threshold to determine the dynamic causal graph.

[0008] In some embodiments, based on the dynamic causal graph, determine future potential event types and corresponding event occurrence probabilities, including: Based on the set of event nodes in the dynamic causal graph, determine a set of candidate event types; Based on the connection weights of each event node in the dynamic causal graph, determine the node activation probability of each candidate event type; Based on a preset probability threshold, screen the node activation probability to determine the candidate event types that meet the preset probability threshold as future potential event types; Determine the event occurrence probability based on the node activation probability corresponding to future potential event types, where the event occurrence probability has a positive correlation with the node activation probability.

[0009] In some embodiments, determine the event trigger constraint conditions based on the event occurrence probability, including: Determine the probability distribution parameters of each future potential event type based on the event occurrence probability; Determine the time window and spatial region for event triggering based on the preset trigger rules and probability distribution parameters; Determine the dynamic threshold conditions for event triggering based on the time window and spatial region, where the dynamic threshold conditions include the trigger time interval, trigger spatial boundary, and minimum probability threshold; Determine the event trigger constraint conditions based on the dynamic threshold conditions and the event occurrence probability.

[0010] In some embodiments, determine the node weight update parameters of the dynamic causal graph based on the incremental learning framework and real-time input multi-modal historical behavior data, including: Determine the incremental input data set based on the real-time input multi-modal historical behavior data; Determine the distribution difference value between the incremental input data set and the historical data set based on the distribution difference detection algorithm in the incremental learning framework; Determine the model update trigger signal based on the preset distribution difference threshold and the distribution difference value; Determine the knowledge transfer parameters of the online knowledge distillation model based on the model update trigger signal and the incremental input data set; Determine the node weight update parameters based on the knowledge transfer parameters and the current node weights of the dynamic causal graph, where the node weight update parameters are used to adjust the association strength between event nodes and dynamic obstacles in the dynamic causal graph.

[0011] In some embodiments, determine the predicted behavior trajectory of the target object based on the event trigger constraint conditions and the updated dynamic causal graph, including: Determine multiple candidate trajectory branches based on the time window and spatial region; Determine the posterior probability of each candidate trajectory branch based on the node weights in the updated dynamic causal graph; Determine the main predicted trajectory based on the maximum value of the posterior probability; Determine the trajectory conflict probability based on the main predicted trajectory and the positions of dynamic obstacles in the real-time environmental semantic information; Determine the trajectory correction vector based on the trajectory conflict probability and the preset conflict threshold; Determine the predicted behavior trajectory of the target object based on the trajectory correction vector and the preset smoothing algorithm.

[0012] In a second aspect, the present application provides an apparatus for predicting target behavior, including: A target object data acquisition unit, configured to acquire multimodal historical behavior data and real-time environmental semantic information of a target object; A key event node determination unit, configured to determine key causal event nodes based on the multimodal historical behavior data and a spatio-temporal causal model; A dynamic causal graph generation unit, configured to determine a dynamic causal graph including environmental interaction relationships based on the key causal event nodes and the real-time environmental semantic information; A future potential event determination unit, configured to determine future potential event types and corresponding event occurrence probabilities based on the dynamic causal graph; An event constraint condition determination unit, configured to determine event trigger constraint conditions based on the event occurrence probabilities; A node weight parameter update unit, configured to determine node weight update parameters of the dynamic causal graph based on an incremental learning framework and the real-time input multimodal historical behavior data; A target object behavior prediction unit, configured to determine a predicted behavior trajectory of the target object based on the event trigger constraint conditions and the updated dynamic causal graph.

[0013] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program stored in the memory, the steps of the target behavior prediction method according to any one of the first aspects are implemented.

[0014] In summary, the present application improves the accuracy and environmental adaptability of target behavior prediction by integrating multimodal data and spatio-temporal causal reasoning. First, by modeling the environmental interaction relationships based on the dynamic causal graph, the causal event associations between the target object and obstacles can be explicitly captured, solving the prediction deviation problem of traditional methods in complex scenarios. Second, by updating the node weights in real time through the incremental learning framework, the model can continuously adapt to the new data distribution, avoiding prediction failures caused by environmental changes. Finally, by generating a probabilistic trajectory in combination with the event trigger constraint conditions, dynamic avoidance and smooth correction of sudden obstacle conflicts are achieved. Description of the Drawings

[0015] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit the present specification. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a schematic flowchart of a target behavior prediction method provided by an embodiment of the present application; Figure 2Schematic diagram of a target behavior prediction device provided by an embodiment of the present application; Figure 3 Structural diagram of an electronic device for target behavior prediction provided by an embodiment of the present application. Detailed implementation manners

[0016] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0017] Please refer to Figure 1 , which is a schematic diagram of a target behavior prediction process provided by an embodiment of the present application, and specifically may include: S110. Obtain multimodal historical behavior data and real-time environmental semantic information of a target object; Exemplarily, historical behavior data and real-time environmental states of a target object are collected through multi-source sensors. The acquisition of multimodal historical behavior data involves devices such as millimeter-wave radars, vision sensors, and lidar, which are respectively used to capture the motion parameters (such as speed, direction angle) of the target object, morphological features (such as contour, local environmental image), and three-dimensional space information (such as the motion direction annotation of dynamic obstacles). These heterogeneous data are aligned and fused through a preset timestamp synchronization rule to form a unified multimodal dataset, ensuring temporal consistency and spatial correlation, and providing a comprehensive behavioral representation basis for subsequent causal reasoning.

[0018] The acquisition of real-time environmental semantic information focuses on the real-time state monitoring of dynamic obstacles. Based on the three-dimensional environmental topology map generated by lidar, combined with the motion direction annotation information of dynamic obstacles, the spatial structure of the environment where the target object is located and the behavioral characteristics of obstacles are updated in real time. This semantic information not only includes static environmental elements, but also emphasizes the capture of dynamic interaction relationships, such as the movement trend of obstacles and potential conflict areas, thereby providing key environmental context inputs for subsequent construction of a dynamic causal graph and ensuring that the prediction model can adapt to the real-time change requirements of complex scenarios.

[0019] S120. Determine key causal event nodes based on multi-modal historical behavior data and spatio-temporal causal models; Exemplarily, in the process of determining key causal event nodes, it is necessary to conduct a structured analysis of the time series characteristics of multi-modal historical behavior data. Based on the preset historical time segment division rules, continuous behavior data is discretized into time units with independent semantics, and through the event detection rules in the spatio-temporal causal model, significant behavior characteristics within each time segment are extracted, such as speed mutation points, direction turning points, and interaction events with environmental obstacles. These characteristics reflect the abnormal behavior or pattern changes of the target object within a specific spatio-temporal range, providing a preliminary set of candidate event nodes for subsequent causal reasoning.

[0020] Furthermore, through the causal association evaluation mechanism in the spatio-temporal causal model, the candidate event nodes are screened and optimized. Based on the preset causal association threshold, isolated or weakly associated events are eliminated, and key nodes with significant causal relationships with the behavior evolution of the target object are retained. At the same time, combined with the statistical causal test algorithm, the temporal causality between event nodes is verified, and pseudo-correlation interference is excluded. Finally, the core causal event nodes that can characterize the dynamic evolution of the target behavior are determined, laying a foundation for constructing a dynamic causal graph.

[0021] S130. Determine a dynamic causal graph containing environmental interaction relationships based on the key causal event nodes and real-time environmental semantic information; Exemplarily, when constructing a dynamic causal graph, first, it is necessary to integrate the key causal event nodes and real-time environmental semantic information to establish the interaction relationship between the event nodes and dynamic obstacles. Based on the obstacle movement direction and position data provided by the real-time environmental semantic information, combined with the set of key causal event nodes, the environmental interaction weight is calculated to quantify the association strength between the event nodes and obstacles. Through weight assignment, an initial weighted causal graph is generated, where the nodes represent key events and the connecting edges reflect the causal influence degree of environmental interaction, thus initially depicting the interaction network between the target object and the dynamic environment.

[0022] Furthermore, the initial causal graph is dynamically optimized through a preset causal redundancy threshold to eliminate redundant connecting edges with weights lower than the threshold. This process retains strongly associated causal paths, ensuring that the causal graph only contains significant environmental interaction relationships and avoiding noise interference. The finally formed dynamic causal graph can reflect the association evolution between the target object's behavior and the dynamics of obstacles in real time, providing a structured causal reasoning framework for subsequent event probability prediction.

[0023] S140. Determine future potential event types and the corresponding event occurrence probabilities based on the dynamic causal graph; Exemplarily, when predicting future potential event types, based on the topological structure and connection weights of event nodes in the dynamic causality diagram, the activation possibility of candidate event types is deduced. The event nodes in the dynamic causality diagram represent the causal associations between historical behaviors and environmental interactions, and their connection weights reflect the intensity of causal influence between events. By aggregating the weight transfer effects of adjacent nodes, the node activation probabilities of each candidate event type are calculated, initially forming a set of event types and probability distributions that may occur in the future, providing a basis for subsequent probability screening.

[0024] Furthermore, the candidate event types are probabilistically screened and sorted in combination with a preset probability threshold. By comparing the node activation probability with the threshold, low-probability events are eliminated, and potential event types with significant occurrence possibilities are retained. Finally, based on the screened event types and their corresponding activation probabilities, a quantitative expression of the event occurrence probability is generated, ensuring that the prediction results not only conform to the dynamic evolution law of the causality diagram but also satisfy the feasibility constraints of the actual scenario.

[0025] S150. Determine the event trigger constraint conditions based on the event occurrence probability; Exemplarily, when determining the event trigger constraint conditions, first, based on the distribution characteristics of the event occurrence probability, the probability distribution parameters of future potential event types (such as mean, variance, or confidence interval) are extracted, and the temporal and spatial trigger boundaries of the event are quantified in combination with preset trigger rules (such as safety thresholds, priority strategies). For example, high-probability events may be assigned shorter time windows and more precise spatial regions to ensure the matching of the prediction results with the dynamics of the real-time environment. Through the integration of probability distribution parameters and trigger rules, the temporal window and spatial region constraints for event triggering are initially generated, providing a quantitative basis for the definition of dynamic threshold conditions.

[0026] Furthermore, by integrating the time window, spatial region, and minimum probability threshold, a multi-dimensional dynamic threshold condition is constructed. The time window restricts the triggering timeliness of the event, the spatial region limits the physical range where the event occurs, and the minimum probability threshold filters out candidate events with high confidence. These threshold conditions jointly form a comprehensive constraint framework for event triggering, ensuring both the consistency between the prediction behavior and the probability distribution and taking into account the real-time adaptation requirements of environmental changes, providing an operable trigger logic boundary for subsequent predictions.

[0027] S160. Determine the node weight update parameters of the dynamic causality diagram based on the incremental learning framework and real-time input multi-modal historical behavior data; Exemplarily, during the node weight update process of the dynamic causal graph, multi-modal historical behavior data is first accessed in real time through an incremental learning framework to detect the distribution difference between the new data and the historical data set. When the distribution difference of the incremental input data exceeds a preset threshold, the model update mechanism is triggered to generate knowledge transfer parameters. This mechanism ensures that the model can adapt to environmental changes, avoid the degradation of prediction performance caused by data distribution drift, and at the same time maintain the real-time consistency between the node weights in the dynamic causal graph and the environmental interaction intensity.

[0028] Furthermore, based on the online knowledge distillation model, the new knowledge in the incremental data is transferred to the dynamic causal graph. By fusing the knowledge transfer parameters with the current node weights, the association strength between the event nodes and the dynamic obstacles is readjusted. This process locally optimizes the node weight parameters without destroying the original causal structure, enabling the dynamic causal graph to continuously reflect the latest interaction characteristics between the target object and the environment, and providing a dynamically updated causal reasoning basis for the generation of subsequent predicted behavior trajectories.

[0029] S170. Determine the predicted behavior trajectory of the target object based on the event trigger constraint conditions and the updated dynamic causal graph.

[0030] Exemplarily, when generating the predicted behavior trajectory of the target object, multiple candidate trajectory branches are first generated based on the time window and spatial region in the event trigger constraint conditions. Through the weight distribution of the event nodes in the updated dynamic causal graph, the posterior probability of each candidate trajectory branch is calculated, reflecting its matching degree with the historical causal association and the environmental interaction intensity. Finally, the candidate trajectory with the highest posterior probability is selected as the main predicted trajectory, preliminarily representing the most likely behavior path of the target object in the future spatio-temporal range.

[0031] Furthermore, in combination with the position of the dynamic obstacle in the real-time environmental semantic information, the conflict probability between the main predicted trajectory and the obstacle is evaluated. If the trajectory conflict probability exceeds the preset threshold, the spatial path of the main predicted trajectory is adjusted through the trajectory correction vector, and the trajectory mutation is eliminated using a smoothing algorithm to ensure that the corrected predicted behavior trajectory not only satisfies the high probability of causal association but also meets the safety obstacle avoidance requirements of the dynamic environment, and finally outputs a smooth and environment-adaptive behavior prediction result.

[0032] In summary, in the embodiments of the present application, by fusing multi-modal historical behavior data and real-time environmental semantic information, and combining a spatio-temporal causal model to construct a dynamic causal graph, the accuracy of target behavior prediction and scene adaptability are improved. The dynamic causal graph explicitly models the environmental interaction relationship between the target object and dynamic obstacles. Through the weight association and redundancy elimination mechanism of causal event nodes, the key causal links in complex scenarios are effectively captured, overcoming the prediction deviation problem caused by traditional methods ignoring the dynamic interaction of the environment. The incremental learning framework updates the node weight parameters in real time, enabling the model to continuously adapt to the new data distribution, avoiding prediction failures caused by sudden environmental changes, and at the same time ensuring the stability of the causal structure. Based on the trigger constraint conditions of event occurrence probability and the posterior probability screening of candidate trajectories, both high-confidence prediction and real-time environmental conflict avoidance are considered. Finally, a safe and coherent behavior trajectory is generated through trajectory correction and smoothing algorithms, realizing dynamic avoidance and the robustness of behavior prediction in scenarios with sudden obstacles.

[0033] In some instances, obtaining the multi-modal historical behavior data and real-time environmental semantic information of the target object includes: Based on the detection signal of the millimeter-wave radar, determining the motion speed and direction angle data of the target object; Based on the image acquisition result of the vision sensor, determining the contour feature and local environment image of the target object; Based on the point cloud data of the lidar, determining a three-dimensional environmental topology map, where the three-dimensional environmental topology map includes the motion direction annotation information of dynamic obstacles; Based on the preset timestamp synchronization rule, aligning the motion speed, direction angle data, contour feature, local environment image, and three-dimensional environmental topology map to determine the multi-modal historical behavior data; Based on the motion direction annotation information, determining the real-time environmental semantic information.

[0034] Exemplarily, a millimeter-wave radar calculates the radial velocity of a target object by transmitting high-frequency electromagnetic waves and receiving the signals reflected by the target object, based on the Doppler effect. Specifically, when the electromagnetic waves encounter a moving target, the frequency of its reflected signal will shift due to the radial movement of the target relative to the radar. This frequency shift is linearly related to the velocity of the target, and the instantaneous radial velocity of the target can be obtained by solving the frequency shift. To obtain the direction angle information of the target, the millimeter-wave radar adopts a multiple-input multiple-output (MIMO) array antenna structure and combines beamforming technology. By analyzing the phase difference between different receiving antennas, the spatial distribution of the target in the horizontal azimuth angle is calculated. The beamforming algorithm synthesizes the signals received by each antenna through weighting to generate a directional beam. When the target is located at different azimuths, the peak value of the received signal intensity corresponds to its direction angle. In addition, the time-domain signal is subjected to spectral analysis through fast Fourier transform (FFT) to extract the characteristics of the target in the range and velocity domains, and further separate the velocity and angle information in a multi-target scenario. Finally, the motion parameters output by the millimeter-wave radar include the real-time velocity value and direction angle data of the target object. These parameters are stored in a sequential form indexed by timestamps, forming a continuous time-series data set, providing high-precision and low-latency underlying motion characteristics for subsequent behavior analysis.

[0035] A vision sensor (such as an RGB camera or an infrared camera) obtains its contour features and local environment images by capturing visible light or infrared spectral images of the target object. Specifically, the original images collected by the vision sensor first go through preprocessing steps (including denoising, contrast enhancement, and geometric correction) to eliminate the interference of illumination changes, motion blur, or sensor noise on feature extraction. Subsequently, an edge detection algorithm (such as the Canny operator) or a deep learning-based feature extraction model (such as a convolutional neural network) is used to identify the contour features of the target object from the preprocessed images. For targets such as pedestrians or vehicles, the contour features are specifically manifested as key morphological structures such as limb joint connection points and the geometric shape of the vehicle exterior. These features are characterized by pixel-level coordinates or vectorized geometric parameters. At the same time, the acquisition of the local environment image depends on the focusing and segmentation of the area around the target object in the image. For example, through the region of interest (ROI) delineation algorithm, a local image block containing the target object and its adjacent environment is cropped from the complete image. This local image not only retains the spatial relative position information of the target and surrounding objects (such as road markings, adjacent obstacles), but also describes the environmental details through visual features such as texture and color, forming context information related to the target behavior. Through the above steps, the contour features and local environment images output by the vision sensor provide a visual representation for multi-modal historical behavior data, complementing the physical parameters of millimeter-wave radars and lidars.

[0036] The determination of the 3D environmental topological map is realized based on the high-density 3D point cloud data collected by lidar. First, the lidar emits laser pulses and receives reflected signals to obtain the distance information of the surfaces of various objects in the target scene, generating the original point cloud data. After preprocessing these point cloud data (such as denoising and ground segmentation), the point cloud clustering algorithm (such as the density-based DBSCAN algorithm) is used to group the discrete points and separate independent target entities (such as vehicles, pedestrians, trees, etc.). Among them, the DBSCAN algorithm clusters the point clouds with dense spatial distribution into independent objects by presetting the neighborhood radius and minimum point number threshold, so as to distinguish the static environmental structure from the dynamic obstacles. Subsequently, for dynamic obstacles, dynamic target tracking techniques (such as Kalman filtering or particle filtering) are adopted, combined with the temporal correlation of continuous frame point cloud data, to calculate their motion vectors (including instantaneous velocity, acceleration, and direction angle). This process updates the position and motion state of the obstacles in real time through prediction and correction models, and eliminates noise interference based on the preset motion consistency criterion (such as the acceleration limit threshold). Finally, the real-time positions and motion direction vectors of the static environmental structure (such as road boundaries, traffic signs) and dynamic obstacles are integrated into the 3D coordinate system to form a topological map with spatio-temporal correlation. Through coordinate system transformation and rasterization processing, this map maps the point cloud data into a structured environmental model, where the motion direction of dynamic obstacles is marked by vector arrows, and the static structure is presented in the form of geometric grids, thus providing a 3D environmental representation basis for subsequent behavior prediction.

[0037] In the process of determining multi-modal historical behavior data, aiming at the differences in the heterogeneous data acquisition frequencies of millimeter-wave radars, vision sensors, and lidar, spatio-temporal alignment of multi-source data is achieved through preset timestamp synchronization rules. Specifically, the millimeter-wave radar samples the motion speed and direction angle data of the target object at a high frequency, the vision sensor captures contour features and local environment images at a fixed frame rate, and the lidar generates a discrete three-dimensional environmental topology map through pulse scanning. Since the data acquisition periods of different sensors are different, a hardware clock synchronization mechanism is adopted to assign a unified timestamp reference for each sensor's data, ensuring the initial alignment of the raw data in the time dimension. For the asynchronous data segments caused by inconsistent sensor sampling intervals, such as when the lidar has not completed a full scan at a certain moment while the vision sensor has generated an image frame, a sliding window alignment strategy is used to delimit the time window range (such as ±10 ms), and linear interpolation is performed on the parameter values at the missing moments within the window to make the motion speed, direction angle, contour features, and three-dimensional map data strictly match on the unified time axis. In addition, for the motion direction annotation information of dynamic obstacles, combined with its timestamp and the behavior data of the target object, a quadratic interpolation algorithm is used to fit the continuous motion trajectory to eliminate the data breaks caused by discrete sampling. Through the above synchronization rules, heterogeneous sensor data is fused into a multi-modal historical behavior dataset with consistent time series and spatial correlation, ensuring the time coherence and spatial integrity of subsequent analysis.

[0038] The determination of real-time environmental semantic information is based on the motion direction annotation and the real-time updated position coordinates of dynamic obstacles in the three-dimensional environmental topology map. Specifically, through the high-density point cloud data obtained by the lidar, a method such as Kalman filtering is used to real-time identify and annotate the motion direction of obstacles, generating direction parameters including the magnitude and angle of the velocity vector. At the same time, combined with the real-time position coordinates of the obstacles in the three-dimensional space, the instantaneous state of its motion trajectory is deduced through spatial geometric calculations. On this basis, through vector decomposition and kinematic models, the real-time motion characteristics of the obstacles (such as speed, acceleration, and steering angle) are quantified and encoded into semantic labels, such as behavior patterns like "moving straight at a constant speed" and "accelerating and turning left". In addition, based on the spatio-temporal projection of the current position and motion direction of the obstacle, its potential motion path in the short term in the future is calculated, and combined with the predicted trajectory of the target object, the interaction area between the two (such as the collision risk area or intersection point) is identified. Finally, by integrating the above motion direction, position coordinates, behavior patterns, and interaction area information, a real-time environmental semantic description with spatio-temporal consistency is constructed, providing quantitative interaction features of dynamic obstacles for subsequent causal reasoning.

[0039] In some instances, based on the multi-modal historical behavior data and the spatio-temporal causal model, key causal event nodes are determined, including: Determining multiple historical time segments based on the time series characteristics of the multi-modal historical behavior data; Based on the event detection rules in the spatio-temporal causal model, feature extraction is performed on each historical time segment to determine speed mutation points, direction turning points, and environmental interaction events; Based on a preset causal association threshold, the speed mutation points, direction turning points, and environmental interaction events are marked to determine a set of candidate causal event nodes; Based on the Granger causality test algorithm in the spatio-temporal causal model, the set of candidate causal event nodes is associated and eliminated to determine key causal event nodes.

[0040] Exemplarily, when determining the historical time segments, first, based on the time series features of the multi-modal historical behavior data, the sliding window algorithm is used to segment the continuous time series data. Specifically, a preset time segment length is used as the sliding benchmark of the window, and heterogeneous data such as the movement speed, direction angle, contour features, and environmental topological information of the target object are divided into multiple independent time segments according to a fixed duration. Each time segment contains the complete behavior representation of the target object during that period, such as the continuous sampling values of the speed sequence, the instantaneous change curve of the direction angle, the geometric evolution trajectory of the contour features, and the relative positions and movement trends of the dynamic obstacles in the environmental topological information. To ensure the continuity of the time segment division and avoid the loss of key behavior features at the segment boundaries, a preset proportion of overlapping regions is set between adjacent time segments, and the complete time series data stream is covered through the progressive advancement of the sliding window. This overlapping mechanism effectively eliminates the risk of data fragmentation, enabling the continuous capture of the behavior patterns of the target object during the transition stage between adjacent segments (such as speed mutations during acceleration or angle fine-tuning at the initial stage of turning), thus providing a spatio-temporally coherent multi-modal data basis for the subsequent extraction of causal event nodes.

[0041] When extracting candidate event features, according to the divided historical time segments and based on the event detection rules preset by the spatio-temporal causal model, speed mutation points, direction turning points and environmental interaction events are determined respectively. The spatio-temporal causal model is a mathematical model that combines time series analysis and spatial correlation reasoning, aiming to reveal the causal relationship between the behavior of the target object and dynamic obstacles in the environment. For the detection of speed mutation points, the differential threshold method is used to analyze the speed sequence of the target object, that is, to calculate the absolute value of the speed difference between adjacent time points. If this difference exceeds the preset mutation threshold, it is determined that a speed mutation event occurs at the current moment, and its occurrence timestamp and mutation amplitude are recorded. The identification of direction turning points is achieved by calculating the rate of change of the direction angle, specifically, the change amount of the direction angle within a unit time. If this change amount exceeds the preset turning threshold, it is marked as a direction turning event, and the turning angle and duration are recorded at the same time. The detection of environmental interaction events needs to combine the spatio-temporal relationship between the target object and dynamic obstacles, that is, by calculating the Euclidean distance between the two in real time. If the distance continuously decreases within the time segment and is lower than the preset safety distance, and the included angle between their motion vectors is less than the preset angle threshold, it is determined as an environmental interaction event, and the obstacle identifier, interaction distance and vector included angle parameters are recorded. The above detection results together constitute the event feature set within the time segment, providing basic data for subsequent causal association analysis.

[0042] When determining the set of candidate causal event nodes, a causal relevance assessment is performed on the extracted speed mutation points, direction turning points and environmental interaction events. Specifically, based on the preset causal association threshold (for example, using the Pearson correlation coefficient or mutual information as a quantification index), the association strength between the event and the evolution of the target behavior is analyzed. For a speed mutation point, if the time interval between its occurrence time and subsequent direction turning events or environmental interaction events is within the preset time window, and the correlation coefficient between the two exceeds the preset threshold, it is determined that there is a causal association between this speed mutation point and the subsequent event; similarly, if a direction turning event is closely related to a subsequent environmental interaction event in space and time (for example, the obstacle position changes synchronously within the same time window), and the correlation coefficient between the two exceeds the preset threshold, it is marked as a candidate node with causal association. By traversing all historical time segments, the event nodes that meet the above causal association conditions (including event type, occurrence timestamp and association strength parameters) are summarized into the set of candidate causal event nodes. This set forms a preliminary causal network structure by integrating the associated events within multiple time segments, providing data support for the subsequent screening of key nodes.

[0043] When determining the key causal event nodes, based on the Granger causality test algorithm integrated in the spatio-temporal causal model, causal verification and redundancy elimination are performed on the generated set of candidate causal event nodes. Specifically, the spatio-temporal causal model quantifies the temporal causal relationship between event nodes by introducing the Granger causality test to determine whether an event significantly affects the occurrence of another event statistically. For each pair of candidate event nodes (e.g., event A and event B), first, a vector autoregressive (VAR) model is constructed, with the historical temporal data of event A as the independent variable and the future temporal data of event B as the dependent variable, and the model parameters are fitted by the least squares method. Subsequently, an F-test is used to conduct a hypothesis test on the model, and the significance level (p-value) of the temporal information of event A on the predictive ability of event B is calculated. If the p-value of the test result is lower than the preset significance level (e.g., 0.05), then event A is determined to be the Granger cause of event B, indicating that event A has a causal driving effect on event B in terms of time series; conversely, if the p-value does not meet the significance condition, then this causal association is eliminated to avoid the interference of spurious correlation. By traversing all event pairs in the candidate node set and iteratively performing the above tests and screening processes, finally, the causal association nodes with statistical significance are retained to form the key causal event nodes. This node not only includes the event type, timestamp, and correlation strength parameters but also strengthens the credibility of the causal relationship through the verification results of the Granger causality test, providing a high-confidence spatio-temporal causal reasoning basis for the construction of the dynamic causal graph.

[0044] In some instances, based on the key causal event nodes and real-time environmental semantic information, a dynamic causal graph including environmental interaction relationships is determined, including: Based on the key causal event nodes, a set of event nodes is determined; Based on the real-time environmental semantic information, the movement direction and position of the dynamic obstacle are determined; Based on the set of event nodes, the movement direction, and the position, the environmental interaction weight is determined, where the environmental interaction weight is used to characterize the association strength between the event node and the dynamic obstacle; Based on the set of event nodes and the environmental interaction weight, an initial causal graph with weights is determined; Based on a preset causal redundancy threshold, the connection edges with weights less than the causal redundancy threshold in the initial causal graph are eliminated to determine the dynamic causal graph.

[0045] Exemplarily, the key causal event nodes are screened by the spatio-temporal causal model through Granger causality tests. Each node represents an event that has a significant causal relationship with the behavior evolution of the target object (such as speed mutation, direction turn, or environmental interaction event). The event node set is formed by integrating all the key causal event nodes that pass the test. Each node contains the event type (such as "speed mutation"), the occurrence timestamp, the association strength parameter (such as the p-value of the Granger causality test), and the quantitative feature corresponding to the event (such as the mutation amplitude or the turning angle). For example, if a speed mutation event is verified as the Granger cause of a subsequent direction turn event, then this speed mutation event will be included in the event node set, and its interaction relationship with the obstacle will be marked. The determination of the event node set provides the basic elements for the construction of the dynamic causal graph.

[0046] The real-time environmental semantic information includes the real-time state of dynamic obstacles annotated in the three-dimensional environmental topology map generated by the lidar, including the position coordinates of the obstacles (such as the x, y, and z-axis positions in three-dimensional space) and the motion direction vector (such as the magnitude and angle of the velocity vector). By analyzing the motion direction annotation in the semantic information, the instantaneous motion parameters and spatial positions of each dynamic obstacle are extracted and mapped to a unified three-dimensional coordinate system. For example, the motion direction of the obstacle is represented by a vector arrow, and its position is accurately marked by rasterized coordinates, forming a real-time state description of the dynamic obstacle, which provides the input for the subsequent calculation of the environmental interaction weight.

[0047] The environmental interaction weight is used to quantify the association strength between the event node and the dynamic obstacle. The calculation rule is that for each event node, the Euclidean distance between the target object and the dynamic obstacle at the moment of its occurrence is calculated. The closer the distance, the higher the weight. Secondly, the included angle between the motion vectors of the target object and the obstacle is analyzed. If the included angle is less than the preset angle threshold, it indicates that there is a potential conflict in their motion directions, and the weight value is correspondingly increased. In addition, combined with the type of the event node (such as the weight of the environmental interaction event is higher than that of the speed mutation event), it is weighted by a preset weight distribution coefficient. Finally, the environmental interaction weight is mapped to the [0,1] interval through a normalization formula (such as the Sigmoid function) to ensure the comparability and consistency of the weight values.

[0048] The initial causal graph constructs a directed weighted graph with a set of event nodes as vertices and the environmental interaction weight as the edge weight value. Specifically, if there is a causal association between two event nodes in space-time (such as event A being the Granger cause of event B), a directed edge is drawn from event A to event B, and the environmental interaction weight is assigned to this edge. The weight value of the edge comprehensively reflects the strength of the causal relationship between events and the degree of influence of environmental interaction. For example, if event A (sudden speed change) causes event B (environmental interaction) and the two are close and have conflicting movement directions, the weight value of the edge approaches 1; conversely, if the association is weak, the weight value approaches 0. The initial causal graph is stored through an adjacency matrix or a graph structure database, supporting subsequent dynamic optimization and real-time updates.

[0049] Through a preset causal redundancy threshold, the connecting edges in the initial causal graph are screened. Specifically, all directed edges are traversed. If the weight value of a certain edge is lower than the causal redundancy threshold, it is determined as a redundant connection and removed from the graph. For example, if the edge weight from event A to event B is 0.25, which is lower than the threshold of 0.3, this edge is deleted to eliminate noise interference. This process retains the strong causal paths with high weights, ensuring that the dynamic causal graph only contains significant environmental interaction relationships. The finally generated dynamic causal graph is continuously optimized through a real-time update mechanism (such as an incremental learning framework), and can reflect the dynamic evolution of the causal association between the target object and the obstacle, providing a structured reasoning framework for subsequent event probability prediction and behavior trajectory generation.

[0050] In some instances, based on the dynamic causal graph, the future potential event types and the corresponding event occurrence probabilities are determined, including: Based on the set of event nodes in the dynamic causal graph, a set of candidate event types is determined; Based on the connection weights of each event node in the dynamic causal graph, the node activation probability of each candidate event type is determined; Based on a preset probability threshold, the node activation probabilities are screened, and the candidate event types that meet the preset probability threshold are determined as future potential event types; Based on the node activation probabilities corresponding to the future potential event types, the event occurrence probability is determined, where the event occurrence probability is positively correlated with the node activation probability.

[0051] Exemplarily, a dynamic causal graph consists of a set of event nodes (including speed mutation, direction turning, and environmental interaction events) and their weighted connecting edges. First, traverse all event nodes in the dynamic causal graph, extract their event types (such as "environmental interaction" or "speed mutation"), and classify them according to a preset event classification rule (for example, semantic division based on event labels) to form a set of candidate event types. For example, if there are multiple environmental interaction event nodes in the causal graph, they are classified as the "environmental interaction" type; if there are multiple direction turning event nodes, they are classified as the "direction turning" type. The set of candidate event types needs to cover all event types that the target object may trigger in its historical behavior to ensure the comprehensiveness of the prediction.

[0052] For each candidate event type, its node activation probability is calculated by aggregating the connection weights of relevant event nodes in the dynamic causal graph. Specifically, for a certain candidate event type (such as "environmental interaction"), traverse all event nodes belonging to this type in the causal graph, and extract the weight values of all its incoming edges (that is, the causal connection edges pointing to this node). The comprehensive activation probability of this event type is obtained through weighted summation or mean calculation (for example, normalization using the Softmax function). For example, if an environmental interaction event node has two incoming edges with weights of 0.8 and 0.6 respectively, its activation probability can be calculated as (0.8 + 0.6) / 2 = 0.7. This process reflects the cumulative impact of historical causal associations on the future occurrence of event types.

[0053] The node activation probabilities of candidate event types are screened through a preset probability threshold. If the activation probability of an event type exceeds the threshold, it is determined as a future potential event type; otherwise, it is excluded. For example, if the activation probability of the "environmental interaction" type is 0.7 (exceeding the threshold of 0.5), it is retained as a potential event type; if the activation probability of the "speed mutation" type is 0.4 (lower than the threshold), it is excluded. This step ensures that only high-confidence event types are retained, avoiding interference from low-probability noise on the prediction results. The set of screened potential event types will be used as the input for subsequent probability quantification.

[0054] According to the screened future potential event types and their corresponding node activation probabilities, a quantitative expression of the event occurrence probability is generated. The event occurrence probability is positively correlated with the node activation probability, which is specifically realized through a linear mapping or a probability distribution model (such as the Poisson distribution). For example, if the activation probability of the "environmental interaction" type is 0.7, its occurrence probability is preset as the normalized result of the activation probability (such as 0.7 / sum of total activation probabilities), or directly use the activation probability as the event occurrence probability. The finally generated probability value will be used to define subsequent event triggering constraint conditions to ensure that the prediction results not only conform to the strength of causal associations but also meet the feasibility requirements of the actual scenario.

[0055] In some instances, based on the event occurrence probability, event trigger constraint conditions are determined, including: Based on the event occurrence probability, probability distribution parameters of each future potential event type are determined; Based on a preset trigger rule and the probability distribution parameters, the time window and spatial region for event triggering are determined; Based on the time window and spatial region, dynamic threshold conditions for event triggering are determined, where the dynamic threshold conditions include a trigger time interval, a trigger spatial boundary, and a minimum probability threshold; Based on the dynamic threshold conditions and the event occurrence probability, event trigger constraint conditions are determined.

[0056] Exemplarily, for the filtered future potential event types (such as "environmental interaction" or "direction change"), distribution characteristics of their event occurrence probabilities are extracted, including parameters such as mean, variance, and confidence interval. For example, if the occurrence probability of a certain event type is 0.7, its mean can directly take the probability value, and the variance reflects the fluctuation range of this probability in historical data. For the confidence interval, the upper and lower bounds of the probability value are calculated through statistical methods (such as Bootstrap sampling). For example, at a 95% confidence level, the probability interval is [0.65, 0.75]. These parameters quantify the stability and reliability of event occurrence and provide a basis for formulating subsequent trigger rules.

[0057] According to the preset trigger rules (such as safety thresholds, priority strategies), the probability distribution parameters are mapped to spatio-temporal constraints for event triggering. For example, if the preset rule requires high-probability events to be responded to within a short time, a shorter time window (such as the next 2 seconds) is allocated according to the probability mean (such as 0.7); at the same time, the accuracy range of the spatial region is determined based on the upper limit of the confidence interval (such as 0.75) (such as a trigger region within a radius of 1 meter). For low-priority events, the time window may be relaxed (such as the next 5 seconds) and the spatial region may be expanded (such as a radius of 3 meters). By fusing the probability parameters and the trigger rules, a preliminary time window and spatial region are generated to ensure the adaptability of the prediction results to the real-time environmental dynamics.

[0058] Combining the time window, the spatial region, and a preset minimum probability threshold (such as 0.6), multi-dimensional dynamic threshold conditions are constructed, including a trigger time interval, a trigger spatial boundary, and a minimum probability threshold; the trigger time interval limits that the event is only valid within the time window; the trigger spatial boundary defines the physical range where the event occurs, for example, the event is triggered only when the target object enters the spatial region; the minimum probability threshold filters out event types with occurrence probabilities higher than the threshold, for example, only events with a probability ≥ 0.6 are retained. The dynamic threshold conditions combine the above dimensions through a logical AND relationship to ensure the strictness and operability of the trigger conditions.

[0059] Match the dynamic threshold conditions with the event occurrence probabilities to generate the final event trigger constraint conditions. Specifically, traverse all future potential event types. If the probability distribution parameter of an event (such as the mean value of 0.7) meets the dynamic threshold conditions (such as time window, spatial region, and minimum probability threshold), then include it in the trigger constraint set. For example, if the occurrence probability of an environmental interaction event is 0.7, within the time window, and within the preset spatial region, then it is determined to meet the trigger conditions, and its trigger priority is recorded (such as sorted according to the probability value). The finally generated trigger constraint conditions are stored in structured data (such as JSON or XML format), including event type, trigger time, spatial range, and probability threshold, providing a clear trigger logic boundary for subsequent behavior trajectory prediction.

[0060] In some instances, based on the incremental learning framework and the real-time input multi-modal historical behavior data, determine the node weight update parameters of the dynamic causal graph, including: Based on the real-time input multi-modal historical behavior data, determine the incremental input data set; Based on the distribution difference detection algorithm in the incremental learning framework, determine the distribution difference value between the incremental input data set and the historical data set; Based on the preset distribution difference threshold and the distribution difference value, determine the model update trigger signal; Based on the model update trigger signal and the incremental input data set, determine the knowledge transfer parameters of the online knowledge distillation model; Based on the knowledge transfer parameters and the current node weights of the dynamic causal graph, determine the node weight update parameters, which are used to adjust the association strength between the event nodes and the dynamic obstacles in the dynamic causal graph.

[0061] Exemplarily, the incremental input data set is generated after preprocessing and alignment of the real-time input multi-modal historical behavior data (including the motion parameters of the millimeter-wave radar, the contour features of the vision sensor, and the three-dimensional environmental topology data of the lidar). Specifically, through the preset timestamp synchronization rule, align the real-time collected motion speed, direction angle, contour features, and environmental topology information along the unified time axis to eliminate the timing misalignment caused by the sensor sampling frequency difference. For the missing data segments, use the linear interpolation or sliding window filling algorithm to fill them to ensure the temporal integrity of the incremental data set. The finally generated incremental input data set contains the behavior characteristics and environmental interaction states of the target object within the latest time window, providing input for subsequent distribution difference detection.

[0062] In the incremental learning framework, the Kullback-Leibler (KL) divergence or Jensen-Shannon (JS) divergence algorithm is adopted to quantify the distribution difference between the incremental input dataset and the historical dataset. Specifically, the probability distributions of the historical dataset and the incremental dataset in key feature dimensions (such as speed mutation frequency, direction angle change rate, obstacle interaction distance) are extracted, and the distribution difference value is obtained through divergence calculation. For example, if the mean value of speed mutation in the historical dataset is 1.5, while that in the incremental dataset is 2.0, the KL divergence will reflect this offset. The distribution difference value is mapped to the [0, 1] interval through normalization processing to determine whether the data distribution has changed significantly.

[0063] A preset distribution difference threshold (such as 0.3) is set. When the calculated distribution difference value exceeds this threshold, it is determined that there is a significant offset between the incremental data and the historical data distribution, and the model update mechanism is triggered. For example, if the distribution difference value is 0.45 (exceeding the threshold of 0.3), an update trigger signal is generated; conversely, if the difference value is 0.2, the current model is maintained. The trigger signal is transmitted to the subsequent processing module through a binary flag (such as "1" for update and "0" for maintenance) or a priority parameter (such as the larger the difference value, the higher the update priority), ensuring that the model only starts the update process when necessary and avoiding stability problems caused by frequent adjustments.

[0064] When the model update trigger signal takes effect, an online knowledge distillation model is used to extract new knowledge from the incremental dataset. Specifically, the incremental data is input into the teacher model (such as a pre-trained deep neural network) to generate soft labels or feature embedding vectors as the representation of new knowledge. At the same time, the student model (i.e., the dynamic causal graph) calculates the knowledge transfer loss (such as cross-entropy loss or mean square error) by comparing the output of the teacher model with its own prediction results. The student model optimizes the loss function through the gradient descent algorithm to obtain knowledge transfer parameters (such as weight adjustment coefficients or bias terms) to quantify the correction intensity of new data on the original causal relationship. For example, if the incremental data shows a significant increase in the obstacle interaction frequency, the transfer parameter will increase the weight coefficient of the environmental interaction event.

[0065] Fuse the knowledge transfer parameter with the current node weight of the dynamic causal graph, and recalculate the association strength between the event node and the dynamic obstacle. Specifically, for each event node, the formula for its updated weight is: new weight = original weight × (1 - decay coefficient) + transfer parameter × learning rate; where the decay coefficient (such as 0.1) is used to control the retention ratio of the historical weight, and the learning rate (such as 0.05) adjusts the incorporation speed of new knowledge. For example, if the original weight of an environmental interaction node is 0.8 and the transfer parameter is 0.3, then the new weight is 0.8×0.9 + 0.3×0.05 = 0.735. This process locally adjusts the node weights without destroying the original causal graph topology, enabling the model to adapt to the dynamic changes of the environment. The finally generated node weight update parameters are written into the dynamic causal graph through the graph database or the adjacency matrix update interface to complete the real-time optimization of the model.

[0066] In some instances, based on the event-triggering constraint conditions and the updated dynamic causal graph, determine the predicted behavior trajectory of the target object, including: Based on the time window and the spatial region, determine multiple candidate trajectory branches; Based on the node weights in the updated dynamic causal graph, determine the posterior probability of each candidate trajectory branch; Based on the maximum value of the posterior probability, determine the main predicted trajectory; Based on the main predicted trajectory and the position of the dynamic obstacle in the real-time environmental semantic information, determine the trajectory conflict probability; Based on the trajectory conflict probability and the preset conflict threshold, determine the trajectory correction vector; Based on the trajectory correction vector and the preset smoothing algorithm, determine the predicted behavior trajectory of the target object.

[0067] Exemplarily, First, according to the time window and the spatial region defined in the event-triggering constraint conditions, generate multiple candidate trajectory branches. The generation of candidate trajectories depends on the historical behavior pattern of the target object and the causal relevance of the event nodes in the dynamic causal graph. For example, based on the current movement speed and direction angle data of the target object, deduce its possible paths within the time window through a kinematic model (such as a uniform motion model or an acceleration model), and combine with the obstacle distribution in the environmental semantic information to exclude invalid paths that overlap with static obstacles. Each candidate trajectory branch is represented parametrically, recording its time and spatial coordinate sequences, forming a trajectory set covering the potential movement directions of the target object.

[0068] For each candidate trajectory branch, calculate its matching degree with historical causal associations based on the weight distribution of event nodes in the updated dynamic causal graph. Specifically, traverse the spatio-temporal region passed by the trajectory branch, and extract the event nodes and their connection weights in the dynamic causal graph within this region. For example, if a trajectory passes through an environmental interaction event node (weight 0.8) and a speed mutation event node (weight 0.6), then calculate the posterior probability of this trajectory through weighted summation or Bayesian network inference. The quantization formula for the posterior probability is: where is the posterior probability of the th candidate trajectory branch, that is, the likelihood of this trajectory occurring in the future, quantifying the matching degree between the trajectory and historical causal associations and the intensity of environmental interactions. The higher the probability value, the more consistent the trajectory is with the evolution law of the dynamic causal graph; perform a summation operation on all event nodes in the dynamic causal graph to comprehensively consider the influence of all relevant event nodes on the current trajectory and reflect the cumulative effect of causal associations; is the weight of the th event node, which comes from the updated connection edge weights in the dynamic causal graph, characterizing the correlation strength between the event node and the behavior evolution of the target object. The higher the weight value, the more significant the causal impact of this event on the trajectory; is used to determine whether the candidate trajectory covers the spatio-temporal range of the event node . If the spatio-temporal path of the trajectory intersects with the spatio-temporal range of the event , it is 1, otherwise it is 0, screening the causal event nodes related to the trajectory. Only when the trajectory passes through the spatio-temporal region of the event node, the weight of this event will be included in the posterior probability.

[0069] By comparing the posterior probability values of all candidate trajectory branches, select the trajectory with the highest probability value as the main predicted trajectory. For example, if the posterior probability of candidate trajectory A is 0.85 and that of trajectory B is 0.72, then trajectory A is determined as the main predicted trajectory. The spatio-temporal coordinate sequence of the main predicted trajectory represents the most likely behavior path of the target object within the future time window, including speed, direction, and position change trends. At the same time, record the probability value of the main predicted trajectory and the information of associated event nodes, providing input for subsequent conflict detection.

[0070] Based on the positions of dynamic obstacles (such as three-dimensional coordinates and motion directions) in the real-time environmental semantic information, calculate the conflict probability between the main predicted trajectory and the obstacles. Specifically, through a spatial geometric model (such as bounding box detection or distance field analysis), determine whether there is an intersection between the main predicted trajectory and the predicted path of the obstacles within the time window. For example, if the obstacle moves at a speed Moving in the direction of the main predicted trajectory, the relative distance to the target will decrease over time, and the conflict risk should increase with the speed. The conflict probability calculated by the time-to-collision formula is expressed as: Where is a preset safety distance threshold, representing the minimum safe distance allowed between the target object and the obstacle; is the Euclidean distance calculated in real time, that is, the actual distance between the target object and the obstacle; is a regulation coefficient, controlling the slope of the Sigmoid function and determining the sensitivity of the probability to the distance change ( when, the influence of the distance difference on the probability is more significant); is the moving speed of the obstacle relative to the target object; is the prediction time window; The above formula calculates the possible reduced distance of the obstacle within time ; The corrected actual distance is (if the obstacle is approaching, then is a positive value, reducing the distance); when the corrected actual distance is less than or equal to the preset safety distance threshold , the conflict probability increases significantly.

[0071] When the trajectory conflict probability exceeds the preset threshold (0.2), a trajectory correction vector is generated through an optimization algorithm (such as gradient descent or geometric projection). The direction of the correction vector is determined by the relative relationship between the obstacle position and the main predicted trajectory: if the obstacle is on the left side of the trajectory, a lateral offset vector to the right is generated to avoid the conflict. If it is in front, a deceleration or longitudinal avoidance vector is generated. The magnitude of the correction vector is controlled by the difference ratio between the conflict probability and the preset threshold. For example, when the conflict probability is 0.8, the magnitude of the correction vector is (0.8 - 0.2) / 0.2 × the maximum correction amount. Through vector superposition, the correction vector is applied to the path points of the main predicted trajectory to generate an adjusted candidate trajectory.

[0072] The corrected trajectory may have problems such as path mutation or discontinuity. The corrected trajectory is optimized through a preset smoothing algorithm (such as Kalman filtering or spline interpolation) to eliminate path mutations and enhance trajectory continuity. For example, the cubic spline interpolation algorithm is used, with the corrected key path points as control points, to generate a smooth trajectory curve. The smoothed trajectory needs to simultaneously satisfy maintaining a safe distance from dynamic obstacles within the time window; conforming to the kinematic characteristics of the target object (such as speed and acceleration limits); and maintaining a causal association with the high-weight event nodes in the dynamic causal graph. The finally output predicted behavior trajectory is characterized by a spatio-temporal coordinate sequence, providing a highly reliable behavior prediction result for autonomous driving or robot navigation.

[0073] Please refer to Figure 2, which is a schematic structural diagram of an object behavior prediction device provided by an embodiment of the present application, including: A target object data acquisition unit 21, configured to acquire multi-modal historical behavior data and real-time environmental semantic information of a target object; A key event node determination unit 22, configured to determine key causal event nodes based on multi-modal historical behavior data and a spatio-temporal causal model; A dynamic causal graph generation unit 23, configured to determine a dynamic causal graph including environmental interaction relationships based on the key causal event nodes and real-time environmental semantic information; A future potential event determination unit 24, configured to determine future potential event types and corresponding event occurrence probabilities based on the dynamic causal graph; An event constraint condition determination unit 25, configured to determine event trigger constraint conditions based on event occurrence probabilities; A node weight parameter update unit 26, configured to determine node weight update parameters of the dynamic causal graph based on an incremental learning framework and real-time input multi-modal historical behavior data; A target object behavior prediction unit 27, configured to determine a predicted behavior trajectory of the target object based on the event trigger constraint conditions and the updated dynamic causal graph.

[0074] Please refer to Figure 3 , an embodiment of the present application further provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, the steps of any method for target behavior prediction are implemented.

[0075] Since the electronic device introduced in this embodiment is the device used to implement an object behavior prediction device in an embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiment of the present application belongs to the scope protected by the present application.

[0076] In a specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0077] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0078] Those skilled in the art should understand that the embodiments of the present application may provide a method, a system or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program code.

[0079] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0082] The embodiments of the present application also provide a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a target behavior prediction method in the corresponding embodiment.

[0083] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, they wholly or partly generate a process or function in accordance with an embodiment of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium, etc.

[0084] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0085] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, indirect couplings or communication connections of devices or units, and may be in electrical, mechanical, or other forms.

[0086] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0087] In addition, the functional units in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated units may be implemented in the form of hardware and / or software functional units.

[0088] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods in various embodiments of this application.

[0089] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

[0090] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0091] Obviously, those skilled in the art can make various changes and deformations to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and deformations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification also intends to include these modifications and deformations.

Claims

1. A target behavior prediction method, characterized in that: include: Obtain the multimodal historical behavior data and real-time environmental semantic information of the target object; Determining key causal event nodes based on the multimodal historical behavior data and the spatiotemporal causal model; Determining a dynamic causal graph including environmental interaction relationships based on the key causal event nodes and the real-time environmental semantic information; Based on the dynamic causal graph, determining the types of potential future events and the corresponding probability of occurrence of the events; Based on the probability of occurrence of the event, determining event triggering constraint conditions; Determining node weight update parameters of the dynamic causal graph based on an incremental learning framework and multimodal historical behavior data input in real time; Based on the event triggering constraint conditions and the updated dynamic causal graph, a predicted behavior trajectory of the target object is determined.

2. The method according to claim 1, characterized in that The step of obtaining the multimodal historical behavior data and real-time environment semantic information of the target object includes: Determine the moving speed and angular direction data of the target object based on the detection signal of the millimeter wave radar; Determine the contour features and local environment image of the target object based on the image acquisition results of the visual sensor; Determine a three-dimensional environment topology map based on the point cloud data of the laser radar, wherein the three-dimensional environment topology map includes movement direction annotation information of dynamic obstacles; Based on a preset timestamp synchronization rule, the movement speed, the directional angle data, the contour features, the local environment image and the three-dimensional environment topology map are aligned to determine multimodal historical behavior data; Based on the movement direction annotation information, real-time environmental semantic information is determined.

3. The method according to claim 1, characterized in that Determining key causal event nodes based on the multimodal historical behavior data and the spatiotemporal causal model includes: Determining a plurality of historical time segments based on the time series characteristics of the multimodal historical behavior data; Based on the event detection rules in the spatiotemporal causal model, feature extraction is performed on each of the historical time segments to determine the speed mutation point, direction turning point and environmental interaction event; Based on a preset causal association threshold, the speed mutation point, the direction turning point and the environment interaction event are marked to determine a set of candidate causal event nodes; Based on the Granger causality test algorithm in the spatiotemporal causal model, the candidate causal event node set is eliminated by association to determine the key causal event nodes.

4. The method according to claim 1, characterized in that: The step of determining a dynamic causal graph including an environmental interaction relationship based on the key causal event nodes and the real-time environmental semantic information includes: Based on the key causal event nodes, determining an event node set; Based on the real-time environment semantic information, determine the movement direction and position of the dynamic obstacle; Determining an environment interaction weight based on the event node set, the movement direction and the position, wherein the environment interaction weight is used to characterize the association strength between the event node and the dynamic obstacle; Determining an initial causal graph with weights based on the event node set and the environment interaction weights; Based on a preset causal redundancy threshold, connecting edges whose weights are less than the causal redundancy threshold in the initial causal graph are removed to determine a dynamic causal graph.

5. The method according to claim 1, characterized in that Determining the future potential event types and corresponding event occurrence probabilities based on the dynamic causal graph includes: Determine a set of candidate event types based on a set of event nodes in the dynamic causal graph; Determining a node activation probability for each candidate event type based on a connection weight of each event node in the dynamic causal graph; Based on a preset probability threshold, the node activation probability is screened to determine the candidate event type that meets the preset probability threshold as a future potential event type; Based on the node activation probability corresponding to the future potential event type, the event occurrence probability is determined, wherein the event occurrence probability is positively correlated with the node activation probability.

6. The method according to claim 1, characterized in that The determining of event triggering constraint conditions based on the event occurrence probability includes: Determining probability distribution parameters for each future potential event type based on the probability of occurrence of the event; Based on the preset triggering rules and the probability distribution parameters, determine the time window and spatial area for event triggering; Based on the time window and the spatial region, determining a dynamic threshold condition for event triggering, wherein the dynamic threshold condition includes a triggering time interval, a triggering spatial boundary, and a minimum probability threshold; Based on the dynamic threshold condition and the event occurrence probability, an event triggering constraint condition is determined.

7. The method according to claim 6, characterized in that The step of determining the node weight update parameters of the dynamic causal graph based on the incremental learning framework and the multimodal historical behavior data input in real time includes: Determine the incremental input data set based on the multimodal historical behavior data input in real time; Determine the distribution difference value between the incremental input data set and the historical data set based on the distribution difference detection algorithm in the incremental learning framework; Determining a model update trigger signal based on a preset distribution difference threshold and the distribution difference value; Determining knowledge migration parameters of an online knowledge distillation model based on the model update trigger signal and the incremental input data set; Based on the knowledge transfer parameters and the current node weights of the dynamic causal graph, node weight update parameters are determined, wherein the node weight update parameters are used to adjust the association strength between event nodes and dynamic obstacles in the dynamic causal graph.

8. The method according to claim 7, characterized in that The step of determining the predicted behavior trajectory of the target object based on the event triggering constraint condition and the updated dynamic causal graph includes: Based on the time window and the spatial region, determining a plurality of candidate trajectory branches; Determining the posterior probability of each candidate trajectory branch based on the node weights in the updated dynamic causal graph; Determining a main prediction trajectory based on the maximum value of the posterior probability; Determining a trajectory conflict probability based on the main predicted trajectory and the dynamic obstacle position in the real-time environmental semantic information; Determining a trajectory correction vector based on the trajectory conflict probability and a preset conflict threshold; Based on the trajectory correction vector and a preset smoothing algorithm, a predicted behavior trajectory of the target object is determined.

9. A target behavior prediction device, characterized in that: include: A target object data acquisition unit, used to acquire multimodal historical behavior data and real-time environmental semantic information of the target object; A key event node determination unit, which determines key causal event nodes based on the multimodal historical behavior data and the spatiotemporal causal model; A dynamic causal graph generating unit, which determines a dynamic causal graph including environmental interaction relationships based on the key causal event nodes and the real-time environmental semantic information; A future potential event determination unit, which determines the type of future potential events and the corresponding probability of occurrence of the events based on the dynamic causal graph; An event constraint determination unit, which determines an event trigger constraint based on the event occurrence probability; A node weight parameter updating unit, which determines the node weight update parameters of the dynamic causal graph based on an incremental learning framework and multimodal historical behavior data input in real time; The target object behavior prediction unit determines the predicted behavior trajectory of the target object based on the event triggering constraint condition and the updated dynamic causal graph.

10. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the target behavior prediction method as described in any one of claims 1 to 8 when executing the computer program stored in the memory.

Citation Information

Patent Citations

  • Vehicle track prediction modeling method and device based on dynamic space-time interaction diagram

    CN116595871A

  • Mine electric locomotive unmanned comprehensive control method based on artificial intelligence

    CN119840612A

  • Method and Apparatus for Predicting Motion Track of Obstacle and Autonomous Vehicle

    US20230060005A1

Cited By

  • Smart home control method and system based on multi-modal fusion and deep learning

    CN120630745A

  • Dynamic environment adaptive control method and device, equipment, medium and program product

    CN120803276A

  • Drought and flood sharp turning key factor prediction method and system based on causal inference

    CN120951063A

  • Visible light infrared image assisted vehicle-mounted laser point cloud semantic segmentation method and system

    CN121033411A

  • Historical debugging-oriented multi-thread race condition accurate reproduction method

    CN121349841A