Car following reminding method based on multi-modal perception and behavior prediction, vehicle and medium

By using multimodal perception and behavior prediction methods, the system acquires image sequences of the vehicle in front and the real-time status of the current vehicle, predicts the starting intention of the vehicle in front, solves the problem of following delay caused by brief driver distraction, and achieves accurate and timely following reminders, thereby improving traffic efficiency and order.

CN122035040BActive Publication Date: 2026-06-19JIQI INTELLIGENT TECHNOLOGY (TIANJIN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIQI INTELLIGENT TECHNOLOGY (TIANJIN) CO LTD
Filing Date
2026-04-17
Publication Date
2026-06-19

Smart Images

  • Figure CN122035040B_ABST
    Figure CN122035040B_ABST
Patent Text Reader

Abstract

This application provides a following warning method, vehicle, and medium based on multimodal perception and behavior prediction, relating to the fields of intelligent transportation and automotive assisted driving technology. The method includes: acquiring information about the vehicle's forward environment and its real-time motion state; extracting visual features from a sequence of images of the preceding vehicle to obtain its visual micro-behavioral features; fusing features using a pre-defined multimodal feature fusion network based on the preceding vehicle's visual micro-behavioral features and the vehicle's real-time motion state to obtain the preceding vehicle's dynamic features relative to the vehicle; predicting the preceding vehicle's starting behavior using a pre-defined behavior prediction network based on the preceding vehicle's visual micro-behavioral features and dynamic features, to obtain the preceding vehicle's starting behavior prediction result; and generating following warning information for the vehicle based on the preceding vehicle's starting behavior prediction result. This method can accurately, timely, and reliably determine the preceding vehicle's starting behavior prediction result, thereby making the generated following warning information for the vehicle more reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of intelligent transportation and automotive assisted driving technology, and more specifically, to a following warning method, vehicle, and medium based on multimodal perception and behavior prediction. Background Technology

[0002] In urban driving scenarios with frequent stop-and-go traffic, it's extremely common for drivers to delay following other vehicles due to brief moments of inattention. This not only reduces the efficiency of individual lanes and exacerbates overall congestion, but can also increase driver stress due to honking from vehicles behind, and even lead to "cutting in," disrupting traffic order. Therefore, providing drivers with following reminders is becoming increasingly important.

[0003] In related technologies, the determination of whether the vehicle in front has started is achieved by comparing changes in the outline area of ​​the vehicle in front in consecutive image frames. However, this method is susceptible to changes in lighting, vehicle shadows, interference from other vehicles, or camera shake. Furthermore, it only issues an alarm after a significant vehicle movement, resulting in an inherent delay and a poor user experience. In short, these technologies suffer from unreliable following alerts. Summary of the Invention

[0004] The purpose of this application is to address the shortcomings of the prior art by providing a following warning method, vehicle, and medium based on multimodal perception and behavior prediction, so as to solve the aforementioned technical problems in the related technologies.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, embodiments of this application provide a following alert method based on multimodal perception and behavior prediction, including:

[0007] Acquire information about the environment ahead of the vehicle and the real-time motion status of the vehicle, wherein the information about the environment ahead includes: a sequence of images of the vehicle in front;

[0008] Visual features are extracted from the preceding vehicle image sequence to obtain the visual micro-behavioral features of the preceding vehicle;

[0009] Based on the visual micro-behavioral features of the preceding vehicle and the real-time motion state of the current vehicle, a preset multimodal feature fusion network is used to perform feature fusion to obtain the dynamic features of the preceding vehicle relative to the current vehicle.

[0010] Based on the visual micro-behavioral features and dynamic features of the preceding vehicle, a preset behavior prediction network is used to predict the starting behavior of the preceding vehicle, and the starting behavior prediction result of the preceding vehicle is obtained. The starting behavior prediction result of the preceding vehicle includes: the probability of the starting behavior intention of the preceding vehicle and the prediction time window.

[0011] Based on the predicted starting behavior of the vehicle in front, a following reminder message for the current vehicle is generated.

[0012] Optionally, the preceding vehicle's visual micro-behavioral features include: the preceding vehicle's physical dynamics features and the preceding vehicle's semantic micro-behavioral features; the step of extracting visual features from the preceding vehicle's image sequence to obtain the preceding vehicle's visual micro-behavioral features includes:

[0013] Physical dynamic features of the preceding vehicle are extracted from the preceding vehicle image sequence to obtain the preceding vehicle physical dynamic features; wherein, the preceding vehicle physical dynamic features include at least one of the following: the centroidal axial acceleration of the preceding vehicle's bounding box, the dense optical flow vector field at the bottom edge corner of the preceding vehicle, and the divergence.

[0014] Semantic micro-behavioral features are extracted from the preceding vehicle image sequence to obtain the preceding vehicle semantic micro-behavioral features; wherein, the preceding vehicle semantic micro-behavioral features include at least one of the following: the preceding vehicle taillight state encoding, the preceding vehicle wheel region texture motion energy spectrum, and the preceding vehicle attitude micro-change features.

[0015] Optionally, the preset behavior prediction network includes: a temporal causal prediction network, a transient pattern prediction network, and a dual-path fusion network;

[0016] The step of predicting the starting behavior of the preceding vehicle using a preset behavior prediction network based on the visual micro-behavioral features and dynamic features of the preceding vehicle, and obtaining the prediction result of the starting behavior of the preceding vehicle, includes:

[0017] Using the aforementioned temporal causal prediction network, the temporal dependencies of vehicle state evolution are determined based on the visual micro-behavioral features and dynamic features of the preceding vehicle.

[0018] Using the instantaneous pattern prediction network, feature information of the preceding vehicle in the instant before its starting intention is generated is determined based on the preceding vehicle's visual micro-behavioral features and the preceding vehicle's dynamic features;

[0019] The dual-path fusion network is used to fuse the temporal dependencies of the vehicle state evolution and the feature information of the instant before the preceding vehicle's starting intention, to generate the probability of the preceding vehicle's starting behavior intention and the prediction time window.

[0020] Optionally, generating the following warning information for the current vehicle based on the prediction result of the preceding vehicle's starting behavior includes:

[0021] Based on the preceding vehicle image sequence and current weather information, obtain the current environmental confidence level;

[0022] The driver's attention level of the vehicle is collected based on the driver status monitoring system of the vehicle.

[0023] Based on the current environment confidence level and the driver attention level, a dynamic start-up decision threshold for the vehicle is generated;

[0024] If the probability of the preceding vehicle's intention to start is greater than the dynamic start decision threshold, and the preceding vehicle starts within the prediction time window at the current moment;

[0025] The following alert information is generated based on the preceding vehicle's start-up event and the driver's attention level.

[0026] Optionally, generating the following warning information based on the preceding vehicle's start-up event and the driver's attention level includes:

[0027] Calculate the overall confidence level based on the current weather information, the driver's attention level, and the probability of the preceding vehicle's intention to start.

[0028] A reinforcement learning strategy based on finite state machines is adopted to output a reminder pattern matrix based on the preceding vehicle's start-up event, the overall confidence level, the driver's attention level, and the rate of change of the driver's attention level of the current vehicle.

[0029] The following vehicle reminder information is determined based on the driver's status, scene complexity, and time confidence in the reminder mode matrix.

[0030] Optionally, before generating the following warning information for the current vehicle based on the prediction result of the preceding vehicle's starting behavior, the method further includes:

[0031] Obtain the relative distance sequence between the vehicle and the vehicle in front;

[0032] Calculate the speed of the preceding vehicle relative to this vehicle based on the relative distance sequence;

[0033] Calculate the vehicle's acceleration based on its motion state;

[0034] Update the probability of the preceding vehicle's starting behavior intention based on the speed of the preceding vehicle relative to the current vehicle and the acceleration of the current vehicle;

[0035] And / or,

[0036] The probability of the preceding vehicle's intention to start is updated based on the preceding vehicle's start signal broadcast by the preceding vehicle.

[0037] Optionally, before extracting visual features from the preceding vehicle image sequence to obtain the preceding vehicle's visual micro-behavioral features, the method further includes:

[0038] If the vehicle's speed is consistently lower than a preset speed threshold during real-time motion, and the preceding vehicle is continuously detected in the preceding vehicle image sequence, then the congestion following monitoring mode is activated.

[0039] Optionally, the method further includes:

[0040] A warning signal is generated based on the physical and dynamic characteristics of the vehicle in front.

[0041] Secondly, embodiments of this application provide a vehicle, including: a vehicle body and a processing module, wherein the processing module is disposed in the vehicle body, and the processing module executes a computer program to implement the following warning method based on multimodal perception and behavior prediction as described in any of the first aspects above.

[0042] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when read and executed, implements the following alert method based on multimodal perception and behavior prediction as described in any of the first aspects above.

[0043] The beneficial effects of this application are as follows: This application provides a following warning method based on multimodal perception and behavior prediction. The method includes: acquiring the front environment information of the vehicle and the real-time motion state of the vehicle, wherein the front environment information includes: a sequence of images of the preceding vehicle; extracting visual features from the sequence of images of the preceding vehicle to obtain the visual micro-behavioral features of the preceding vehicle; performing feature fusion using a preset multimodal feature fusion network based on the visual micro-behavioral features of the preceding vehicle and the real-time motion state of the vehicle to obtain the dynamic features of the preceding vehicle relative to the vehicle; predicting the starting behavior of the preceding vehicle using a preset behavior prediction network based on the visual micro-behavioral features and the dynamic features of the preceding vehicle to obtain the starting behavior prediction result of the preceding vehicle, wherein the starting behavior prediction result of the preceding vehicle includes: the probability of the starting behavior intention of the preceding vehicle and the prediction time window; and generating following warning information for the vehicle based on the starting behavior prediction result of the preceding vehicle. By acquiring multimodal perception information, such as the frontal environment information and the real-time motion status of the vehicle, and extracting the visual micro-behavioral features of the preceding vehicle from the image sequence, rich and comprehensive features of the preceding vehicle can be obtained. Furthermore, by combining the real-time motion status of the vehicle itself, the dynamic features of the preceding vehicle relative to the vehicle are determined, making the determined dynamic features of the preceding vehicle more accurate. Through detailed analysis of multimodal perception information, this embodiment of the application can accurately, timely, and reliably determine the prediction results of the starting behavior of the preceding vehicle, thereby making the generated following warning information of the vehicle more reliable. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A schematic diagram of the structure of a vehicle provided in an embodiment of this application;

[0046] Figure 2 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 1 ;

[0047] Figure 3 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 2 ;

[0048] Figure 4 This is a schematic diagram of the structure of a preset behavior prediction network provided in an embodiment of this application;

[0049] Figure 5 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 3 ;

[0050] Figure 6 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 4 ;

[0051] Figure 7 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 5 ;

[0052] Figure 8 A schematic diagram of a framework for generating following vehicle alert information is provided in an embodiment of this application;

[0053] Figure 9 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 6 ;

[0054] Figure 10 A schematic diagram of the structure of a following warning device based on multimodal perception and behavior prediction provided in an embodiment of this application;

[0055] Figure 11 This is a schematic diagram of the structure of a processing module provided in an embodiment of this application. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of this application, but not all embodiments.

[0057] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0058] In the description of this application, it should be noted that if the terms "upper", "lower", etc. appear to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship that the product of this application is usually placed in, it is only for the convenience of describing this application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0059] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0060] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.

[0061] Figure 1 This application provides a schematic diagram of the structure of a vehicle, as shown in the embodiment of the present application. Figure 1 As shown, the vehicle includes a vehicle body and a processing module, with the processing module housed within the vehicle body. The processing module includes a microprocessor and an algorithm program stored therein.

[0062] like Figure 1As shown, the vehicle may also include an inertial measurement unit (IMU) and a visual perception unit. The processing module is connected to the inertial measurement unit and the visual perception unit respectively. The inertial measurement unit is used to collect and send the real-time motion status of the vehicle to the processing module, and the visual perception unit is used to collect and send the image sequence of the preceding vehicle to the processing module.

[0063] This application also provides a following alert method based on multimodal perception and behavior prediction, applied to the processing module in the aforementioned vehicle. The following explains the following following description of the following following alert method based on multimodal perception and behavior prediction provided by this application.

[0064] Figure 2 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 1 ,like Figure 2 As shown, the method may include:

[0065] S101. Obtain information about the environment ahead of the vehicle and the real-time motion status of the vehicle.

[0066] The information about the environment ahead includes: a sequence of images of the vehicle in front.

[0067] In some implementations, a visual perception unit is used to acquire a sequence of images of the vehicle ahead and send the sequence of images to the processing module; an inertial measurement unit is used to acquire the real-time motion state of the vehicle and send the real-time motion state of the vehicle to the processing module. Accordingly, the processing module can receive the sequence of images of the vehicle ahead and the real-time motion state of the vehicle.

[0068] It should be noted that the visual perception unit is a front-facing monocular camera, which captures image sequences of the vehicle in front in a low-power mode.

[0069] S102. Visual features are extracted from the image sequence of the preceding vehicle to obtain the visual micro-behavioral features of the preceding vehicle.

[0070] In this embodiment of the application, a spatiotemporal synchronization protocol can be initiated to perform hardware timestamp alignment and coordinate system unification of the preceding vehicle image sequence and the real-time motion state of the current vehicle. Then, visual features are extracted from the preceding vehicle image sequence at different levels, making the visual and behavioral features of the preceding vehicle richer and more comprehensive.

[0071] S103. Based on the visual micro-behavioral features of the preceding vehicle and the real-time motion state of this vehicle, a preset multimodal feature fusion network is used to perform feature fusion to obtain the dynamic features of the preceding vehicle relative to this vehicle.

[0072] The real-time motion state of the vehicle is used to characterize the micro-vibration spectrum features of the vehicle.

[0073] In some implementations, the visual micro-behavioral features of the preceding vehicle and the micro-vibration spectrum features of the current vehicle are input into a multimodal feature fusion network based on channel attention. This multimodal feature fusion network dynamically evaluates the reliability weight of each feature channel in the current environment to generate a unified and robust dynamic feature of the preceding vehicle relative to the current vehicle.

[0074] In addition, the dynamic characteristics of the vehicle in front relative to the vehicle can also be called the environment-vehicle dynamic representation vector, which is used to represent the relative speed and positional relationship between the vehicle in front and the vehicle itself, as well as whether there is a following action.

[0075] S104. Based on the visual micro-behavioral characteristics and dynamic characteristics of the preceding vehicle, a preset behavior prediction network is used to predict the starting behavior of the preceding vehicle, and the prediction result of the starting behavior of the preceding vehicle is obtained.

[0076] The prediction results for the preceding vehicle's starting behavior include: the probability of the preceding vehicle's starting behavior and the prediction time window. The probability of the preceding vehicle's starting behavior is used to characterize the probability of the preceding vehicle starting, and the preset time window is used to characterize the time interval corresponding to the probability of the preceding vehicle's starting behavior.

[0077] It is worth noting that the probability of the preceding vehicle's intention to start can be represented as P_intent, and the prediction time window can be represented as [T_early, T_late], where T_early represents the earliest moment and T_late represents the latest moment.

[0078] In some implementations, the preset behavior prediction network is a high-precision behavior prediction network. The preset behavior prediction network includes different network paths. The visual micro-behavioral features and dynamic features of the preceding vehicle are input into the different network paths of the preset behavior prediction network to predict the starting behavior of the preceding vehicle, thereby obtaining the prediction result of the starting behavior of the preceding vehicle.

[0079] S105. Based on the prediction results of the starting behavior of the vehicle in front, generate a following reminder message for this vehicle.

[0080] In this embodiment of the application, a vehicle start-up event is generated based on the prediction result of the vehicle start-up behavior, and then a following reminder information for the current vehicle is generated based on the vehicle start-up event.

[0081] In practical applications, the following reminder information of this vehicle is used to remind the driver of this vehicle to follow the vehicle in front in a timely manner, so as to avoid the driver's delay due to a brief lapse in attention. This can improve the traffic efficiency of a single lane, alleviate overall congestion, reduce the driving pressure caused by the honking of the following vehicle, reduce the probability of cutting in line, and establish good traffic order.

[0082] In summary, this application provides a following alert method based on multimodal perception and behavior prediction. The method includes: acquiring information about the vehicle's front environment and its real-time motion state, wherein the front environment information includes a sequence of images of a preceding vehicle; extracting visual features from the preceding vehicle image sequence to obtain visual micro-behavioral features of the preceding vehicle; fusing features using a preset multimodal feature fusion network based on the preceding vehicle's visual micro-behavioral features and the vehicle's real-time motion state to obtain dynamic features of the preceding vehicle relative to the vehicle; predicting the preceding vehicle's starting behavior using a preset behavior prediction network based on the preceding vehicle's visual micro-behavioral features and dynamic features, obtaining a preceding vehicle starting behavior prediction result, wherein the preceding vehicle starting behavior prediction result includes the probability of the preceding vehicle's starting behavior intent and a prediction time window; and generating following alert information for the vehicle based on the preceding vehicle starting behavior prediction result. By acquiring multimodal perception information, such as the frontal environment information and the real-time motion status of the vehicle, and extracting the visual micro-behavioral features of the preceding vehicle from the image sequence, rich and comprehensive features of the preceding vehicle can be obtained. Furthermore, by combining the real-time motion status of the vehicle itself, the dynamic features of the preceding vehicle relative to the vehicle are determined, making the determined dynamic features of the preceding vehicle more accurate. Through detailed analysis of multimodal perception information, this embodiment of the application can accurately, timely, and reliably determine the prediction results of the starting behavior of the preceding vehicle, thereby making the generated following warning information of the vehicle more reliable.

[0083] Optionally, the visual micro-behavioral features of the preceding vehicle include: the physical dynamic features of the preceding vehicle and the semantic micro-behavioral features of the preceding vehicle; Figure 3 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 2 ,like Figure 3 As shown, the process of extracting visual features from the preceding vehicle image sequence to obtain the visual micro-behavioral features of the preceding vehicle in S102 above may include:

[0084] S201. Extract the physical dynamic features of the preceding vehicle from the image sequence to obtain the physical dynamic features of the preceding vehicle.

[0085] The physical dynamics features of the preceding vehicle include at least one of the following: the centroidal axial acceleration of the preceding vehicle's bounding box, the dense optical flow vector field at the bottom edge corners of the preceding vehicle, and its divergence. The process of extracting physical dynamics features can also be referred to as the process of extracting low-level physical dynamics features.

[0086] It should be noted that in the image sequence of the preceding vehicle, the preceding vehicle is enclosed by a first rectangular frame, which can be considered the bounding box of the preceding vehicle, with its centroid at the center point of the first rectangular frame. The axial acceleration of the centroid of the preceding vehicle's bounding box is used to characterize whether the preceding vehicle is accelerating away from the following vehicle or braking, causing the distance between the preceding and following vehicles to shorten.

[0087] Furthermore, in the image sequence of the preceding vehicle, the bottom edge corners of the preceding vehicle are enclosed by a second rectangle. When the preceding vehicle is farther away from the current vehicle, the four corners of the second rectangle are more densely distributed; when the preceding vehicle is farther away from the current vehicle, the four corners of the second rectangle are more sparsely distributed. The dense optical flow vector field at the bottom edge corners of the preceding vehicle can describe this feature. Both can describe the degree of expansion or contraction of the dense optical flow vector field.

[0088] S202. Extract semantic micro-behavioral features from the image sequence of the preceding vehicle to obtain the semantic micro-behavioral features of the preceding vehicle.

[0089] Among them, the semantic micro-behavioral features of the preceding vehicle include at least one of the following: the state encoding of the preceding vehicle's taillights, the texture motion spectrum of the preceding vehicle's wheel area, and the micro-change features of the preceding vehicle's posture.

[0090] In some implementations, a lightweight multi-task convolutional neural network is used to extract semantic micro-behavioral features of the preceding vehicle in parallel from the region of interest in the preceding vehicle image sequence.

[0091] It is worth noting that S201 can be executed first and then S202, or S202 can be executed first and then S201, or S201 and S202 can be executed simultaneously. This application embodiment does not impose specific restrictions on this.

[0092] Furthermore, the texture motion energy spectrum of the front vehicle's wheel area is used to characterize the energy distribution characteristics of wheel rotation and deformation in the frequency domain. It can be used to infer the wheel's rotational state (speed, slip ratio) and force conditions (braking, acceleration), thereby determining the vehicle's stability. The texture motion energy spectrum of the front vehicle's wheel area is also used to characterize the vehicle's pitch angle, roll angle, and displacement changes, which can be used to perceive the vehicle's dynamic trends (such as "nose-diving" during sudden braking, "nose-jumping" during rapid acceleration, and "roll" during cornering), allowing for early prediction of driving intentions.

[0093] Optionally, Figure 4 This is a schematic diagram of the structure of a preset behavior prediction network provided in an embodiment of this application, such as... Figure 4 As shown, the preset behavior prediction network includes: a time-series causal prediction network, an instantaneous pattern prediction network, and a dual-path fusion network; the preset behavior prediction network is also known as a dual-path heterogeneous prediction network, and the preset behavior prediction network can form a high-precision prediction path.

[0094] Figure 5 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 3 ,like Figure 5 As shown, the process in S104 above, which uses a preset behavior prediction network to predict the starting behavior of the vehicle in front based on the visual micro-behavioral features and dynamic features of the vehicle in front, and obtains the prediction result of the starting behavior of the vehicle in front, may include:

[0095] S301. A temporal causal prediction network is adopted to determine the temporal dependency of vehicle state evolution based on the visual micro-behavioral characteristics and dynamic characteristics of the preceding vehicle.

[0096] Among them, the temporal causal prediction network refers to a recurrent neural network with gated recurrent units, which can form a temporal causal path.

[0097] Additionally, time-series causal prediction networks include multiple gated recurrent units (GRUs) connected sequentially, such as... Figure 4 As shown, a temporal causal prediction network may include GRU1, GRU2, and GRU3 connected in sequence. It should be understood that... Figure 4 The number of GRUs is merely an example, and this application does not impose a specific limit on the number of GRUs. In the multiple sequentially connected gating loop units, the input of the first gating loop unit is the state information of different parts of the preceding vehicle and the relative state information between the preceding vehicle and the current vehicle. The output of the preceding gating loop unit is the input of the following gating loop unit, and the last gating loop unit is used to output the temporal dependency relationship of the vehicle state evolution.

[0098] In this embodiment, the gating loop unit transmits information step by step along the timeline through a gating mechanism, ultimately preserving in the hidden state those key events that occurred a few seconds ago but still have an impact on the present (such as braking, waiting, and the movement of the vehicle in front). This cross-time point information association is the temporal dependency relationship of the vehicle state evolution.

[0099] S302. An instantaneous mode prediction network is adopted to determine the feature information of the moment before the starting intention of the vehicle in front is generated, based on the visual micro-behavioral features and dynamic features of the vehicle in front.

[0100] The instantaneous pattern prediction network is a multi-layer, one-dimensional temporal convolutional network that can form instantaneous pattern pathways. For example, the instant before the preceding vehicle's intention to start can refer to 0.5 to 1 second before the preceding vehicle's intention to start.

[0101] The transient pattern prediction network consists of multiple sequentially connected temporal convolutional layers, all of which are one-dimensional convolutional layers (i.e., 1D Conv), such as... Figure 4 As shown, the instantaneous mode prediction network may include: 1D Conv1, 1D Conv2, and 1D Conv3 connected in sequence. It should be understood that... Figure 4The number of 1D Convs is merely an example, and this application does not impose a specific limitation on the number of 1D Convs. In the sequentially connected multiple temporal convolutional layers, the input to the first temporal convolutional layer is the state information of different parts of the preceding vehicle and the relative state information between the preceding vehicle and the current vehicle. The output of the previous temporal convolutional layer is the input to the next temporal convolutional layer, and the last temporal convolutional layer is used to output the feature information of the instant before the preceding vehicle's intention to start is generated.

[0102] In some implementations, features are determined just before the vehicle's intention to start is generated, based on the vehicle's visual micro-behavioral features and dynamic features. The first temporal convolutional layer slides along the time axis to extract short-term patterns (such as texture changes within 50 milliseconds after the brake lights turn off). Subsequent temporal convolutional layers expand the receptive field by stacking, combining the features at the lower level (such as brightness jumps) into more complex information at the higher level (such as the temporal combination of brake lights turning off, the front of the car rising, and wheel micro-movements), ultimately obtaining the feature information just before the vehicle's intention to start is generated.

[0103] In practical applications, the characteristic information of the vehicle just before it begins to move can include visual and kinematic features. For example, visual features might include: brake lights turning off, a slight tilt of the front of the vehicle upwards, and the wheel treads becoming blurred / the high-frequency components of the energy spectrum increasing. Kinematic features might include: minor fluctuations in optical flow divergence and changes in the vertical acceleration of the center of mass.

[0104] It should be noted that the processes S301 and S302 described above can be executed in parallel or sequentially, and this application embodiment does not impose specific restrictions on this.

[0105] S303. A dual-path fusion network is adopted to fuse the temporal dependency of vehicle state evolution and the feature information of the instant before the preceding vehicle's starting intention, and to generate the probability of the preceding vehicle's starting behavior intention and the prediction time window.

[0106] Among them, the dual-path fusion network is a learnable adaptive aggregation layer.

[0107] In this embodiment, a learnable adaptive aggregation layer is used to fuse the temporal dependencies of vehicle state evolution output by the temporal causal path and the feature information of the instantaneous mode path output and the moment before the preceding vehicle's starting intention, to generate a high-precision probability of the preceding vehicle's starting behavior intention and a prediction time window.

[0108] Optionally, Figure 6 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 4 ,like Figure 6As shown, the process in S105 above, which generates the following warning information for the vehicle based on the prediction result of the preceding vehicle's starting behavior, may include:

[0109] S401. Obtain the current environmental confidence level based on the preceding vehicle image sequence and current weather information.

[0110] Among these methods, the image sharpness is determined based on the sequence of images of the vehicle ahead, and the confidence level of the current environment can be determined based on the image sharpness and current weather information.

[0111] S402. The driver's attention level of this vehicle is collected based on the driver status monitoring system of this vehicle.

[0112] like Figure 1 As shown, the vehicle is also equipped with a Driver Monitor System (DMS). The Driver Monitor System uses a built-in camera facing the driver to run a lightweight algorithm to estimate the driver's attention level and the rate of change of the driver's attention level in real time.

[0113] In some implementations, the driver status monitoring system acquires images of the driver's face; it employs a lightweight convolutional neural network to detect in real time the coordinates of the driver's eye centers and their opening and closing states, the three-dimensional rotation angle of the head, and the facial orientation vector based on the driver's facial images.

[0114] In this embodiment, attention level A[0,100] is defined, with lower values ​​indicating less concentration. The driver's instantaneous attention score for each frame of the driver's facial image is calculated based on a weighted sum of the following three dimensions:

[0115] Aframe=W1×Sgaze+W2×Shead+W3×Sblink+W4×Sprion

[0116] Among them, Sgaze is the gaze deviation score, which is calculated by detecting the angle between the center of the eyeball and the normal to the face. Linear deductions are made when the horizontal gaze deviation exceeds ±15° or the vertical gaze deviation exceeds ±10°. Shead is the head posture score, with deductions made when the absolute value of the head yaw angle exceeds 20° or the absolute value of the pitch angle exceeds 15°, and zeroed out for severe head tilt (>45°). Sblink is the blink frequency and eye-closing duration score, using a blink detection network to count the number of blinks per second. Normal blinks (0.2~0.4 seconds / blink) do not incur deductions, while prolonged eye closure (>0.5 seconds) or abnormally high-frequency blinks (>3 times / second) incur deductions. Sprion is the historical attention smoothing item, taking the average of the previous 5 frames to prevent drastic changes in the level due to momentary jitter. Aframe is the driver's instantaneous attention score, i.e., the driver's attention level.

[0117] In addition, the weights W1, W2, W3, and W4 were obtained by fitting a calibration dataset collected from actual vehicles. For example, W1=0.4, W2=0.3, W3=0.2, and W4=0.1.

[0118] In this embodiment of the application, the rate of change of driver attention level is defined. (Unit: Levels / Second) is:

[0119]

[0120] in, Take 1 second. A positive rate of change indicates that attention is recovering, while a negative rate indicates that attention is rapidly declining. When the level is 1 / second and lasts for more than 2 seconds, the system determines that the driver is in a state of rapidly declining attention and triggers the highest priority warning. This represents the driver's attention level at time t. express Driver attention level at all times.

[0121] For example, in a red light queue scenario, if the driver is looking straight ahead, A=85. The driver looked down at his phone (head tilt angle -35°, line of sight deviated -40°), and A dropped from 85 to 20 within 1 second. Rating / second. The pilot briefly turns his head to speak with the passenger (yaw angle +25°, lasting 2 seconds), A drops to 50, rate of change. Levels per second. After the driver turns around and their vision returns, A rises from 50 to 80, at a rate of change of [missing value]. Levels per second.

[0122] S403. Based on the current environmental confidence level and the driver's attention level, generate the dynamic start-up decision threshold for this vehicle.

[0123] In some implementations, a dynamic start-up decision threshold for the vehicle is generated based on the current environment confidence level, driver attention level, weighting coefficients, and a preset baseline threshold. This can be expressed as: Th_start = α × (1 - C_env) + β × (1 - A_driver) + Th_base, where α and β are weighting coefficients, Th_base is the preset baseline threshold, C_env is the current environment confidence level, A_driver is the driver attention level, and Th_start is the dynamic start-up decision threshold for the vehicle.

[0124] S404. If the probability of the preceding vehicle's intention to start is greater than the dynamic start decision threshold, and the current time is within the prediction time window, generate the preceding vehicle start event.

[0125] Specifically, if P_intent > Th_start and the current time falls within the prediction time window, a preceding vehicle start event is generated. The processes S401 to S404 described above constitute the generation process of the dynamic decision threshold.

[0126] S405. Generate following alert information based on the preceding vehicle's start-up event and the driver's attention level.

[0127] In this embodiment of the application, after generating the following vehicle reminder information, the following vehicle reminder information can be output. After receiving the following vehicle reminder information, the user can promptly perform the following vehicle operation.

[0128] Optionally, Figure 7 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 5 ,like Figure 7 As shown, the process in S405 above, which generates following warning information based on the preceding vehicle's start-up event and the driver's attention level, may include:

[0129] S501. Calculate the overall confidence level based on current weather information, driver attention level, and probability of the preceding vehicle's intention to start.

[0130] S502. Employs a reinforcement learning strategy based on finite state machines to output a reminder pattern matrix based on the preceding vehicle's start-up event, overall confidence level, driver attention level, and the rate of change of the driver's attention level in this vehicle.

[0131] Among them, a finite state machine and a reinforcement learning strategy are adopted to form a reinforcement learning strategy based on a finite state machine.

[0132] Figure 8 This application provides a schematic diagram of a framework for generating following vehicle alert information, as shown in the embodiments of this application. Figure 8 As shown, the preceding vehicle's start-up event, overall confidence level, driver attention level, rate of change of the driver's attention level in this vehicle, and personalized response model are input into a finite state machine and a reinforcement learning policy. The finite state machine and reinforcement learning policy can output a reminder pattern matrix. Figure 8 As shown, the alert mode matrix includes information in three dimensions: the driver's status, the scene complexity, and the event confidence level.

[0133] It should be noted that a personalized response model is formed by recording the driver's historical response data to different alerts, i.e., the historical following behavior in response to historical following alert information. When the preceding vehicle starts moving, the personalized response model is first used to provide a personalized tolerance time Δt to output the following alert information, rather than immediately using a fixed delay to escalate the alert.

[0134] S503. Determine the following reminder information based on the driver's status, scene complexity, and event confidence in the reminder mode matrix.

[0135] In some implementations, when the vehicle is at a complex intersection or when the confidence level of the event is low, even if the driver is in a good condition, a dual-mode visual and auditory following alert can be used to increase the redundancy of the following alert and ensure safety.

[0136] In this application embodiment, the reminder execution method differs for different following vehicle reminder messages, such as... Figure 8 As shown, reminders can be delivered through visual, auditory, or tactile means. Figure 8 As shown, the personalized response model can also be updated based on the historical response data corresponding to the reminder execution method.

[0137] Optionally, Figure 9 A flowchart illustrating a following alert method based on multimodal perception and behavior prediction provided in this application embodiment. Figure 6 ,like Figure 9 As shown, before generating the following warning information for this vehicle based on the prediction result of the preceding vehicle's starting behavior in step S105, the method may further include:

[0138] S601. Obtain the relative distance sequence between this vehicle and the vehicle in front.

[0139] like Figure 1 As shown, the vehicle may also include a TOF ranging sensor (single-point time-of-flight ranging sensor), which spatially coordinates the TOF ranging sensor and the camera of the visual perception unit to obtain the absolute distance sequence with the vehicle in front at critical decision moments.

[0140] In addition, the TOF ranging sensor is a precise ranging verification unit, which together with the visual perception unit constitutes an environmental perception module that can collect information about the environment in front of the vehicle.

[0141] It should be noted that before performing specific processing, this application initiates a spatiotemporal synchronization protocol, performs hardware timestamp alignment and coordinate system unification on the TOF ranging points corresponding to the preceding vehicle image sequence, the real-time motion state of the current vehicle, and the relative distance sequence, and constructs a spatiotemporally aligned multimodal data stream.

[0142] S602. Calculate the speed of the preceding vehicle relative to the current vehicle based on the relative distance sequence.

[0143] The relative distance sequence is a high-frequency relative distance. Based on this high-frequency relative distance, the speed of the preceding vehicle relative to the current vehicle is calculated, and the speed of the preceding vehicle relative to the current vehicle can be represented as v_rel.

[0144] S603. Calculate the acceleration of the vehicle based on its motion state.

[0145] The acceleration of this vehicle can be represented as a_ego.

[0146] S604. Update the probability of the preceding vehicle's starting intention based on the speed of the preceding vehicle relative to the current vehicle and the current vehicle's acceleration.

[0147] In some implementations, driving intent differential calculation is performed, and cross-analysis is conducted on the speed of the preceding vehicle relative to the current vehicle and the acceleration of the current vehicle. If the speed of the preceding vehicle relative to the current vehicle is greater than 0, and the acceleration of the current vehicle is approximately equal to 0, i.e., v_rel>0 and a_ego≈0, then strong evidence indicates that the preceding vehicle actively started, and the probability of the preceding vehicle's starting behavior is updated.

[0148] Furthermore, if the speed change of the vehicle in front relative to this vehicle is synchronized with the acceleration of this vehicle, then it is determined that this vehicle is rolling backward. The cross-analysis provided in the embodiments of this application can eliminate the vast majority of interference scenarios.

[0149] S605, and / or, update the probability of the preceding vehicle's intention to start based on the preceding vehicle's start signal broadcast by the preceding vehicle.

[0150] like Figure 1 As shown, the vehicle also includes a vehicle-to-infrastructure (V2X) communication unit, which can receive basic safety messages broadcast from the preceding vehicle or roadside facilities, including a preceding vehicle start signal broadcast by the preceding vehicle. In practical applications, after the preceding vehicle starts, the V2X unit deployed on the preceding vehicle can broadcast a start signal, and the V2X unit deployed on this vehicle can receive the start signal broadcast by the preceding vehicle to obtain the preceding vehicle start signal. A Bayesian update is performed on the probability of the preceding vehicle's starting behavior intention based on the preceding vehicle start signal, significantly improving the confidence level.

[0151] It should be noted that the vehicle-road cooperative communication unit and the inertial measurement unit can constitute a vehicle status and cooperative communication module, which can obtain the status information of the vehicle and surrounding traffic participants.

[0152] In this embodiment, the preceding vehicle's start signal broadcast by the preceding vehicle and the driving intention differential can also participate in the determination of the comprehensive confidence level, i.e., the dynamic confidence level. Processes S601 to S604 above constitute the ranging reliability analysis and driving intention differential process, while process S605 above constitutes the dynamic decision threshold generation process. The ranging reliability analysis and driving intention differential process, the dynamic decision threshold generation process, and the dynamic decision threshold generation process constitute a multi-source game verification and dynamic confidence level synthesis process.

[0153] Optionally, before the process of extracting visual features from the preceding vehicle image sequence to obtain the visual micro-behavioral features of the preceding vehicle in S102 above, the method may further include:

[0154] If the vehicle's speed remains below a preset speed threshold during real-time motion, and the vehicle in front is continuously detected in the preceding vehicle's image sequence, then the congestion following monitoring mode will be activated.

[0155] The process of S102 is executed after the congestion following monitoring mode is activated.

[0156] Optionally, the method further includes:

[0157] A warning signal is generated based on the physical dynamic characteristics of the vehicle in front.

[0158] In this embodiment, a rapid sensing path is adopted to generate a low-confidence rapid warning signal based on the centroidal axial acceleration of the bounding box of the vehicle in front, the dense optical flow vector field and divergence of the bottom edge corner of the vehicle in front, and the physical acceleration threshold.

[0159] Among them, the rapid sensing pathway and the high-precision prediction pathway constitute hierarchical prediction.

[0160] It should be noted that the warning signal also participates in the generation of the overall confidence score. The overall confidence score calculated after the warning signal is generated is higher than the overall confidence score calculated when the warning signal is not generated. When the warning signal is generated earlier than the earliest time of the prediction time window, a follow-up reminder can be issued earlier than the earliest time.

[0161] In this embodiment, an incremental self-evolutionary learning unit is also included, which anonymically caches the deviation data between the predicted and actual results locally, paying particular attention to difficult samples with false positives and false negatives. Through a contrastive learning method, the prediction network is periodically fine-tuned in the cloud or on the vehicle to better distinguish between "true starts" and "similar interference," achieving targeted optimization based on the user's usual routes and driving habits.

[0162] The following provides an example illustrating a specific application scenario of a following alert method based on multimodal perception and behavior prediction provided in this application.

[0163] All hardware modules in this embodiment are integrated into a smart rearview mirror, with the camera and ToF sensor located in front of the mirror body and factory calibrated. The processing module utilizes the computing power of the rearview mirror's main control chip. In practical applications, when a vehicle stops at a red light, the system activates and enters a spatiotemporal synchronization state. The driver of the vehicle in front prepares to start, causing a slight forward tilt and taillight changes. The system's vision unit captures these micro-features. The "instantaneous pattern path" of the dual-path prediction network first captures the specific waveform of the taillights turning off, while the "temporal causal path" confirms the vehicle's continuous forward tilt. The network fusion outputs a high start probability. Simultaneously, the ToF sensor detects a slow increase in relative distance, while the vehicle's IMU shows zero acceleration. Differential analysis of driving intent confirms this as the vehicle in front actively starting. The system comprehensively judges and generates a high-confidence event. At this time, the DMS module determines that the driver's gaze has deviated from the front. Based on the user's personalized response model (the user typically reacts quickly), the system selects a combination of "LED green breathing flashing + a soft prompt sound." The driver follows up promptly, and the system records this successful interaction for subsequent optimization.

[0164] This application embodiment can be provided as a pre-installed function for commercial vehicles. The system is deeply integrated into the vehicle's infotainment system and enables V2X communication. In practical applications, at a congested logistics park intersection, multiple trucks equipped with this system are queuing. When the lead truck starts, its T-Box broadcasts a "vehicle start" message via V2X. The prediction network of the following vehicles' systems has calculated a high start probability based on visual features. After receiving the V2X message, the cooperative signal verification unit increases the confidence level to nearly 100% through Bayesian updates. The system instantly triggers the highest priority cooperative reminder (such as hazard lights). At the same time, the fleet management cloud platform collects the following response delay data of each vehicle. Analysis reveals that the delay is generally high at a certain intersection. The platform can issue instructions to fine-tune the dynamic decision threshold of the vehicle system in that area, making it more "sensitive" at that intersection, thereby optimizing the overall fleet traffic efficiency.

[0165] For intelligent electric vehicles already equipped with forward-facing cameras, DMS cameras, and vehicle-to-everything (V2X) communication capabilities, this invention can be implemented as a software algorithm package and pushed out via OTA (Over-The-Air) upgrades. When a vehicle is driving in congested rainy nights, the visual system is interfered with. At this time, the channel attention mechanism of the system's feature fusion module automatically reduces the weight of channels susceptible to rain, such as texture features, and increases the reliance on vehicle contour motion and V2V (Vehicle-to-Vehicle) signals. The prediction network relies more on the vibration transmission patterns of the preceding vehicle sensed by the IMU and V2V signals for inference. When the system predicts with moderate confidence that the preceding vehicle may start moving, combined with the DMS monitoring that the driver is in a state of mild fatigue, the system decides to adopt a strong tactile alert mode of "flashing instrument icons + seat pulse vibration" based on the scene complexity (rainy night) and the driver's state, ensuring that the driver's attention can be reliably and effectively awakened even under adverse conditions.

[0166] In summary, the preset behavior prediction network in this embodiment is specifically designed for vehicle start-up micro-behaviors, possessing both long-term inference and short-term pattern capture capabilities. Compared to directly applying general time-series models, it exhibits significant advantages in prediction accuracy and scenario adaptability. The analysis and calculation of driving intention differences and dynamic confidence synthesis demonstrate strong anti-interference capabilities. When generating following alert information, the system comprehensively considers the driver's real-time state, historical behavioral habits, and dynamic scenario pressure, achieving intelligent interaction that adapts the vehicle to the driver, significantly reducing unnecessary disturbances and improving driving experience and safety. Through an incremental optimization mechanism based on comparative learning, the system can learn from errors, continuously optimizing its prediction performance for different vehicle models, driving styles, and specific road conditions, forming a continuously improving data and technological barrier. Moreover, the hardware in this application still uses a low-cost sensor combination, making the solution suitable for both pre-installed standard configurations and aftermarket additions, demonstrating strong versatility.

[0167] The following describes the following device, processing module, and storage medium used to implement the following warning method based on multimodal perception and behavior prediction provided in this application. For the specific implementation process and technical effects, please refer to the relevant content of the following warning method based on multimodal perception and behavior prediction mentioned above, which will not be repeated below.

[0168] Figure 10 The structure of a following warning device based on multimodal perception and behavior prediction provided in this application embodiment is as follows: Figure 10 As shown, the device includes:

[0169] The acquisition module 101 is used to acquire the front environment information of the vehicle and the real-time motion status of the vehicle, wherein the front environment information includes: a sequence of images of the vehicle in front;

[0170] Extraction module 102 is used to extract visual features based on the preceding vehicle image sequence to obtain the preceding vehicle's visual micro-behavioral features;

[0171] The fusion module 103 is used to perform feature fusion using a preset multimodal feature fusion network based on the visual micro-behavioral features of the preceding vehicle and the real-time motion state of the current vehicle, so as to obtain the dynamic features of the preceding vehicle relative to the current vehicle.

[0172] Prediction module 104 is used to predict the starting behavior of the vehicle ahead based on the visual micro-behavioral features and dynamic features of the vehicle ahead, using a preset behavior prediction network to obtain the starting behavior prediction result of the vehicle ahead. The starting behavior prediction result of the vehicle ahead includes: the probability of the starting behavior intention of the vehicle ahead and the prediction time window.

[0173] The generation module 105 is used to generate following reminder information for the current vehicle based on the prediction result of the starting behavior of the preceding vehicle.

[0174] Optionally, the visual micro-behavioral features of the preceding vehicle include: physical dynamic features of the preceding vehicle and semantic micro-behavioral features of the preceding vehicle; the extraction module 102 is specifically used to extract physical dynamic features from the preceding vehicle image sequence to obtain the preceding vehicle physical dynamic features; wherein, the preceding vehicle physical dynamic features include at least one of the following: the centroidal axial acceleration of the preceding vehicle's bounding box, the dense optical flow vector field at the bottom edge corner of the preceding vehicle, and the divergence; and to extract semantic micro-behavioral features from the preceding vehicle image sequence to obtain the preceding vehicle semantic micro-behavioral features; wherein, the preceding vehicle semantic micro-behavioral features include at least one of the following: the preceding vehicle's taillight state encoding, the texture motion energy spectrum of the preceding vehicle's wheel area, and the micro-change features of the preceding vehicle's attitude.

[0175] Optionally, the preset behavior prediction network includes: a temporal causal prediction network, a transient pattern prediction network, and a dual-path fusion network;

[0176] The prediction module 104 is specifically used to employ the temporal causal prediction network to determine the temporal dependency of vehicle state evolution based on the visual micro-behavioral features and dynamic features of the preceding vehicle; employ the instantaneous pattern prediction network to determine the feature information of the preceding vehicle just before its starting intention is generated based on the visual micro-behavioral features and dynamic features of the preceding vehicle; and employ the dual-path fusion network to fuse the temporal dependency of vehicle state evolution and the feature information of the preceding vehicle just before its starting intention is generated, thereby generating the probability of the preceding vehicle's starting behavior intention and a prediction time window.

[0177] Optionally, the generation module 105 is specifically configured to: obtain the current environmental confidence level based on the preceding vehicle image sequence and current weather information; collect the driver attention level of the vehicle based on the driver status monitoring system of the vehicle; generate a dynamic start-up decision threshold for the vehicle based on the current environmental confidence level and the driver attention level; generate a preceding vehicle start-up event if the probability of the preceding vehicle's starting behavior is greater than the dynamic start-up decision threshold, and the current time is within the prediction time window; and generate the following reminder information based on the preceding vehicle start-up event and the driver attention level.

[0178] Optionally, the prediction module 104 is specifically used to calculate a comprehensive confidence level based on the current weather information, the driver's attention level, and the probability of the preceding vehicle's starting behavior; to output a reminder pattern matrix using a reinforcement learning strategy based on a finite state machine, based on the preceding vehicle's starting event, the comprehensive confidence level, the driver's attention level, and the rate of change of the driver's attention level of the current vehicle; and to determine the following reminder information based on the current vehicle's driver state, scene complexity, and time confidence level in the reminder pattern matrix.

[0179] Optionally, the device further includes:

[0180] An update module is used to obtain a relative distance sequence between the current vehicle and the preceding vehicle; calculate the speed of the preceding vehicle relative to the current vehicle based on the relative distance sequence; calculate the acceleration of the current vehicle based on the current vehicle's motion state; update the probability of the preceding vehicle's starting behavior intention based on the speed of the preceding vehicle relative to the current vehicle and the current vehicle's acceleration; and / or update the probability of the preceding vehicle's starting behavior intention based on the preceding vehicle's broadcast starting signal.

[0181] Optionally, the device further includes:

[0182] The determination module is used to determine to start the congestion following monitoring mode if the speed of the vehicle is continuously less than a preset speed threshold in the real-time motion state of the vehicle and the preceding vehicle is continuously detected in the preceding vehicle image sequence.

[0183] Optionally, the generation module 105 is further configured to generate a warning signal based on the physical dynamic characteristics of the vehicle in front.

[0184] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0185] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0186] Figure 11 This is a schematic diagram of the structure of a processing module provided in an embodiment of this application, such as... Figure 11 As shown, the processing module includes: processor 201 and memory 202.

[0187] The memory 202 is used to store programs, and the processor 201 calls the programs stored in the memory 202 to execute the above method embodiments. The specific implementation and technical effects are similar, and will not be described in detail here.

[0188] Optionally, this application also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, performs the above-described method embodiments.

[0189] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0190] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0191] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.

[0192] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0193] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A car-following reminding method based on multi-modal perception and behavior prediction, characterized in that, include: Acquire information about the environment ahead of the vehicle and the real-time motion status of the vehicle, wherein the information about the environment ahead includes: a sequence of images of the vehicle in front; Visual features are extracted from the preceding vehicle image sequence to obtain the visual micro-behavioral features of the preceding vehicle; Based on the visual micro-behavioral features of the preceding vehicle and the real-time motion state of the current vehicle, a preset multimodal feature fusion network is used to perform feature fusion to obtain the dynamic features of the preceding vehicle relative to the current vehicle. Based on the visual micro-behavioral features and dynamic features of the preceding vehicle, a preset behavior prediction network is used to predict the starting behavior of the preceding vehicle, and the starting behavior prediction result of the preceding vehicle is obtained. The starting behavior prediction result of the preceding vehicle includes: the probability of the starting behavior intention of the preceding vehicle and the prediction time window. Based on the prediction results of the preceding vehicle's starting behavior, a following reminder message for the current vehicle is generated; The preceding vehicle's visual micro-behavioral features include: the preceding vehicle's physical dynamics features and the preceding vehicle's semantic micro-behavioral features; the step of extracting visual features from the preceding vehicle's image sequence to obtain the preceding vehicle's visual micro-behavioral features includes: Physical dynamic features of the preceding vehicle are extracted from the preceding vehicle image sequence to obtain the preceding vehicle physical dynamic features; wherein, the preceding vehicle physical dynamic features include at least one of the following: the centroidal axial acceleration of the preceding vehicle's bounding box, the dense optical flow vector field at the bottom edge corner of the preceding vehicle, and the divergence. Semantic micro-behavioral features are extracted from the preceding vehicle image sequence to obtain the preceding vehicle semantic micro-behavioral features; wherein, the preceding vehicle semantic micro-behavioral features include at least one of the following: preceding vehicle taillight state encoding, preceding vehicle wheel region texture motion energy spectrum, and preceding vehicle attitude micro-change features; The preset behavior prediction network includes: a temporal causal prediction network, an instantaneous pattern prediction network, and a dual-path fusion network; the step of using the preset behavior prediction network to predict the starting behavior of the preceding vehicle based on the visual micro-behavioral features and dynamic features of the preceding vehicle, and obtaining the starting behavior prediction result of the preceding vehicle, includes: Using the aforementioned temporal causal prediction network, the temporal dependencies of vehicle state evolution are determined based on the visual micro-behavioral features and dynamic features of the preceding vehicle. Using the instantaneous pattern prediction network, feature information of the preceding vehicle in the instant before its starting intention is generated is determined based on the preceding vehicle's visual micro-behavioral features and the preceding vehicle's dynamic features; The dual-path fusion network is used to fuse the temporal dependencies of the vehicle state evolution and the feature information of the instant before the preceding vehicle's starting intention, to generate the probability of the preceding vehicle's starting behavior intention and the prediction time window.

2. The method of claim 1, wherein, The step of generating the following warning information for the current vehicle based on the prediction result of the preceding vehicle's starting behavior includes: Based on the preceding vehicle image sequence and current weather information, obtain the current environmental confidence level; The driver's attention level of the vehicle is collected based on the driver status monitoring system of the vehicle. Based on the current environment confidence level and the driver attention level, a dynamic start-up decision threshold for the vehicle is generated; If the probability of the preceding vehicle's intention to start is greater than the dynamic start decision threshold, and the preceding vehicle starts within the prediction time window at the current moment; The following alert information is generated based on the preceding vehicle's start-up event and the driver's attention level.

3. The method of claim 2, wherein, The step of generating the following warning information based on the preceding vehicle's start-up event and the driver's attention level includes: Calculate the overall confidence level based on the current weather information, the driver's attention level, and the probability of the preceding vehicle's intention to start. A reinforcement learning strategy based on finite state machines is adopted to output a reminder pattern matrix based on the preceding vehicle's start-up event, the overall confidence level, the driver's attention level, and the rate of change of the driver's attention level of the current vehicle. The following vehicle reminder information is determined based on the driver's status, scene complexity, and time confidence in the reminder mode matrix.

4. The method according to claim 1, characterized in that, Before generating the following warning information for the current vehicle based on the prediction result of the preceding vehicle's starting behavior, the method further includes: Obtain the relative distance sequence between the vehicle and the vehicle in front; Calculate the speed of the preceding vehicle relative to this vehicle based on the relative distance sequence; Calculate the vehicle's acceleration based on its motion state; Update the probability of the preceding vehicle's starting behavior intention based on the speed of the preceding vehicle relative to the current vehicle and the acceleration of the current vehicle; And / or, The probability of the preceding vehicle's intention to start is updated based on the preceding vehicle's start signal broadcast by the preceding vehicle.

5. The method of claim 1, wherein, Before extracting visual features from the preceding vehicle image sequence to obtain the preceding vehicle's visual micro-behavioral features, the method further includes: If the vehicle's speed is consistently lower than a preset speed threshold during real-time motion, and the preceding vehicle is continuously detected in the preceding vehicle image sequence, then the congestion following monitoring mode is activated.

6. The method of claim 1, wherein, The method further includes: A warning signal is generated based on the physical and dynamic characteristics of the vehicle in front.

7. A vehicle characterized by comprising: include: The vehicle body and the processing module are provided, wherein the processing module is disposed in the vehicle body, and the processing module implements the following warning method based on multimodal perception and behavior prediction as described in any one of claims 1-6 when executing a computer program.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when read and executed, implements the following warning method based on multimodal perception and behavior prediction as described in any one of claims 1-6.