Fusion perception and risk judgment method for collision early warning of networked vehicles

By predicting the motion field and updating the query vector in the connected vehicle collision warning system, and combining multi-source perception features and historical trajectories, the problem of insufficient perception caused by the uncertainty of the future trajectory of objects in the blind spot is solved, thereby improving the robustness and accuracy of collision warning.

CN121640764APending Publication Date: 2026-03-10TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In existing technologies, collision warning systems for connected vehicles do not take into account the uncertainty of the future trajectory of objects in blind spots due to the direct fusion of detection results. This results in weak robustness to communication delays, insufficient perception of potential traffic participants, and a tendency to produce false alarms and missed alarms, thus affecting the availability of warnings.

Method used

By predicting the motion field based on the sampling time of the target coordinate system and multi-source perception features, the query vector is updated to aggregate historical perception features, reducing the dependence on previous perception algorithms. Furthermore, by fusing multi-source perception features, the fusion perception results of each traffic participant are generated. Combined with historical trajectories, future trajectories are predicted to identify collision risks and avoid the evaluation bias of deterministic models.

Benefits of technology

It improves robustness to communication latency, enhances the ability to detect objects in blind spots, reduces false alarms and missed alarms, and improves the accuracy and availability of collision warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640764A_ABST
    Figure CN121640764A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle safety, in particular to a fusion perception and risk judgment method for networked vehicle collision early warning, and the method comprises the steps: predicting motion fields of feature sequences of different sources, and calculating and updating query vectors, so as to obtain query vectors of respective sources after aggregation of historical perception features, and then multi-source perception feature fusion is carried out on the query vector to obtain a fusion query vector, and a fusion perception result of each traffic participant is generated according to the fusion query vector, so that a blind area object set is identified, and a future trajectory is predicted to identify a collision risk. Therefore, the problems that due to the fact that the detection result is directly fused and the uncertainty of the future trajectory of the blind area object is not considered in the related technology, the robustness of communication time delay is not high, the perception ability of potential traffic participants is insufficient, risk assessment is prone to false alarm and missing alarm, and the availability of network connection vehicle collision early warning is restricted are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle safety technology, and in particular to a fusion perception and risk assessment method for collision warning of connected vehicles. Background Technology

[0002] In related technologies, to avoid the impact of blind spots on a single vehicle driving on the road, making it difficult to perceive potential risks and causing traffic accidents, multi-source detection results can be obtained through pre-sensing algorithms. Then, based on the vehicle-road-cloud integrated system, multi-source sensing information can be obtained through network communication and fused sensing can be achieved in the edge cloud to expand the sensing range. Subsequently, potential risks can be identified based on deterministic prediction models, thereby realizing collision warning for connected vehicles.

[0003] However, in related technologies, on the one hand, vehicle collision fusion perception often directly fuses detection results, resulting in perception effectiveness being limited by preceding perception algorithms and weak robustness to communication latency, which in turn leads to insufficient perception capability for potential traffic participants in blind spots. On the other hand, vehicle collision risk assessment is often based on simple deterministic prediction models that directly output collision risk values ​​without considering the uncertainty of the future trajectory of objects in blind spots, resulting in overly conservative or aggressive risk assessments, which easily lead to false alarms and missed alarms, restricting the availability of connected vehicle collision warnings, and urgently needs to be addressed. Summary of the Invention

[0004] This application provides a fusion perception and risk assessment method for collision warning of connected vehicles, in order to solve the problem that related technologies, due to the direct fusion of detection results, do not consider the uncertainty of the future trajectory of objects in the blind spot, resulting in weak robustness to communication delays and insufficient perception of potential traffic participants, making risk assessment prone to false alarms and missed alarms, thus restricting the usability of collision warning for connected vehicles.

[0005] The first aspect of this application provides a fusion perception and risk assessment method for collision warning of connected vehicles, comprising the following steps: based on the sampling time of the target coordinate system and perception features from multiple sources, predicting the motion field of feature sequences from different sources, and calculating and updating query vectors according to the feature sequences from different sources and the motion field to obtain query vectors after aggregating historical perception features from each source; performing multi-source perception feature fusion on the query vectors after aggregating historical perception features from each source to obtain a fusion query vector, and generating fusion perception results for each traffic participant according to the fusion query vector; based on the fusion perception results, identifying a set of blind spot objects, and predicting future trajectories according to the set of blind spot objects and the historical trajectories of each blind spot object, so as to identify collision risks using the navigation path uploaded by the connected vehicle and the future trajectory.

[0006] Through the above technical means, the embodiments of this application can predict the motion field based on the target coordinate system and the sampling time of multi-source perception features, update the query vector to aggregate historical perception features, reduce the dependence on previous perception algorithms, and improve robustness to communication latency. At the same time, by fusing multi-source perception features on the query vectors aggregated from their respective sources, a fused perception result for each traffic participant is generated, fully integrating the advantages of multiple sources to fill blind spots. In addition, by identifying the set of objects in the blind spot based on the fused perception result, predicting future trajectories by combining historical trajectories, and associating them with the navigation path of connected vehicles, collision risks are identified, avoiding the evaluation bias of deterministic models and improving the reliability of risk warning.

[0007] Optionally, in one embodiment of this application, the step of calculating and updating the query vector based on the feature sequences from different sources includes: initializing the query vector in the perceptual features at the earliest acquisition time based on the feature sequences from different sources; calculating the corresponding position of the initialized query vector in the perceptual features at the next time time using the motion field; performing attention calculation on the key-value pair vector based on the corresponding position; and updating the query vector.

[0008] Through the above technical means, the embodiments of this application can calculate the corresponding position of the query vector in the perceived features at the next moment by calculating the motion field. This can determine the spatial correspondence between the perceived features at the previous moment and the perceived features at the next moment, avoid the error of the perception feature association caused by time misalignment, and enhance the robustness to communication latency. Furthermore, by selecting key-value pairs near the corresponding position to perform attention calculation, the query vector can be updated. Instead of passively receiving the detection results of the preceding perception algorithm, it dynamically focuses on the key information in the perceived features, which can make up for the errors of missed detection and false detection of the preceding perception algorithm, reduce the dependence on the preceding algorithm, and improve the reliability of fusion perception.

[0009] Optionally, in one embodiment of this application, the update formula for the query vector may be, but is not limited to, the following: , in, For time index, The length of the historical information sequence stored at the edge cloud. For the first feature sequence The query vector at each time point. For the first feature sequence The query vector at time +1, For source The motion field of the characteristic sequence, for Source in The perceived features are collected in real time and then undergo coordinate alignment and time-delay encoding. A function to update the position of the query vector in the perceived features at the next time step. This is a function that extracts the key vector from the feature vector at a given time. ATT is a function that extracts a value vector from a feature vector at a given time.

[0010] Through the above technical means, the embodiments of this application can iteratively update the query vector, allowing the initial query vector at the earliest collection time to gradually integrate the feature information of subsequent times, so that the final query vector can integrate historical information from multiple times from the same source, effectively aggregate historical perception features, and enhance the temporal coherence of perception features.

[0011] Optionally, in one embodiment of this application, the formula for calculating the collision risk may be, but is not limited to, the following: , in, For a moment, The time decay factor, For the identification of blind spot objects Collision risk value with connected vehicles, For the set of objects in the blind spot, For the first A blind spot object, The navigation path uploaded by the connected vehicle. For the first i The future trajectory of an object in a blind spot. For the first The uncertainty of the future trajectory of an object in a blind spot This is the function for determining the probability of trajectory overlap. This marks the start time for risk assessment. The duration of risk assessment.

[0012] Through the above technical means, the embodiments of this application can perform integral calculation on the product of the trajectory overlap probability judgment function and the time decay factor to replace the simple deterministic prediction model, thereby obtaining a quantitative value of the collision risk in the future time period, improving the accuracy and robustness of collision risk judgment, avoiding false alarms and missed alarms, and improving the availability of collision warning for connected vehicles.

[0013] Optionally, in one embodiment of this application, the method further includes: calculating the collision risk value of the blind spot object; and, if the collision risk value of the blind spot object is greater than or equal to a preset threshold, issuing a collision warning to the connected vehicle and broadcasting the position and motion status of the blind spot object.

[0014] Through the above technical means, the embodiments of this application calculate the quantified collision risk value of objects in the blind spot and compare it with a preset threshold to trigger a warning. On the one hand, it can accurately identify collision risk scenarios that require warning and improve the accuracy of collision warning. On the other hand, when the warning is triggered, the position and movement status of objects in the blind spot are broadcast simultaneously, which allows connected vehicles or drivers to grasp the risk information in a timely and clear manner, providing a direct basis for obstacle avoidance operations, effectively improving the practicality and operability of collision warning, and enhancing the usability of the collision warning system.

[0015] Optionally, in one embodiment of this application, the calculation formula for the sports field may be, but is not limited to, the following: , in, For source The motion field of the characteristic sequence, for Source in The perceived features are collected in real time and then undergo coordinate alignment and time-delay encoding. The length of the historical information sequence stored at the edge cloud. For dimension stacking operations, This is the input operation for spatiotemporal convolutional networks.

[0016] Through the above technical means, the embodiments of this application generate a motion field by inputting the feature sequence into a spatiotemporal convolutional network, which can provide a motion benchmark for subsequent query vector updates and multi-source fusion, effectively compensate for the latency in network communication, solve the problem of feature spatiotemporal misalignment caused by latency, and thus improve the robustness of multi-source perception fusion to communication latency.

[0017] A second aspect of this application provides a fusion perception and risk assessment device for collision warning of connected vehicles, comprising: a calculation module, configured to predict the motion field of feature sequences from different sources based on the sampling time of the target coordinate system and the perception features from multiple sources, and to calculate and update a query vector according to the feature sequences from different sources and the motion field to obtain a query vector after aggregating historical perception features from each source; a perception module, configured to perform multi-source perception feature fusion on the query vector after aggregating historical perception features from each source to obtain a fusion query vector, and to generate a fusion perception result for each traffic participant according to the fusion query vector; and a judgment module, configured to identify a set of blind spot objects based on the fusion perception result, and to predict future trajectories according to the set of blind spot objects and the historical trajectories of each blind spot object, so as to identify collision risks using the navigation path uploaded by the connected vehicle and the future trajectory.

[0018] Through the above technical means, the embodiments of this application can predict the motion field based on the target coordinate system and the sampling time of multi-source perception features, update the query vector to aggregate historical perception features, reduce the dependence on previous perception algorithms, and improve robustness to communication latency. At the same time, by fusing multi-source perception features on the query vectors aggregated from their respective sources, a fused perception result for each traffic participant is generated, fully integrating the advantages of multiple sources to fill blind spots. In addition, by identifying the set of objects in the blind spot based on the fused perception result, predicting future trajectories by combining historical trajectories, and associating them with the navigation path of connected vehicles, collision risks are identified, avoiding the evaluation bias of deterministic models and improving the reliability of risk warning.

[0019] Optionally, in one embodiment of this application, the calculation module includes: a first calculation unit, configured to initialize a query vector in the perceptual features at the earliest acquisition time based on the feature sequences from different sources; and a second calculation unit, configured to calculate the corresponding position of the initialized query vector in the perceptual features at the next time time using the motion field, and to perform attention calculation on the key-value pair vector based on the corresponding position to update the query vector.

[0020] Through the above technical means, the embodiments of this application can calculate the corresponding position of the query vector in the perceived features at the next moment by calculating the motion field. This can determine the spatial correspondence between the perceived features at the previous moment and the perceived features at the next moment, avoid the error of the perception feature association caused by time misalignment, and enhance the robustness to communication latency. Furthermore, by selecting key-value pairs near the corresponding position to perform attention calculation, the query vector can be updated. Instead of passively receiving the detection results of the preceding perception algorithm, it dynamically focuses on the key information in the perceived features, which can make up for the errors of missed detection and false detection of the preceding perception algorithm, reduce the dependence on the preceding algorithm, and improve the reliability of fusion perception.

[0021] Optionally, in one embodiment of this application, the update formula for the query vector may be, but is not limited to, the following: , in, For time index, The length of the historical information sequence stored at the edge cloud. For the first feature sequence The query vector at each time point. For the first feature sequence The query vector at time +1, For source The motion field of the characteristic sequence, for Source in The perceived features are collected in real time and then undergo coordinate alignment and time-delay encoding. A function to update the position of the query vector in the perceived features at the next time step. This is a function that extracts the key vector from the feature vector at a given time. ATT is a function that extracts a value vector from a feature vector at a given time.

[0022] Through the above technical means, the embodiments of this application can iteratively update the query vector, allowing the initial query vector at the earliest collection time to gradually integrate the feature information of subsequent times, so that the final query vector can integrate historical information from multiple times from the same source, effectively aggregate historical perception features, and enhance the temporal coherence of perception features.

[0023] Optionally, in one embodiment of this application, the formula for calculating the collision risk may be, but is not limited to, the following: , in, For a moment, The time decay factor, For the identification of blind spot objects Collision risk value with connected vehicles, For the set of objects in the blind spot, For the first A blind spot object, The navigation path uploaded by the connected vehicle. For the first The future trajectory of an object in a blind spot. For the first The uncertainty of the future trajectory of an object in a blind spot This is the function for determining the probability of trajectory overlap. This marks the start time for risk assessment. The duration of risk assessment.

[0024] Through the above technical means, the embodiments of this application can perform integral calculation on the product of the trajectory overlap probability judgment function and the time decay factor to replace the simple deterministic prediction model, thereby obtaining a quantitative value of the collision risk in the future time period, improving the accuracy and robustness of collision risk judgment, avoiding false alarms and missed alarms, and improving the availability of collision warning for connected vehicles.

[0025] Optionally, in one embodiment of this application, it further includes: a calculation module for calculating the collision risk value of the blind spot object; and a warning module for issuing a collision warning to the connected vehicle and broadcasting the position and motion status of the blind spot object when the collision risk value of the blind spot object is greater than or equal to a preset threshold.

[0026] Through the above technical means, the embodiments of this application calculate the quantified collision risk value of objects in the blind spot and compare it with a preset threshold to trigger a warning. On the one hand, it can accurately identify collision risk scenarios that require warning and improve the accuracy of collision warning. On the other hand, when the warning is triggered, the position and movement status of objects in the blind spot are broadcast simultaneously, which allows connected vehicles or drivers to grasp the risk information in a timely and clear manner, providing a direct basis for obstacle avoidance operations, effectively improving the practicality and operability of collision warning, and enhancing the usability of the collision warning system.

[0027] Optionally, in one embodiment of this application, the calculation formula for the sports field may be, but is not limited to, the following: , in, For source The motion field of the characteristic sequence, for Source in The perceived features are collected in real time and then undergo coordinate alignment and time-delay encoding. The length of the historical information sequence stored at the edge cloud. For dimension stacking operations, This is the input operation for spatiotemporal convolutional networks.

[0028] Through the above technical means, the embodiments of this application generate a motion field by inputting the feature sequence into a spatiotemporal convolutional network, which can provide a motion benchmark for subsequent query vector updates and multi-source fusion, effectively compensate for the latency in network communication, solve the problem of feature spatiotemporal misalignment caused by latency, and thus improve the robustness of multi-source perception fusion to communication latency.

[0029] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to implement the fusion perception and risk assessment method for collision warning of connected vehicles as described in the above embodiments.

[0030] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for fusion perception and risk assessment for collision warning of connected vehicles.

[0031] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, implements the above-described method for fusion perception and risk assessment for collision warning of connected vehicles.

[0032] This application's embodiments predict the motion field based on the target coordinate system and multi-source sensing feature sampling time, update the query vector to aggregate historical sensing features, reduce reliance on preceding sensing algorithms, and improve robustness to communication latency. Simultaneously, by fusing multi-source sensing features from the aggregated query vectors, a fused sensing result for each traffic participant is generated, fully integrating the advantages of multiple sources to fill blind spots. Furthermore, by identifying the set of objects in the blind spots based on the fused sensing result, predicting future trajectories by combining historical trajectories, and associating with the navigation path of connected vehicles, collision risks are identified, avoiding evaluation biases of deterministic models and improving the reliability of risk warnings. Therefore, this solves the problem that related technologies, due to direct fusion of detection results without considering the uncertainty of future trajectories of objects in blind spots, suffer from weak robustness to communication latency, insufficient perception of potential traffic participants, and are prone to false alarms and missed alarms in risk assessment, thus limiting the usability of connected vehicle collision warnings.

[0033] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0034] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the system architecture for the execution process of a fusion perception and risk assessment method for collision warning of connected vehicles according to an embodiment of this application; Figure 2 This is a flowchart of a fusion perception and risk assessment method for collision warning of connected vehicles provided in an embodiment of this application; Figure 3 This is a schematic diagram of a collision warning system for a connected vehicle according to an embodiment of this application; Figure 4 A flowchart illustrating the working principle of a fusion perception and risk assessment method for collision warning of connected vehicles according to an embodiment of this application; Figure 5 This is a block diagram of a fusion perception and risk assessment device for collision warning of connected vehicles provided in accordance with an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application.

[0035] Figure label: 101-Roadside unit, 102-Edge cloud, 103-Vehicle unit; 50-Fusion perception and risk assessment device for collision warning of connected vehicles; 100-Computing module, 200-Perception module, 300-Judgment module; 601-Memory, 602-Processor, 603-Communication interface. Detailed Implementation

[0036] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0037] The following describes, with reference to the accompanying drawings, a fusion perception and risk assessment method for collision warning of connected vehicles according to embodiments of this application. Addressing the problems mentioned in the background art, where direct fusion of detection results fails to consider the uncertainty of future trajectories of objects in blind spots, resulting in weak robustness to communication delays and insufficient perception of potential traffic participants, leading to false alarms and missed alarms in risk assessment and hindering the usability of collision warnings for connected vehicles, this application provides a fusion perception and risk assessment method for collision warning of connected vehicles. In this method, the motion field can be predicted based on the target coordinate system and the sampling time of multi-source perception features, updating the query vector to aggregate historical perception features, reducing reliance on preceding perception algorithms and improving robustness to communication delays. Simultaneously, by fusing multi-source perception features from the aggregated query vectors, fusion perception results for each traffic participant are generated, fully integrating the advantages of multiple sources to fill blind spots. Furthermore, by identifying the set of objects in blind spots based on the fusion perception results, predicting future trajectories by combining historical trajectories, and associating with the connected vehicle navigation path, collision risks are identified, avoiding evaluation biases of deterministic models and improving the reliability of risk warnings. This solves the problem that related technologies, by directly fusing detection results without considering the uncertainty of the future trajectory of objects in blind spots, have weak robustness to communication delays and insufficient perception of potential traffic participants, making risk assessment prone to false alarms and missed alarms, thus restricting the availability of collision warnings for connected vehicles.

[0038] Before introducing the fusion perception and risk assessment method for collision warning of connected vehicles provided in the embodiments of this application, we will first briefly introduce the system architecture of the fusion perception and risk assessment method for collision warning of connected vehicles.

[0039] Figure 1 This is a schematic diagram of the system architecture for the execution process of a fusion perception and risk assessment method for collision warning of connected vehicles according to an embodiment of this application.

[0040] The system architecture for the execution process of the fusion perception and risk assessment method for collision warning of connected vehicles provided in this application embodiment may include, but is not limited to, roadside unit 101, edge cloud 102 and vehicle unit 103.

[0041] Among them, the roadside unit 101 refers to the sensing devices (such as cameras, radars, etc.) deployed on the side of the road, which are used to collect sensing data of the road environment and traffic participants, and transmit the reported data to the edge cloud 102.

[0042] The edge cloud 102 may include, but is not limited to, an edge cloud data transfer subsystem, an edge cloud fusion perception subsystem, and an edge cloud collaborative decision-making subsystem. Specifically, the edge cloud data transfer subsystem can receive and store reported data from the roadside unit 101 and the vehicle unit 103 to provide data support for fusion perception and collaborative decision-making; the edge cloud fusion perception subsystem can perform fusion processing on the stored reported data to generate fusion perception results; and the edge cloud collaborative decision-making subsystem can make decisions based on the fusion perception results (such as collision risk assessment) and transmit the decision results as distributed data to the vehicle unit 103.

[0043] The vehicle unit 103 refers to the device (such as a vehicle sensor) installed on the connected vehicle, which is used to collect the vehicle's own perception data and transmit the reported data to the edge cloud 102. At the same time, it receives the data sent down from the edge cloud 102 to realize the vehicle's collaborative decision-making and control.

[0044] Specifically, Figure 2 This is a flowchart of a fusion perception and risk assessment method for collision warning of connected vehicles provided according to an embodiment of this application.

[0045] like Figure 2 As shown, the fusion perception and risk assessment method for collision warning of connected vehicles includes the following steps: In step S201, based on the target coordinate system and the sampling time of the sensing features from multiple sources, the motion field of the feature sequences from different sources is predicted, and the query vector is updated according to the feature sequences and motion fields from different sources to obtain the query vector after aggregating the historical sensing features from each source.

[0046] In the embodiments of this application, the target coordinate system refers to a coordinate system used to unify the spatial reference of sensing features from multiple sources, and to provide a basic reference for spatial alignment of sensing features from different sources, thereby ensuring the spatial consistency of sensing information.

[0047] In addition, perception features refer to raw perception data collected from different channels such as roadside sensors and connected vehicle sensors, which includes information such as the position, speed, and shape of objects. These data are used as the basic input for fusion perception in connected vehicle collision warning.

[0048] In addition, sampling time refers to the time when each of the sensing features from multiple sources is collected. It is used to characterize the temporal attributes of sensing features, handle network communication latency, and achieve time alignment of multi-source sensing features.

[0049] Furthermore, a feature sequence refers to a time-series set of sensing features from the same source arranged in chronological order of sampling time (such as a sequence formed by sorting sensing features collected by a roadside sensor at different times in chronological order). It is used to carry historical sensing information from the corresponding source and to provide a time-series data foundation for subsequent information aggregation of query vectors.

[0050] In addition, the sports field refers to the information predicted based on the sampling time of the target coordinate system and the perceived features from multiple sources. It describes the movement and change of the perceived features in the feature sequence over time (such as the displacement and velocity trends of the perceived features) and is used as an important basis for subsequent updating of the query vector.

[0051] In addition, the query vector refers to a vector tool used to focus on and aggregate historical information in a feature sequence. In this application embodiment, the query vector can be initialized in the perceptual features at the earliest acquisition time and updated through motion field calculation.

[0052] In actual implementation, when the embodiments of this application receive sensing features from different sources at the edge cloud, such as sensing features reported by roadside sensors and sensing features reported by connected vehicles, the sensing features from different sources can be uniformly transformed to the target coordinate system according to the relative positional relationship of each sensor. Then, based on the difference between the sampling time of each sensing feature and the current time, the sensing features are time-delay encoded. Subsequently, the sensing features can be sequentially input into a spatiotemporal convolutional network, from the earliest acquisition time to the latest acquisition time of all sensing features. Predict the motion field of feature sequences from different sources.

[0053] Furthermore, in this embodiment of the application, after predicting the motion field of feature sequences from different sources, the query vector is updated based on the feature sequences and motion fields from different sources.

[0054] Furthermore, in this embodiment, the above steps can be executed sequentially until the query vectors from different sources are all updated to the latest acquisition time of the perceived features from all sources. At this point, the query vectors obtained are the aggregated historical perception features from their respective sources. .

[0055] Specifically, in one embodiment of this application, calculating and updating the query vector based on feature sequences from different sources includes: initializing the query vector in the perceptual features at the earliest acquisition time based on feature sequences from different sources; calculating the corresponding position of the initial query vector in the perceptual features at the next time time using the motion field; performing attention calculation on the key-value pair vector based on the corresponding position; and updating the query vector.

[0056] It should be noted that prior sensing algorithms can detect the raw sensing information from each source before multi-source sensing information fusion. For example, the target detection algorithm of a roadside camera can output the detection results such as the size and speed of a vehicle / pedestrian at a certain location, and the obstacle recognition algorithm of vehicle radar can output the detection results such as distance / direction angle / speed. However, related technologies often directly fuse the detection results, which will overly rely on the accuracy of the prior sensing algorithm. When the prior sensing algorithm has missed detections or false detections, the subsequent fusion will directly inherit the errors of missed detections and false detections. At the same time, it will lack in-depth mining of sensing features. That is, the detection results of the prior sensing algorithm are processed conclusions and may lack useful details of sensing features. Subsequent fusion will find it difficult to correct errors and reduce the accuracy of the sensing results.

[0057] In practical implementation, to compensate for the shortcomings of directly fusing the detection results of preceding sensing algorithms, this application embodiment initializes the query vector in the sensing features at the earliest acquisition time for feature sequences from different sources, and utilizes the motion field... The corresponding position of the initial query vector in the perceived features at the next time step is calculated. Then, key-value pair vectors can be selected near the corresponding position to perform attention calculation and update the query vector.

[0058] This application embodiment can initialize the query vector in the sensing features at the next moment by calculating the motion field, which can determine the spatial correspondence between the sensing features at the previous moment and the sensing features at the next moment, avoid the sensing feature association error caused by time misalignment, and enhance the robustness to communication latency. Furthermore, by selecting key-value pairs near the corresponding position for attention calculation, the query vector can be updated. Instead of passively receiving the detection results of the preceding sensing algorithm, it dynamically focuses on the key information in the sensing features, which can make up for the errors of missed detection and false detection of the preceding sensing algorithm, reduce the dependence on the preceding algorithm, and improve the reliability of fusion sensing.

[0059] Optionally, in one embodiment of this application, the update formula for the query vector may be, but is not limited to, the following: , in, For time index, The length of the historical information sequence stored at the edge cloud. For the first feature sequence The query vector at each time point. For the first feature sequence The query vector at time +1, For source The motion field of the characteristic sequence, for Source in Perceptual features acquired in real time and after coordinate alignment and time-delay encoding A function to update the position of the query vector in the perceived features at the next time step. This is a function that extracts the key vector from the feature vector at a given time. ATT is a function that extracts a value vector from a feature vector at a given time.

[0060] In actual implementation, the embodiments of this application can be based on the first feature sequence. i The query vector at each time point Combined with sports field ,pass The function updates the position of the query vector in the perceived features at the next time step, and through... Functions and The function extracts a key-value pair vector, which is then used to compute and update the query vector via the Attention Calculation Network (ATT).

[0061] This application embodiment can iteratively update the query vector, allowing the initial query vector from the earliest collection time to gradually integrate feature information from subsequent times, so that the final query vector can integrate historical information from multiple times from the same source, effectively aggregating historical sensing features and enhancing the temporal coherence of sensing features.

[0062] Optionally, in one embodiment of this application, the formula for calculating the sports field may be, but is not limited to, the following: , in, For source The motion field of the characteristic sequence, for Source in Perceptual features acquired in real time and after coordinate alignment and time-delay encoding The length of the historical information sequence stored at the edge cloud. For dimension stacking operations, This is the input operation for spatiotemporal convolutional networks.

[0063] In actual implementation, regarding the source In this embodiment, the perceptual features of multiple consecutive time points are first collected after coordinate alignment and time-delay encoding. Then, a feature sequence is formed by dimension stacking. Subsequently, a spatiotemporal convolutional network is used to extract the spatiotemporal correlation information in the feature sequence, thereby outputting a result reflecting the source. A motion field whose perceptual features change over time.

[0064] This application embodiment generates a motion field by inputting the feature sequence into a spatiotemporal convolutional network, which can provide a motion benchmark for subsequent query vector updates and multi-source fusion, effectively compensate for latency in network communication, solve the problem of feature spatiotemporal misalignment caused by latency, and thus improve the robustness of multi-source perception fusion to communication latency.

[0065] Based on the description of other related embodiments, and exemplarily, the steps of generating a query vector from aggregated historical perceptual features from multiple sources in this application embodiment are as follows: (1) such as Figure 3 As shown, when a connected vehicle approaches an intersection with blind spots, it continuously sends information such as vehicle perception data, vehicle perception results, and vehicle status to the edge cloud via its onboard unit.

[0066] At this time, the roadside unit continuously uploads roadside sensing data to the edge cloud.

[0067] (2) The edge cloud receives perception data from vehicle-mounted units and roadside units, generates a motion field based on current perception features and historical features, guides the perception features to perform time delay compensation and fusion, and achieves higher precision and coverage of blind spots in fused perception.

[0068] (2-1) The edge cloud establishes a target coordinate system with the current location of the connected vehicle as the origin. Then, based on the positional relationships between the various sensors received from the edge cloud and the connected vehicle, a coordinate system transformation matrix is ​​constructed. Furthermore, coordinate system transformation is performed on perceptual features from different sources and historical perceptual features to unify them to the target coordinate system. The expression for the transformation process can be, but is not limited to, as follows: , in, for Source of time The original perceptual characteristics, After alignment to the target coordinate system The perceptual characteristics of time.

[0069] For the historical perception features of connected vehicles, the edge cloud can perform spatiotemporal alignment based on the vehicle's historical motion trajectory to eliminate the feature offset caused by vehicle motion. Similarly, based on the position and heading angle of the connected vehicle at each historical moment and the current position and heading angle, a coordinate system transformation matrix is ​​constructed to transform the historical features of the connected vehicle to unify them to the target coordinate system.

[0070] After the above processing, the embodiments of this application can unify the current and historical perception features from different sources (roadside sensors, connected vehicles) into a target coordinate system with the connected vehicle at the current moment as the origin, eliminating the feature offset caused by the inconsistency of the coordinate system, thereby supporting the correct performance of subsequent motion field calculation and feature fusion.

[0071] (2-2) For the edge cloud's perception features from different sources, calculate the sampling time and the latest sampling time. The difference is used for time-delay coding. The expression for time-delay coding can be, but is not limited to, as follows: , in, For source Sampling time of perceived features For the dimensions of perceived features, This is the index along the dimension.

[0072] It is understandable that, since delay coding records the delay information between the sampling time and the latest sampling time of sensing features from different sources, the embodiments of this application can directly add the sampling time to each sensing feature after a linear transformation operation, so that the sensing features have delay sensing characteristics. Its expression can be, but is not limited to, as follows: , in, For source exist Perceptual features are collected in real time and then processed with coordinate alignment and time-delay encoding. This is a linear transformation operation.

[0073] Furthermore, in this embodiment, feature sequences from different sources are stacked along the time dimension and then input into a spatiotemporal convolutional network to calculate the motion field of the perceptual features from different sources. The expression for the motion field can be, but is not limited to, as follows: .

[0074] Understandably, due to the sports field The motion vector containing the perceptual features at each location in the feature map, pointing to the perceptual features at the next time step, therefore... It can be used to guide the aggregation of historical features.

[0075] (2-3) After obtaining the motion field of the sensing features from different sources by edge cloud computing, for the feature sequences from different sources, the query vector is initialized in the sensing features of their earliest acquisition time. The motion field is used to calculate the corresponding position of the initial query vector in the perceived features at the next time step. The expression can be, but is not limited to, as follows: , in, and To initialize the horizontal and vertical positions of the query vector in the feature map, and To determine the horizontal and vertical positions of the calculated initial query vector within the perceived features at the next time step. and These are the horizontal and vertical motion vectors recorded in the motion graph.

[0076] For ease of explanation, in subsequent embodiments, the process of calculating the corresponding position will be abbreviated as... .

[0077] Furthermore, after obtaining the corresponding location in the edge cloud, key-value pair vectors are extracted from the feature map, attention calculations are performed, and the query vector is updated. Subsequently, the next time step location is calculated again based on the motion map, key-value pair vectors are extracted, attention calculations are performed, and the above process is repeated until the query vector is updated to the latest sampling time of all perceptual features. At this point, the query vector after aggregating historical perception features can be obtained. .

[0078] It should be noted that the update formula for the query vector can be, but is not limited to, the following: .

[0079] In step S202, the query vectors after aggregating historical perception features from their respective sources are fused with multi-source perception features to obtain a fused query vector, and the fused perception results of each traffic participant are generated based on the fused query vector.

[0080] In the embodiments of this application, the fusion query vector refers to the vector obtained by fusion of multi-source perception features after stacking the query vectors of aggregated historical features from various sources in the feature dimensions, and then using a multilayer perceptron to generate the fusion perception results of each traffic participant.

[0081] In addition, the fusion perception result refers to the information such as the category, location, and motion state of each traffic participant obtained by decoding the fusion query vector through the detection head, which is used to identify the set of objects in the blind spot and provide a basis for collision risk assessment.

[0082] In actual implementation, the embodiments of this application first perform stacking operations on the feature dimension for the query vector after aggregating historical perception features from their respective sources, and then use a multilayer perceptron to fuse multi-source perception features to obtain a fused query vector. After decoding the fused query vector using a detection head, the fused perception results of each traffic participant can be obtained.

[0083] Based on the description of other related embodiments, and exemplarily, the steps of generating the fused perception results of each traffic participant from the query vector after aggregating historical perception features from their respective sources in this application embodiment are as follows: The edge cloud aggregates query vectors from different sources based on historical sensing features, first stacks them along the feature dimension, and then uses a multilayer perceptron to fuse these multi-source sensing features to obtain a fused query vector. The expression for multi-source sensing feature fusion can be, but is not limited to, as follows: , in, To merge query vectors, It is a multilayer perceptron. and For different sources Aggregated query vector.

[0084] Furthermore, edge cloud utilizes detection heads to... Decoding yields a fusion perception result for each traffic participant, including object size, position, speed, direction, and category.

[0085] In summary, the multi-source data received by the edge cloud is sent to the edge cloud fusion perception subsystem via the edge cloud data transfer subsystem. The fusion perception subsystem then performs multi-source information fusion perception, taking into account the impact of communication latency, and sends the fusion perception results back to the edge cloud data transfer subsystem.

[0086] In step S203, based on the fusion perception results, a set of blind spot objects is identified, and the future trajectory is predicted according to the set of blind spot objects and the historical trajectory of each blind spot object, so as to identify collision risks by using the navigation path and future trajectory uploaded by the connected vehicle.

[0087] In this embodiment of the application, the blind spot object set refers to the set of unmatched detection results after the fusion perception results are matched with the detection results of the connected vehicle itself, which is used for subsequent trajectory prediction and collision risk assessment of blind spot objects.

[0088] In addition, historical trajectory refers to the past movement path record of objects in the blind spot object set, which is used to predict the future trajectory of the object.

[0089] In addition, the future trajectory refers to the future movement path of an object predicted based on the historical trajectory of the object in the blind spot, which is used to identify collision risks in conjunction with the navigation path uploaded by connected vehicles.

[0090] Furthermore, collision risk refers to the cumulative collision risk value between blind spot objects and connected vehicles, which is quantified by an integral calculation formula, combined with the navigation trajectory of connected vehicles, the future trajectory of objects in blind spots, and trajectory uncertainties. This value is used as the basis for decision-making in collision warnings for connected vehicles.

[0091] In actual implementation, after the fusion perception is completed at the edge cloud, the vehicle detection results reported by the connected vehicles are first transformed to the target coordinate system, and then the fusion perception results are processed. And vehicle inspection results A matching process is performed, and unmatched results are marked as a possible set of blind spot objects. The expression for the possible set of blind spot objects can be, but is not limited to, as follows: , in, For the set of possible blind spot objects, To match the confidence level, This is a matching algorithm used to output a matching result with a specified confidence level, given different fusion perception results, vehicle detection results, and matching confidence levels.

[0092] Furthermore, in this embodiment, the above operations are performed on the detection results of N consecutive frames, for those that persist... Objects are marked as blind spot objects for connected vehicles, resulting in the final set of blind spot objects. Then, based on the historical trajectories of objects in each blind spot, future trajectories are predicted, and collision risks are assessed using navigation paths uploaded by connected vehicles.

[0093] Optionally, in one embodiment of this application, the formula for calculating collision risk may be, but is not limited to: , in, For a moment, The time decay factor, For the identification of blind spot objects Collision risk value with connected vehicles, For the set of objects in the blind spot, For the first A blind spot object, The navigation route uploaded for connected vehicles. For the first The future trajectory of an object in a blind spot For the first The uncertainty of the future trajectory of an object in a blind spot This is the function for determining the probability of trajectory overlap. This marks the start time for risk assessment. The duration of risk assessment.

[0094] It should be noted that related technologies often output collision risk values ​​based on simple deterministic prediction models, without taking into account the uncertainty of the future trajectory of objects in the blind spot. This leads to overly conservative or aggressive risk assessments, which can easily result in false alarms or missed alarms, thus limiting the availability of collision warnings for connected vehicles.

[0095] In practical implementation, to avoid limiting the availability of vehicle collision warnings due to the use of deterministic prediction models, the embodiments of this application can utilize a trajectory overlap probability judgment function. Navigation routes uploaded by integrated connected vehicles The future trajectory of objects in the blind spot and the uncertainty of the future trajectory of objects in the blind spot The probability of trajectory overlap is assessed in a probabilistic manner, and a time decay factor can be introduced. By reducing the weight of long-term forecasts to address the uncertainty of long-term forecasts, collision risks in future time periods can be accumulated and quantified through integral calculations.

[0096] The embodiments of this application can replace the simple deterministic prediction model by performing an integral operation on the product of the trajectory overlap probability judgment function and the time decay factor, thereby obtaining a quantitative value of the collision risk in the future time period, improving the accuracy and robustness of collision risk judgment, avoiding false alarms and missed alarms, and improving the availability of collision warning for connected vehicles.

[0097] Furthermore, in one embodiment of this application, the method further includes: calculating the collision risk value of the blind spot object; and issuing a collision warning to the connected vehicle and broadcasting the position and motion status of the blind spot object when the collision risk value of the blind spot object is greater than or equal to a preset threshold.

[0098] In the embodiments of this application, the collision risk value refers to the quantified cumulative collision risk value between blind spot objects and connected vehicles, which is used to measure the probability of collision between the two and provide a quantitative basis for collision warning decisions.

[0099] In addition, the preset threshold refers to a pre-set collision risk threshold, which is used to determine whether the collision risk has reached a level that requires triggering a warning.

[0100] In actual implementation, the embodiments of this application can calculate the collision risk value of blind spot objects, and then compare the collision risk value of blind spot objects with a preset threshold. If the collision risk value of blind spot objects is greater than or equal to the preset threshold, a collision warning is issued to the connected vehicle. At the same time, the position and movement status of the blind spot objects are broadcast to help the connected vehicle avoid possible collisions in advance.

[0101] This application embodiment calculates a quantified collision risk value for objects in the blind spot and compares it with a preset threshold to trigger a warning. On the one hand, it can accurately identify collision risk scenarios that require warning and improve the accuracy of collision warnings. On the other hand, when a warning is triggered, the position and movement status of objects in the blind spot are broadcast simultaneously, allowing connected vehicles or drivers to grasp risk information in a timely and clear manner, providing a direct basis for obstacle avoidance operations, effectively improving the practicality and operability of collision warnings, and enhancing the usability of the collision warning system.

[0102] Based on the description of other related embodiments, and by way of example, the steps for determining collision risk from the fused perception results and the vehicle detection results in this application embodiment are as follows: (1) Based on the fusion perception results and the vehicle detection results reported by the connected vehicles, the edge cloud identifies possible blind spot objects. Based on the multi-frame recognition results, objects that continuously exist in the blind spot are recorded in the blind spot object set. Then, based on the historical trajectory of each blind spot object, the future trajectory is predicted for subsequent collision risk assessment.

[0103] (1-1) The edge cloud determines the confidence level based on the fusion perception results and the vehicle detection results. Then, the Hungarian algorithm is used to match the perception results. The expression for the matching cost function can be, but is not limited to, as follows: , in, To integrate perception to identify objects and vehicle-mounted object recognition The matching cost between them To identify the intersection-union ratio (IU) of 3D bounding boxes for two objects. and Let these be the coordinates of the center positions of the two objects. Indicates an indicator function, and For the categories of the two objects, These are the weighting coefficients.

[0104] It should be noted that when the conditions At the time of its establishment, The value is 1 when the condition is met. When not valid, It is 0.

[0105] The edge cloud uses the Hungarian algorithm to obtain matching result pairs, filtering out those whose matching cost exceeds the confidence level. The result is the matching result. Then, the matching results are filtered out from the fused perception results, thus obtaining the possible set of blind spot objects. The expression for the possible set of blind spot objects can be, but is not limited to, as follows: , in, For the set of possible blind spot objects, To integrate the perception results, For the vehicle inspection results, This is a matching algorithm.

[0106] Furthermore, the edge cloud performs the above operations on the detection results of N consecutive frames, for those that persist... The objects are identified, their identifiers are recorded, and they are marked as blind spot objects for connected vehicles, resulting in the final set of blind spot objects. .

[0107] (1-2) The edge cloud predicts the future trajectory of objects in the blind spot based on the fusion perception results and the set of objects in the blind spot.

[0108] The edge cloud first uses the historical trajectories of each traffic participant to perform time-series aggregation. The expression for time-series aggregation can be, but is not limited to, the following: , in, For traffic participants The aggregation characteristics, For traffic participants exist Motion status information at any given time. It should be noted that each traffic participant can be understood as all traffic participants perceived through fusion, which may include, but is not limited to, objects in blind spots.

[0109] Subsequently, the edge cloud constructs the interaction relationships among various traffic participants and updates features using an attention mechanism. The expression for this can be, but is not limited to, as follows: , in, For traffic participants Based on the updated aggregated features according to the interaction effects For traffic participants and The interaction between them can be represented by their relative positional relationships.

[0110] Furthermore, by decoding the updated aggregated features of each traffic participant at the edge cloud using a multilayer perceptron, the future trajectories and uncertainties of these trajectories can be obtained. The expression for these uncertainties can be, but is not limited to, as follows: , in, For traffic participants The future trajectory, It indicates the uncertainty of the future trajectory.

[0111] It should be noted that in (1-2), the fusion perception results and the vehicle detection results are sent to the edge cloud collaborative decision-making subsystem via the edge cloud data transfer subsystem. After the edge cloud collaborative decision-making subsystem identifies potential blind spot objects, it judges the blind spot objects by combining the identification results of multiple historical frames, predicts the future trajectory of the blind spot objects based on their historical trajectories, and then sends the set of blind spot objects and the corresponding prediction results of their future trajectories back to the edge cloud data transfer subsystem.

[0112] (2) The edge cloud uses the identification results of the blind spot object set and the prediction results of the future trajectory, as well as the navigation path uploaded by the connected vehicle, to determine the location of the object. It calculates the collision risk between connected vehicles and objects in blind spots at future moments.

[0113] Edge cloud for connected vehicles and blind spot objects at a certain moment The probability of trajectory overlap, using Calculate the mean position distribution of objects in the blind spot relative to connected vehicles. Covariance Since the trajectory uncertainty follows a Gaussian distribution, the probability of overlap at a certain moment can be calculated, and its expression can be, but is not limited to, as follows: , in, Let be the cumulative distribution function of the Gaussian distribution. objects in the blind spot The equivalent safety radius for connected vehicles can be approximated by a circle based on the shape of the object. For along The unit direction vector.

[0114] Furthermore, edge cloud is based on the future period of time. The probability of trajectory overlap within the area is used to calculate the collision risk using a weighted average. The formula for calculating the collision risk can be, but is not limited to, the following: , in, For the identification of blind spot objects Collision risk value with connected vehicles, and These represent the start time and duration of the risk assessment, respectively.

[0115] The collision risk calculation formula can calculate the cumulative collision risk over a future period of time, and the risk decreases over time, avoiding system misjudgment due to the uncertainty of long-term prediction.

[0116] In addition, the edge cloud compares the calculated collision risk value of each blind spot object with a preset threshold. If the collision risk value is greater than or equal to the preset threshold, a collision warning is issued to the connected vehicle. At the same time, the position and movement status of the blind spot object are broadcast to help the connected vehicle avoid possible collisions in advance.

[0117] The working principle of the fusion perception and risk assessment method for collision warning of connected vehicles proposed in this application is illustrated below with a specific embodiment.

[0118] Figure 4 This is a flowchart illustrating the working principle of a fusion perception and risk assessment method for collision warning of connected vehicles according to an embodiment of this application.

[0119] Step S401: There is a blind spot ahead of the vehicle on the road.

[0120] In this embodiment of the application, when there is a blind spot in front of the vehicle on the road, a process for triggering a collision warning for connected vehicles is executed, and step S402 is performed.

[0121] Step S402: Perform multi-source information fusion sensing at the edge cloud, taking into account the impact of communication latency.

[0122] In this embodiment, the motion field of the feature sequence from different sources can be predicted based on the perception features from multiple sources, and the query vector can be updated according to the motion field to obtain the query vector after aggregating the historical perception features of each source. Then, the query vector after aggregating the historical perception features of each source is fused with multi-source perception features to obtain the fused query vector, and the fused perception result of each traffic participant is generated according to the fused query vector.

[0123] Step S403: The edge cloud identifies objects in the blind spot based on the fusion perception results and the vehicle detection results.

[0124] Step S404: Predict the future trajectory of objects in the blind zone at the edge cloud.

[0125] Step S405: The edge cloud determines the collision risk based on the future trajectory of objects in the blind spot.

[0126] In this embodiment of the application, the collision risk value of objects in the blind spot can be calculated by judging the collision risk.

[0127] Step S406: Is the collision risk value of the object in the blind spot greater than or equal to a preset threshold?

[0128] In this embodiment, it can determine whether the collision risk value of the blind spot object is greater than a preset threshold. If the collision risk value of the blind spot object is greater than or equal to the preset threshold, step S407 is executed. If the collision risk value of the blind spot object is less than the preset threshold, step S402 is executed.

[0129] Step S407: Issue collision warnings from the edge cloud.

[0130] The fusion perception and risk assessment method for collision warning of connected vehicles proposed in this application can predict the motion field based on the target coordinate system and the sampling time of multi-source perception features, update the query vector to aggregate historical perception features, reduce the dependence on previous perception algorithms, and improve robustness to communication latency. Simultaneously, by fusing multi-source perception features from the aggregated query vectors, fusion perception results for each traffic participant are generated, fully integrating the advantages of multiple sources to fill blind spots. Furthermore, by identifying the set of objects in the blind spot based on the fusion perception results, predicting future trajectories by combining historical trajectories, and associating with the connected vehicle navigation path, collision risks are identified, avoiding the evaluation bias of deterministic models and improving the reliability of risk warning. Therefore, this solves the problem that related technologies, due to direct fusion of detection results without considering the uncertainty of future trajectories of objects in the blind spot, suffer from weak robustness to communication latency, insufficient perception of potential traffic participants, and are prone to false alarms and missed alarms in risk assessment, thus restricting the usability of collision warning for connected vehicles.

[0131] Next, referring to the accompanying drawings, a fusion perception and risk assessment device for collision warning of connected vehicles proposed according to an embodiment of this application is described.

[0132] Figure 5 This is a block diagram of a fusion perception and risk assessment device for collision warning of connected vehicles provided according to an embodiment of this application.

[0133] like Figure 5 As shown, the fusion perception and risk assessment device 50 for collision warning of connected vehicles includes: a calculation module 100, a perception module 200, and a judgment module 300.

[0134] The calculation module 100 is used to predict the motion field of feature sequences from different sources based on the sampling time of the target coordinate system and the perception features from multiple sources, and to calculate and update the query vector according to the feature sequences and motion fields from different sources, so as to obtain the query vector after aggregating the historical perception features from each source.

[0135] The perception module 200 is used to perform multi-source perception feature fusion on the query vector after aggregating historical perception features from their respective sources to obtain a fused query vector, and generate the fused perception results of each traffic participant based on the fused query vector.

[0136] The judgment module 300 is used to identify the set of objects in the blind spot based on the fusion perception results, and predict the future trajectory based on the set of objects in the blind spot and the historical trajectory of each object in the blind spot, so as to identify the collision risk by using the navigation path and future trajectory uploaded by the connected vehicle.

[0137] Optionally, in one embodiment of this application, the computing module 100 includes: a first computing unit and a second computing unit.

[0138] The first computing unit is used to initialize a query vector based on the feature sequences from different sources in the perceptual features at the earliest acquisition time.

[0139] The second computing unit is used to calculate the corresponding position of the query vector in the perceived features at the next time step using the motion field, and to select key-value pairs based on the corresponding position to perform attention calculation and update the query vector.

[0140] Optionally, in one embodiment of this application, the update formula for the query vector may be, but is not limited to, the following: , in, For time index, The length of the historical information sequence stored at the edge cloud. For the first feature sequence The query vector at each time point. For the first feature sequence The query vector at time +1, For source The motion field of the characteristic sequence, for Source in Perceptual features acquired in real time and after coordinate alignment and time-delay encoding A function to update the position of the query vector in the perceived features at the next time step. This is a function that extracts the key vector from the feature vector at a given time. ATT is a function that extracts a value vector from a feature vector at a given time.

[0141] Optionally, in one embodiment of this application, the formula for calculating collision risk may be, but is not limited to: , in, For a moment, The time decay factor, For the identification of blind spot objects Collision risk value with connected vehicles, For the set of objects in the blind spot, For the first A blind spot object, The navigation route uploaded for connected vehicles. For the first i The future trajectory of an object in a blind spot For the first The uncertainty of the future trajectory of an object in a blind spot This is the function for determining the probability of trajectory overlap. This marks the start time for risk assessment. The duration of risk assessment.

[0142] Optionally, in one embodiment of this application, it further includes a calculation module and an early warning module.

[0143] The calculation module is used to calculate the collision risk value of objects in the blind spot.

[0144] The warning module is used to issue a collision warning to the connected vehicle when the collision risk value of an object in the blind spot is greater than or equal to a preset threshold, and to broadcast the position and movement status of the object in the blind spot.

[0145] Optionally, in one embodiment of this application, the formula for calculating the sports field may be, but is not limited to, the following: , in, For source The motion field of the characteristic sequence, for Source in Perceptual features acquired in real time and after coordinate alignment and time-delay encoding The length of the historical information sequence stored at the edge cloud. For dimension stacking operations, This is the input operation for spatiotemporal convolutional networks.

[0146] It should be noted that the foregoing explanation of the embodiment of the fusion perception and risk assessment method for collision warning of connected vehicles also applies to the fusion perception and risk assessment device for collision warning of connected vehicles in this embodiment, and will not be repeated here.

[0147] The fusion perception and risk assessment device for collision warning of connected vehicles proposed in this application can predict the motion field based on the target coordinate system and the sampling time of multi-source perception features, update the query vector to aggregate historical perception features, reduce the dependence on previous perception algorithms, and improve robustness to communication latency. Simultaneously, by fusing multi-source perception features from the aggregated query vectors, it generates fusion perception results for each traffic participant, fully integrating the advantages of multiple sources to fill blind spots. Furthermore, by identifying the set of objects in the blind spots based on the fusion perception results, predicting future trajectories by combining historical trajectories, and associating them with the navigation path of connected vehicles, it identifies collision risks, avoids the evaluation bias of deterministic models, and improves the reliability of risk warning. Therefore, it solves the problem that related technologies, due to direct fusion of detection results without considering the uncertainty of future trajectories of objects in the blind spots, suffer from weak robustness to communication latency, insufficient perception capability for potential traffic participants, and are prone to false alarms and missed alarms in risk assessment, thus restricting the usability of collision warnings for connected vehicles.

[0148] Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. The electronic device may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0149] When the processor 602 executes the program, it implements the fusion perception and risk assessment method for collision warning of connected vehicles provided in the above embodiments.

[0150] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.

[0151] The memory 601 is used to store computer programs that can run on the processor 602.

[0152] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0153] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0154] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0155] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0156] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described method for fusion perception and risk assessment for collision warning of connected vehicles.

[0157] This application also provides a computer program product, including a computer program that, when executed, implements the above-described method for fusion perception and risk assessment for collision warning of connected vehicles.

[0158] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific sensing features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific sensing features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the sensing features of different embodiments or examples.

[0159] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technically perceived features indicated. Therefore, a perceived feature defined with "first" or "second" may explicitly or implicitly include at least one of that perceived feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0160] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0161] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0162] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0163] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0164] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0165] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A fusion perception and risk judgment method for connected vehicle collision pre-warning, characterized in that, The method comprises the following steps: Based on the target coordinate system and the sampling time of the perception features from multiple sources, the motion field of the feature sequences from different sources is predicted, and the query vector is updated according to the feature sequences from different sources and the motion field to obtain the query vector after the historical perception features of each source are aggregated; Multi-source perception feature fusion is performed on the query vector after the historical perception features of each source are aggregated to obtain a fused query vector, and a fused perception result of each traffic participant is generated according to the fused query vector; Based on the fused perception result, a blind area object set is identified, and a future trajectory is predicted according to the blind area object set and the historical trajectory of each blind area object, so as to identify a collision risk by using the navigation path uploaded by the connected vehicle and the future trajectory.

2. The method of claim 1, wherein, The calculation formula of the motion field is: , wherein, is the motion field of the characteristic sequence of origin , is the motion field of the characteristic sequence of origin , is the perception feature of origin at time instant after coordinate alignment and time delay encoding, is the length of the historical information sequence stored in the edge cloud, is the dimension stacking operation, is the spatio-temporal convolution network input operation.

3. The method of claim 1, wherein, The updating of the query vector according to the feature sequences from different sources comprises: Initializing the query vector in the perception feature at the earliest collection time based on the feature sequences from different sources; Using the motion field to calculate the corresponding position of the initialized query vector in the perception feature at the next time, and updating the query vector based on the attention calculation of the key-value pair vector selected according to the corresponding position.

4. The method of claim 3, wherein, The updating formula of the query vector is: , wherein, is the time index, is the length of the historical information sequence stored in the edge cloud, is the query vector of the time point in the feature sequence, is the query vector of the +1 time point in the feature sequence, is the motion field of the feature sequence from the source, is the perception feature collected at the time point from the source after coordinate alignment and time delay encoding, is a function for updating the position of the query vector in the perception feature at the next time point, is a function for extracting the key vector from the feature vector at a certain time point, is a function for extracting the value vector from the feature vector at a certain time point, and is an attention calculation network.

5. The method of claim 1, wherein, The calculation formula of the collision risk is: , wherein, is a time, is a time decay factor, is a blind zone object identified is a collision risk value with a connected vehicle, is a set of blind zone objects, is a first blind zone object, is a navigation path uploaded by the connected vehicle, is the future trajectory of a first blind zone object, is the uncertainty of the future trajectory of a first blind zone object, is a trajectory overlap probability judgment function, is a start time of risk judgment, is a duration of risk judgment.

6. The method of claim 1, wherein, Further comprising: Calculating the blind area object collision risk value of the collision risk; In the case where the blind area object collision risk value is greater than or equal to a preset threshold, a collision warning is issued to the connected vehicle, and the position and motion state of the blind area object are broadcasted.