A car-car cooperative perception method based on blind area feature on-demand request

CN122830744APending Publication Date: 2026-09-29BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610979642.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

持续广播和接收全量特征,不仅仍然占据了较高的通信信道,更严重消耗了车端边缘计算单元的车载单元(On-Board Unit,OBU)算力,造成了大量的无用计算功耗

Benefits of technology

在本说明书提供的基于盲区特征按需请求的车车协同感知方法中,目标车辆先对环境进行采集得到BEV特征图,再利用目标车辆在未来时间窗口的预期行驶轨迹和物理遮挡物的位置定位出可能的高风险盲区多边形区域,然后把该区域信息通过向邻近的协同车辆发送,以使协同车辆根据协同感知请求数据包确定仅包含高风险盲区多边形区域的深度语义信息的轻量化隐式特征表示,然后将局部隐式特征张量与第一鸟瞰视角BEV特征图进行融合,从而确定目标车辆的控制指令。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122830744A_ABST
    Figure CN122830744A_ABST
Patent Text Reader

Abstract

The application discloses a car-car cooperative perception method based on blind area feature on-demand request and relates to the technical field of automatic driving. According to the environmental data collected by a target vehicle, a first bird's-eye view (BEV) feature map is generated, and the position of a physical shelter of the target vehicle is identified according to the first BEV feature map; according to the expected driving track of the target vehicle and the position of the physical shelter, a high-risk blind area polygon is determined; the vertex coordinates of the high-risk blind area polygon, the global positioning information of the target vehicle and a timestamp are extracted to form a cooperative perception request data packet, and the cooperative perception request data packet is sent to cooperative vehicles within the communication radius of the target vehicle; a local implicit feature tensor returned by the cooperative vehicles is received; the local implicit feature tensor is fused with the first bird's-eye view BEV feature map to obtain global features, potential obstacle state prediction of the high-risk blind area is performed through the global features, and control instructions of the target vehicle are determined. The method reduces the computing power consumption of cooperative perception of the automatic driving vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to a vehicle-to-vehicle cooperative perception method based on on-demand requests for blind spot features. Background Technology

[0002] With the continuous development of advanced autonomous driving technology, the performance of single-vehicle intelligent environmental perception systems is also becoming increasingly sophisticated. Because electromagnetic waves and optical signals are linear, the phenomenon of "line-of-sight occlusion" is becoming increasingly severe in complex urban roads and highways. In blind spots of large, heavy-duty vehicles or in areas with dense buildings and narrow roads, the sensor's field of view (FOV) can be effectively blocked, creating significant perception blind spots. These blind spots can lead to serious dangerous incidents such as pedestrians suddenly appearing out of sight, posing a direct threat to traffic safety. Pedestrians or other road users may be outside the range of regular monitoring, and the delayed response of the onboard perception and decision-making system reduces its ability to handle emergencies, thus exacerbating the risk situation.

[0003] To address the inherent limitations of single-vehicle intelligent perception capabilities, both industry and academia have focused on vehicle-to-external information interaction as a primary research direction. Collaborative perception solutions relying on data sharing between onboard devices have garnered significant attention due to their unique advantages in resolving environmental occlusion issues. Current research on vehicle-to-vehicle collaborative perception primarily concentrates on the core technology of data fusion. To balance bandwidth consumption with the preservation of environmental context semantics, the latest cutting-edge research proposes feature-level fusion based on bird's-eye view (BEV). The neural network of the collaborating vehicle extracts deep implicit feature tensors of the environment, compresses them, and broadcasts them. However, existing technologies all employ a passive mode of "indiscriminate full broadcasting," sending all features from the omnidirectional field of view. In reality, in most driving scenarios, over 80% of the target vehicle's field of view is unobstructed, with only a few specific occluded areas requiring collaborative supplementation. Continuously broadcasting and receiving all features not only still occupies a significant amount of communication channel space but also severely depletes the computing power of the on-board unit (OBU) of the vehicle's edge computing unit, resulting in substantial amounts of wasted computational power. Summary of the Invention

[0004] Therefore, it is necessary to provide a vehicle-to-vehicle cooperative perception method based on on-demand requests using blind spot features to address the aforementioned technical problems.

[0005] The following technical solution is adopted in this specification: This specification provides a vehicle-to-vehicle cooperative perception method based on blind spot feature on-demand requests, including: Based on the environmental data collected from the target vehicle, a first bird's-eye view BEV feature map is generated, and the location of physical obstructions of the target vehicle is identified based on the first BEV feature map. Based on the current kinematic state of the target vehicle, predict the expected driving trajectory of the target vehicle in the future time window, and determine the high-risk blind zone polygon based on the expected driving trajectory and the position of physical obstructions; the high-risk blind zone polygon represents the area where the expected driving trajectory of the target vehicle overlaps with the blind zone behind the physical obstruction. Extract the vertex coordinates of the high-risk blind zone polygon, the global positioning information of the target vehicle, and the timestamp to form a collaborative perception request data packet, and send it to the collaborative vehicle within the communication radius of the target vehicle. Receive the local implicit feature tensor returned by the cooperative vehicle; the local implicit feature tensor is a lightweight implicit feature representation that contains only the deep semantic information of the high-risk blind zone polygon region, as determined by the cooperative vehicle based on the cooperative perception request data packet. The local implicit feature tensor is fused with the BEV feature map from the first bird's-eye view to obtain global features. The state of potential obstacles in high-risk blind spots is predicted using global features to determine the control commands for the target vehicle.

[0006] Optionally, based on the expected driving trajectory and the location of physical obstructions, a high-risk blind spot polygon is determined, including: Based on the location of the physical occluder, extract the set of visible edge points of the physical occluder in the first BEV feature map; Using the optical center of the target vehicle's sensing sensor as the origin of the light projection, a cluster of rays is emitted into the surrounding unknown space through the set of visible edge points; Based on the geometric intersection data of the ray cluster and the passable road plane, an initial view frustum occlusion space model is established; the initial view frustum occlusion space model represents the three-dimensional view frustum blind zone behind the physical occlusion object; By processing the expected driving trajectory envelope and the initial visual cone occlusion space model through Boolean intersection operations, the overlapping sub-regions of the two are identified as actual high-risk blind spots; The actual high-risk blind spot is projected vertically downwards onto the horizontal ground plane where the target vehicle is located. The set of boundary vertices of the projected area is extracted to generate a closed high-risk blind spot polygon.

[0007] Optionally, the structure of the data fields within the collaborative awareness request data packet includes: Frame header identifier field; The unique identification code field for the target vehicle; The target vehicle's absolute pose field includes the longitude, latitude, altitude, and three-dimensional heading angle matrix output by the fusion of the global navigation satellite system and the inertial measurement unit, i.e., global positioning information; The time synchronization field records the absolute timestamp of the target vehicle when the high-risk blind spot polygon was generated. The polygon vertex sequence field contains the coordinates of multiple two-dimensional vertices of the high-risk blind spot polygon in the target vehicle's local coordinate system. The target feature hierarchy parameters include parameters such as which target information the collaborative vehicle needs to extract from the neural network structure and the depth position of each layer.

[0008] Optionally, before fusing the local implicit feature tensor with the first-view BEV feature map, the method further includes: Extract the generation timestamp carried by the local implicit feature tensor; The time difference between the current system timestamp and the generated timestamp of the target vehicle; If the time difference is greater than the preset maximum communication delay threshold, the local implicit feature tensor is discarded, and the target vehicle is controlled to drive in the blind area at a preset speed, which is less than the first speed threshold. If the time difference is less than or equal to the maximum communication delay threshold, then based on the obstacle-level prior velocity vector sent by the cooperative vehicle, spatial translation compensation based on the optical flow field assumption is performed on the feature pixels in the local implicit feature tensor; the spatially translated local implicit feature tensor is used for feature fusion.

[0009] Optionally, based on the obstacle-level prior velocity vector sent by the cooperative vehicle, spatial translation compensation based on the optical flow field assumption is performed on the feature pixels in the local implicit feature tensor, including: The translational displacement vector is determined based on the obstacle-level prior velocity vector and the time difference. By spatially translating the feature pixels in the local implicit feature tensor using a translation vector, we obtain the spatially translated and compensated local implicit feature tensor.

[0010] Optionally, the local implicit feature tensor is fused with the first bird's-eye view BEV feature map to obtain global features, including: The time decay term is determined based on the time difference between the current system timestamp and the generated timestamp of the target vehicle; The spatial attenuation term is determined based on the physical straight-line distance between the perception sensor of the cooperative vehicle and the geometric center of the high-risk blind zone polygon and the maximum effective detection range of the perception sensor of the cooperative vehicle; By weighting the time decay term and the spatial decay term, the dynamic weights corresponding to the local implicit feature tensor are obtained; The product of the dynamic weights and the local implicit feature tensor is determined as the weighted local implicit feature tensor; The weighted local implicit feature tensor is aligned in the spatial dimension and filled into the coordinate pixel missing region of the corresponding high-risk blind zone polygon in the first BEV feature map to obtain the global feature.

[0011] Optionally, the cooperative vehicle performs the following steps when generating the local implicit feature tensor: Receive the cooperative perception request data packet broadcast or multicast from the target vehicle, and perform protocol parsing on the cooperative perception request data packet to extract the target vehicle's global pose matrix, the target vehicle's absolute timestamp, and the vertex coordinate set of the high-risk blind zone polygon. The sensor environment perception data stream of the cooperative vehicle is acquired, and multi-scale features of the sensor environment perception data stream are extracted through a pre-deployed BEV backbone network to generate a second BEV feature map in the local coordinate system of the cooperative vehicle. Based on the global pose matrix of the cooperative vehicle, the relative pose affine transformation matrix between the target vehicle and the cooperative vehicle is solved. The vertex coordinates of the high-risk blind zone polygon are mapped to the pixel coordinate system of the second BEV feature map using the relative pose affine transformation matrix to obtain the projected polygon. Based on the projected polygon, the second BEV feature map is cropped to retain the effective feature region inside the projected polygon. A multi-channel fusion approach is used to selectively converge effective feature regions, and local implicit feature vectors are obtained through quantization and dimensionality reduction.

[0012] Optionally, based on the global pose matrix of the cooperating vehicle, the relative pose affine transformation matrix between the target vehicle and the cooperating vehicle is solved. The vertex coordinates of the high-risk blind spot polygon are then mapped to the pixel coordinate system of the second BEV feature map using the relative pose affine transformation matrix, resulting in a projected polygon, including: The product of the inverse of the global pose matrix of the cooperative vehicle and the global pose matrix of the target vehicle is determined as the relative pose affine transformation matrix. The fixed-scale scaling from the local physical coordinate system of the cooperative vehicle to the pixel coordinate system of the feature map, the product of the translation intrinsic parameter matrix, the relative pose affine transformation matrix, and the vertex coordinates of the high-risk blind zone polygon, are used to map the vertex coordinates of the high-risk blind zone polygon to the pixel coordinates of the second BEV feature map. Construct a projection polygon based on the mapped pixel coordinates.

[0013] Optionally, the second BEV feature map is cropped based on the projected polygon, retaining the effective feature regions within the projected polygon, including: Generate a two-dimensional spatial mask matrix with the same size as the spatial resolution of the second BEV feature map; Iterate through all pixel coordinates in the two-dimensional spatial mask matrix. If the point is located inside or on the boundary of the projected polygon, assign the element value corresponding to the pixel coordinates to 1; otherwise, assign the value to 0. Perform a Hadamard product operation between the two-dimensional spatial mask matrix and the second BEV feature map to obtain the clipped feature tensor, which is the effective feature region.

[0014] Optionally, a multi-channel fusion approach is used to selectively converge effective feature regions, and local implicit feature vectors are obtained through quantization and dimensionality reduction, including: Extract the subtensor within the smallest bounding rectangle region composed of the non-zero elements of the clipped feature tensor; Dimensionality reduction of the sub-tensor by channel dimension is performed using a pre-trained 1x1 convolutional kernel to obtain floating-point feature data; By quantization mapping and encoding conversion, floating-point feature data is mapped to INT8 integer representation to obtain local implicit feature vectors.

[0015] This specification provides a vehicle-to-vehicle cooperative perception device based on blind spot feature on-demand request, including: The generation module is used to generate a first bird's-eye view BEV feature map based on the environmental data collected from the target vehicle, and to identify the location of physical obstructions of the target vehicle based on the first BEV feature map. The determination module is used to predict the expected driving trajectory of the target vehicle in a future time window based on the current kinematic state of the target vehicle, and to determine the high-risk blind zone polygon based on the expected driving trajectory and the position of physical obstructions; the high-risk blind zone polygon represents the area where the expected driving trajectory of the target vehicle overlaps with the blind zone behind the physical obstruction. The sending module is used to extract the vertex coordinates of the high-risk blind zone polygon, the global positioning information of the target vehicle, and the timestamp, form a collaborative perception request data packet, and send it to the collaborative vehicle within the communication radius of the target vehicle. The receiving module is used to receive the local implicit feature tensor returned by the cooperative vehicle; the local implicit feature tensor is a lightweight implicit feature representation that contains only the deep semantic information of the high-risk blind zone polygon region, as determined by the cooperative vehicle based on the cooperative perception request data packet. The prediction module is used to fuse the local implicit feature tensor with the BEV feature map from the first bird's-eye view to obtain global features. The global features are then used to predict the state of potential obstacles in high-risk blind spots and determine the control commands for the target vehicle.

[0016] This specification provides a vehicle-to-vehicle cooperative perception system based on blind spot feature on-demand request, which is used to implement the above-mentioned vehicle-to-vehicle cooperative perception method based on blind spot feature on-demand request.

[0017] This specification provides an electronic device for implementing a vehicle-to-vehicle cooperative perception method based on blind spot feature on-demand requests.

[0018] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described vehicle-to-vehicle cooperative perception method based on blind spot feature on-demand requests.

[0019] This specification provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described vehicle-to-vehicle cooperative perception method based on blind spot features and on-demand requests.

[0020] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects: In the vehicle-to-vehicle cooperative perception method based on blind spot feature on-demand request provided in this specification, the target vehicle first collects environmental data to obtain a BEV feature map. Then, it uses the expected driving trajectory of the target vehicle in the future time window and the position of physical obstructions to locate possible high-risk blind spot polygonal regions. The information of this region is then sent to nearby cooperative vehicles so that the cooperative vehicles can determine a lightweight implicit feature representation containing only the depth semantic information of the high-risk blind spot polygonal regions based on the cooperative perception request data packet. Then, the local implicit feature tensor is fused with the BEV feature map from the first bird's-eye view to determine the control command of the target vehicle.

[0021] This method first predicts high-risk blind spots and then initiates on-demand requests only to those areas. While ensuring or even improving the ability to perceive sudden dangers such as ghost peeks, it significantly reduces the communication bandwidth usage and edge computing power consumption in vehicle-to-vehicle collaboration, solving the resource waste problem caused by the "full broadcast" of existing technologies. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0023] Figure 1 This document provides a flowchart illustrating a vehicle-to-vehicle cooperative perception method based on on-demand requests using blind spot features. Figure 2 This manual provides a polygonal diagram illustrating a high-risk blind spot in a typical traffic scenario. Figure 3 This document provides a flowchart of a collaborative vehicle-side feature extraction process. Figure 4is a schematic flow diagram of another vehicle-vehicle collaborative perception method based on blind spot feature on-demand request provided in this specification; Figure 5 is a schematic diagram of a computer device for implementing the vehicle-vehicle collaborative perception method based on blind spot feature on-demand request provided in this specification. DETAILED DESCRIPTION

[0024] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of the present application will be described clearly and completely below with reference to specific embodiments of this specification and the corresponding accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0025] The ordinal numbers such as first and second in this patent text do not have priority or significance, and are only symbols used to distinguish different objects and concepts. Without special contextual constraints, the words "a", "an" and "the" should be interpreted as covering both singular and plural meanings, which conforms to the actual use of Chinese vocabulary and the inherent logical rules.

[0026] Current research on vehicle-vehicle collaborative perception mainly focuses on the core technology of data fusion, and can be roughly divided into three categories. However, current research still has many technical difficulties in practical applications.

[0027] The bandwidth and latency disaster existing in raw data level fusion. Early fusion schemes require cooperating vehicles to directly send the raw panoramic video streams or high-density 3D laser point cloud data collected by their sensors to the target vehicle through wireless networks. However, due to the extremely high requirements for environmental resolution in automatic driving, the data throughput generated by a single vehicle is often as high as thousands of gigabits per second (Gbps). Under the existing C-V2X (Cellular V2X) or 5G communication protocols, uplink and downlink bandwidth resources are extremely limited. If multiple vehicles broadcast raw data simultaneously in a dense traffic environment, it will instantly cause channel congestion and serious packet loss. The high data transmission delay (here refers to the total time consumed from data collection, packaging, sending, receiving to unpacking) directly undermines the millisecond-level real-time performance required by automatic driving systems, making this scheme almost impractical in engineering.

[0028] Semantic loss and ghosting illusions caused by target-level fusion. To avoid bandwidth bottlenecks, mass-produced vehicles generally tend to adopt post-fusion schemes. In this scheme, the cooperating vehicle first uses its own computing power to identify obstacles, and then only sends lightweight structural data such as the generated target BBox (Bounding Box) coordinates, category, and speed to the target vehicle. Although this scheme uses very little bandwidth, it faces two fatal flaws: First, the complete loss of environmental semantics. The target-level data only contains the results of independent objects, stripping away road topology, weather lighting, and features of "long-tailed" irregular obstacles that are not recognized by the neural network. If the algorithm of the cooperating vehicle misses a detection, the target vehicle will receive a false safety illusion. Second, severe coordinate system alignment errors. This scheme is extremely dependent on the absolute positioning accuracy of GNSS. In urban canyon or tunnel environments, the absolute positioning error between the two vehicles (referring to the spatial deviation between the actual physical coordinates of the vehicle and the satellite positioning coordinates) often reaches decimeters or even meters. When the target vehicle forcibly transforms the cooperative target BBox with positioning errors to its own coordinate system, the same real physical obstacle will be mistaken by the system as two spatially separated objects, producing "ghosting objects", which in turn causes the autonomous driving system to trigger dangerous phantom braking commands.

[0029] Full feature-level fusion leads to wasted computing power and redundant transmission. To balance bandwidth consumption and preservation of environmental context semantics, recent cutting-edge research has proposed feature-level fusion based on BEVs (Battery-Electric Vehicles). The collaborative vehicle's neural network extracts deep implicit feature tensors of the environment, compresses them, and broadcasts them. However, existing technologies all employ a passive "indiscriminate full broadcast" mode, sending all features from the omnidirectional view. This severely consumes the computing power of the vehicle's edge computing unit (OBU), resulting in a large amount of useless computational power consumption.

[0030] Currently, the inability to simultaneously achieve optimal communication bandwidth, sensing accuracy, and computational latency has disrupted traditional either-or thinking. When wireless channel resources are limited and onboard computing power is weak, ensuring the real-time and reliable transmission of critical data to its destination becomes a significant challenge for achieving high-level autonomous driving. This invention proposes a new technical solution to address these issues, aiming to advance the technology in this field.

[0031] To address the functional limitations of traditional single-vehicle intelligent systems due to limited field of view, and the numerous problems encountered in the application of V2V collaborative perception technology, such as the excessive network bandwidth consumption caused by raw data broadcasting, semantic distortion and visual interference caused by target-level fusion, and the high computational resource consumption of feature-level full data transmission, this invention proposes a vehicle-to-vehicle collaborative perception method based on on-demand requests for blind zone features. As autonomous driving technology matures, the resolution of onboard sensors is increasing, resulting in a massive amount of perception data that exceeds the Shannon Limit of the physical layer of the current C-V2X communication protocol based on the PC5 interface. Using full or feature-level indiscriminate transmission leads to technical challenges such as link congestion and increased latency. This invention utilizes spatial ray tracing algorithms, kinematic model correlation analysis, and binary mask filtering techniques to transform the traditional, non-directional full data transmission method into a directional information transmission mode based on the demand for local features in high-risk blind zones. This scheme ensures the accuracy of underlying physical characteristics and environmental semantics, achieves higher data compression ratios, and enables efficient multimodal information fusion and potential obstacle detection even in limited communication environments. This method can improve perception quality and computing speed, and enhance the overall stability and reliability of the system. It has great application potential and broad prospects in the field of intelligent transportation.

[0032] The main objective of this invention is to create a multi-vehicle cooperative perception technology architecture that combines logical closed-loop control and spatiotemporal redundancy verification. This invention primarily studies the target vehicle (the main entity requesting information), the specific operation mode of the core algorithm, and a detailed analysis of its various technical and theoretical foundations.

[0033] This invention addresses the major problems of low bandwidth, high data redundancy, and high transmission latency in vehicle-to-vehicle (V2V) communication. It employs a combination of blind-spot polygon space clipping and channel-level quantization compression to significantly reduce the proportion of invalid background information and improve the data transmission method of network tensors in high-risk areas. This design overcomes the limitations of traditional transmission methods and reduces the probability of signal loss due to overload, exhibiting excellent resolution and rapid response capabilities in complex or constantly changing traffic conditions. The established real-time communication framework provides safety and reliability for autonomous driving systems, ensuring timely emergency obstacle avoidance during dynamic decision-making. This achievement has significant practical implications and development potential for improving the overall performance of in-vehicle intelligent interconnection systems.

[0034] This invention primarily studies the key issue of ghosting in the field of collaborative perception, aiming to improve the fidelity of semantic information. The proposed method creatively incorporates hidden network feature-level sharing on top of the traditional hierarchical shared target architecture. This mechanism can preserve the texture features of the entire blind zone depth environment, the spatial structure and topological relationships of irregularly shaped objects, and uses a relative homogeneous transformation matrix to align the feature map spaces between different networks. The positioning system designed based on feature tensor dimension fusion exhibits good robustness and anti-interference capabilities, especially for GPS absolute positioning errors. During subsequent decoding, an attention mechanism is used to improve pixel-level alignment, fusing multi-view blind zone features together, thereby significantly reducing the risk of missed or false detections.

[0035] This invention designs an adaptive fault-tolerant algorithm that combines kinematic trajectory analysis and optical flow delay correction to improve the safety of intelligent driving systems. It utilizes 3D geometric modeling technology to identify and analyze safety risks in the target area, making timely adjustments based on actual conditions to ensure reliable operation. From an information processing perspective, it establishes a method combining optical flow field displacement and confidence level weighting to measure data quality and provides the system with the ability to quantitatively analyze the reliability of externally perceived data, thus providing a scientific basis for decision-making. When sensor data exhibits delays or distortion, the adaptive optimization mechanism can immediately change the weight parameters to eliminate misjudgments caused by environmental interference and communication instability, improving the robustness and stability of the autonomous driving system in complex environments.

[0036] This technical solution relies on event-triggered, on-demand request-driven methods to intelligently allocate edge computing resources from the in-vehicle environment, breaking through the inherent constraints of traditional all-time broadcast distribution methods. Its core structure adopts a heterogeneous computing platform design. When no task requests are received from surrounding high-risk blind spots or have not been prioritized, the underlying heterogeneous computing units cease tensor processing operations on the global feature data. Only when there are specific scenario requirements will the system execute matching polygon pose mapping and mask clipping algorithms. This solution significantly improves the concurrent processing capabilities of the in-vehicle computing domain controller. By optimizing energy consumption management and heat dissipation performance, it provides strong technical support for the efficient operation of intelligent connected vehicles, possessing significant practical value and promising prospects for widespread adoption.

[0037] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0038] Figure 1 This is a flowchart illustrating a vehicle-to-vehicle cooperative perception method based on blind spot features and on-demand requests, as described in this specification. The method, with the target vehicle as the executing entity, specifically includes the following steps: S101, Based on the environmental data collected from the target vehicle, generate a first bird's-eye view BEV feature map, and identify the location of physical obstructions of the target vehicle based on the first BEV feature map.

[0039] In one embodiment, after processing the raw data obtained from the multimodal sensors of the target vehicle through an edge-side feature extraction network, a first BEV feature map based on a global perspective is obtained. Based on this first BEV feature map, the physical location and attributes of physical obstructions in the environment surrounding the target vehicle can be found.

[0040] Specifically, multimodal sensors include visual sensors and LiDAR. This invention mainly focuses on the core issues of multimodal spatiotemporal perception technology for target vehicles, primarily studying its role in creating the first-layer BEV (bird's-eye view) feature map. The target vehicle is equipped with a perception terminal integrating heterogeneous sensors, which can collect environmental data in real time. The visual camera captures high-resolution RGB images, while the LiDAR generates three-dimensional depth information. The raw perception data is transmitted to a pre-defined deep learning platform for processing. Multi-level semantic features are extracted using a two-dimensional backbone network based on ResNet or SwinTransformer architecture, and then converted into three-dimensional spatial representations using LSS or multi-view mapping. For LiDAR point cloud data, voxelization and three-dimensional sparse convolution analysis are performed first. The two paths are fused in the target vehicle's local coordinate system to obtain the first BEV feature map with temporal characteristics. This feature map can support subsequent target detection and tracking tasks and, together with the foreground segmentation module, accurately locate important obstacles in the field of view, such as large cargo, green plants, and buildings.

[0041] The target vehicle simultaneously acquires environmental images and point cloud data using a multimodal sensor suite. Relying on hardware-level synchronization through precise timestamps, the raw perception data is transmitted to a pre-deployed deep neural network model on the vehicle for forward computation. After these steps, a first BEV feature map is generated in the local physical coordinate system, providing target detection results, semantic segmentation labels, and precise localization information for occluded objects in the field of view. These functions provide crucial information for path planning and obstacle avoidance decisions in the autonomous driving system, significantly improving system performance and safety.

[0042] S102, based on the current kinematic state of the target vehicle, predict the expected driving trajectory of the target vehicle in the future time window, and determine the high-risk blind zone polygon based on the expected driving trajectory and the position of the physical obstruction; the high-risk blind zone polygon represents the area where the expected driving trajectory of the target vehicle overlaps with the blind zone behind the physical obstruction.

[0043] In this embodiment, based on the physical positioning data acquired by the vehicle-mounted sensors and the pre-defined three-dimensional geometric structure, spatial beam tracing is used to accurately predict and depict high-collision-hazard areas in the three-dimensional environment. A projection transformation algorithm is then used to map the three-dimensional blind spots onto the first bird's-eye view feature map plane, accurately representing its boundary contours in polygon form, thereby completing the visual representation and quantitative analysis of the risk areas.

[0044] Specifically, based on the current kinematic state of the target vehicle, the expected driving trajectory of the target vehicle in a future time window is predicted, including: obtaining the current chassis kinematic state of the target vehicle, and using a vehicle kinematic single-track model to predict the target vehicle's trajectory in the future prediction time window. The expected driving trajectory envelope within.

[0045] To avoid requesting area features from cooperating vehicles that pose no substantial threat to the target vehicle's driving safety, this invention further introduces a vehicle kinematics prediction model based on Ackermann Steering Geometry. The target vehicle's current longitudinal speed is acquired in real-time via the onboard CAN (Controller Area Network) bus. Longitudinal acceleration, yaw rate and steering wheel angle Wait for the real-time status of the chassis, based on the vehicle wheelbase Calculate the instantaneous turning radius of the target vehicle. Based on the aforementioned kinematic state parameters, the centerline of the target vehicle's expected driving trajectory within the future prediction time window is deduced. Subsequently, this centerline is extended laterally by half the vehicle width on both sides to generate the envelope of the expected future driving trajectory.

[0046] This step involves the derivation of high-risk blind zone polygons based on ray tracing and kinematic coupling. This step aims to accurately isolate core blind zones that pose a potentially lethal threat to the target vehicle from a vast, unknown physical space. In one embodiment, the high-risk blind zone polygon is determined based on the expected driving trajectory and the location of physical obstructions, including the following steps:

[0047] S201, Based on the location of the physical occlusion, extract the set of visible edge points of the physical occlusion in the first BEV feature map. ,in The number of edge points, , and These are the pixel values ​​of the edge points in the first BEV feature map.

[0048] like Figure 2 As shown, Figure 2 In the given typical traffic scenario, when the target vehicle is traveling on the road, a large truck in front is the main physical obstruction, significantly interfering with the target vehicle's sensors' perception of the environment ahead. This truck itself creates a large blind spot, preventing potential threats from being detected. Furthermore, the sensor systems of other cooperating vehicles in adjacent lanes or oncoming lanes have their fields of view extending into this blind spot, successfully detecting the presence of the unknown obstacle and thus exposing the target vehicle's current limitations in perception.

[0049] The core of this embodiment lies in the target vehicle actively performing geometric spatial deduction of high-risk blind spots. First, the target vehicle acquires the three-dimensional coordinates of the physical installation of its main sensors. In this coordinate system, These represent the lateral, longitudinal, and vertical height coordinates of the sensor's optical center within the target vehicle's ISO 8855 standard local coordinate system. Using the initial environmental feature map generated from its own sensor data, the target vehicle employs a foreground semantic segmentation algorithm to extract the visible edge contours of physical occlusions, forming a set of edge points in three-dimensional space.

[0050] S202, using the optical center of the target vehicle's perception sensor as the origin of the light projection, passes through the set of visible edge points. It emits clusters of rays into the unknown outer space.

[0051] Specifically, the target vehicle domain controller uses a spatial ray tracing model to extrapolate the view frustum occlusion space. Using coordinates... (abbreviation) ( ) serves as the origin of the ray projection, passing through the set of edge points A cluster of rays is emitted outwards. The three-dimensional spatial parameterization equation for a single ray is defined as follows:

[0052] in, Represents the three-dimensional spatial coordinates of any point along the direction of ray projection. The step size is a scalar parameter and This is used to control the distance the ray extends in physical space, and its maximum value is determined by the maximum effective ranging threshold of the target vehicle sensor. Indicates starting from the origin Pointer to edge point set The three-dimensional normalized direction vector of a specific edge point.

[0053] S203. Based on the geometric intersection data of the ray cluster and the passable road plane, an initial view frustum occlusion space model is established; the initial view frustum occlusion space model represents the three-dimensional view frustum blind zone behind the physical occlusion. The initial view frustum occlusion space model is used to simulate the structure of the three-dimensional scene behind the physical obstacle.

[0054] By solving a system of equations, the geometric intersection between the ray cluster and the pre-fitted road drivable plane equation is calculated, thereby constructing an initial visual cone occlusion space that is located behind the occlusion and appears radial.

[0055] S204 uses Boolean intersection operations to process the expected driving trajectory envelope and the initial view cone occlusion space model, and identifies the overlapping sub-regions as actual high-risk blind spots.

[0056] By performing a spatial Boolean intersection operation between the expected driving trajectory envelope and the aforementioned initial visual cone occlusion space, only the overlapping sub-region where spatial interference occurs is extracted as the actual high-risk blind spot that truly threatens driving safety. This operation precisely reduces the vast physical blind spot to the "high-risk blind spot" that truly threatens the future driving safety of the target vehicle.

[0057] S205 projects the actual high-risk blind spot vertically downwards onto the horizontal ground plane where the target vehicle is located. On the BEV reference plane, extract the set of boundary vertices of the projection region to generate a closed high-risk blind zone polygon. .

[0058] Optionally, the three-dimensional high-risk blind zone can be projected downwards along the direction of gravity to... On the reference ground plane, extract the outer contour boundary (two-dimensional convex hull or concave hull polygon outer boundary) of the projection area to generate a completely closed high-risk blind zone polygon. .

[0059] Please continue reading Figure 2 The polygon clearly defines the precise physical boundary of the local features requested by the target vehicle. The target vehicle initiates V2X on-demand requests to the cooperating vehicle only for the polygon area. After processing, the cooperating vehicle sends the local features corresponding to the polygon area back to the target vehicle.

[0060] In this embodiment, based on the extraction of the occlusion contour and the acquisition of the external attitude parameters of the vehicle-mounted sensors, a method combining ray tracing models and kinematic equations is proposed for the first time to determine the location of dangerous blind spots and use it for localization. The sequence of boundary vertices on the two-dimensional plane is analyzed, and a polygonal geometry algorithm is used to draw the distribution of potential risk areas.

[0061] S103: Extract the vertex coordinates of the high-risk blind zone polygon, the global positioning information of the target vehicle, and the timestamp, form a collaborative perception request data packet, and send it to the collaborative vehicle within the communication radius of the target vehicle.

[0062] The global positioning information can be collected from onboard equipment. Specifically, a data sequence is generated, including the vertex coordinates of the high-risk blind spot polygon, the real-time global positioning information of the onboard equipment, and the system timestamp, to obtain a complete cooperative perception request payload. Using vehicle-to-everything (V2X) communication technology, the cooperative perception request payload is sent to the cooperating vehicles via broadcast or multicast within a specified communication radius.

[0063] The communication radius is a configurable distance threshold defined for the target vehicle. It is the effective communication distance threshold radiating outward from the target vehicle. It indicates that the target vehicle only wants to interact with other vehicles within this radius, thereby achieving a balance between communication effectiveness and network load control.

[0064] Serialization and targeted broadcasting of collaborative perception request packets: To ensure standardized and efficient parsing of V2X communication, the target vehicle converts the generated high-risk blind spot polygons into standardized on-demand request commands. Specifically, the polygons are... The two-dimensional vertex coordinate sequence, the high-precision absolute pose affine matrix output based on deep coupling of GNSS and IMU, and the hardware-level absolute timestamp when the polygon was generated. Extraction is performed. Simultaneously, to inform the collaborating vehicles of the required network layer for the target vehicle, a desired feature depth field is appended to the request. This information is then fed into the protocol stack, serialized and encapsulated according to a specific application-layer communication protocol, generating a collaborative perception request data packet with error correction verification codes. This data packet is transmitted via the underlying radio frequency antenna of the onboard V2X communication module, either through broadcast or geofencing-based multicast, to all potential collaborating vehicles within a surrounding communication distance threshold.

[0065] Optionally, the specific structure of the data fields inside the collaborative awareness request data packet is as follows: Frame header identifier field: used to distinguish between regular BSM (Basic Safety Message) broadcasts and the on-demand request messages of this invention; The target vehicle's unique identification code field contains the target vehicle's unique MAC (Media Access Control) address or digital certificate serial number; The target vehicle's absolute pose field includes the longitude, latitude, and altitude output by the fusion of the Global Navigation Satellite System (GNSS) and Real-time Kinematic (RTK) technology, as well as the three-dimensional heading angle matrix represented by Euler angles or quaternions provided by the Inertial Measurement Unit (IMU), which is the global positioning information. The time synchronization field records the absolute timestamp of the target vehicle when the high-risk blind spot polygon was generated. With accuracy down to the millisecond level; Polygon vertex sequence field, containing high-risk blind zone polygons In the local coordinate system of the target vehicle Two-dimensional vertex coordinates ,in , The number of two-dimensional vertex coordinates for the high-risk blind zone polygon; The target feature level parameters are used to inform the collaborative vehicle at which depth level of its feature pyramid to perform feature clipping, including which target information the collaborative vehicle needs to extract from the neural network structure and the depth position of each level.

[0066] The encapsulated data packet is broadcast to all cooperating vehicles via the PC5 communication interface of the on-board OBU device.

[0067] S104, Receive the local implicit feature tensor returned by the cooperative vehicle; the local implicit feature tensor is a lightweight implicit feature representation that contains only the deep semantic information of the high-risk blind zone polygon region, as determined by the cooperative vehicle based on the cooperative perception request data packet.

[0068] When the baseband chip is in continuous detection mode, once a nearby cooperating vehicle sends a data request signal, it will immediately execute a direct memory access (DMA) operation to quickly send the payload data with local implicit feature tensors into the video memory, so as to ensure that multi-source data fusion processing can be performed later.

[0069] After receiving the cooperative perception request data packet, the cooperative vehicle returns a local implicit feature tensor according to certain triggering conditions. That is, the cooperative vehicle aligns the cooperative perception request data packet with the spatial coordinate system of the second BEV feature map and performs two-dimensional spatial mask clipping. The part that intersects with the polygonal region of the high-risk blind spot is obtained as the intermediate layer feature information of the network and is used as a simplified representation of the original perception data, i.e., the local implicit feature tensor.

[0070] Specifically, such as Figure 3 As shown, Figure 3 The feature extraction process for the collaborative vehicle is presented. When the collaborative vehicle is in an open road environment and driving normally, it automatically switches to collaborative mode after receiving a collaborative working instruction from other target vehicles. The collaborative vehicle performs the following steps when generating local implicit feature tensors:

[0071] S301: The cooperating vehicle continuously monitors the V2X communication network, receives cooperative perception request data packets broadcast or multicast by the target vehicle, and performs protocol parsing on the cooperative perception request data packets to extract the target vehicle's global pose matrix, the target vehicle's absolute timestamp, and the vertex coordinate set of the high-risk blind zone polygon.

[0072] The underlying communication hardware of the collaborative vehicle continuously monitors the V2X communication frequency band. The V2X protocol stack relies on the onboard communication equipment to continuously detect broadcast signals in the environment. Once it receives a collaborative perception request data packet broadcast by the target vehicle, the protocol parsing engine immediately unloads the load and extracts the high-risk blind spot polygon. The vertex set and the global pose matrix of the target vehicle The absolute timestamp of the target vehicle synchronized by the system And the coordinate set of the polygon vertices in high-risk blind spots. Considering that multiple vehicles may issue requests simultaneously in congested traffic conditions, the collaborative vehicle has a request priority scheduling queue. The potential threat level is assessed based on the distance and relative speed between the request source and the receiver. High-risk tasks are placed in the highest priority position according to the priority principle, so as to achieve the purpose of effective and rapid arrangement of work process for resource allocation.

[0073] Optionally, the remaining time to collision (TTC) can be calculated based on the speed and distance of the requesting vehicle, and the emergency request with the shortest TTC can be responded to first.

[0074] S302: Acquire the sensor environment perception data stream of the cooperative vehicle, and extract multi-scale features of the sensor environment perception data stream through a pre-deployed BEV backbone network to generate a second BEV feature map in the local coordinate system of the cooperative vehicle.

[0075] In parallel with receiving requests, the autonomous driving domain controller of the cooperative vehicle uses its configured sensor suite to acquire environmental perception data streams and generates a second BEV feature map in the local physical coordinate system of the cooperative vehicle in real time through its internally deployed BEV backbone feature extraction network. ,in, represents the channel depth dimension of the feature map. The spatial height of the feature map in a two-dimensional plane. Let be the width of the feature map in the two-dimensional plane. The dimension of the tensor data structure of this feature map in memory is defined as... ,in The size of the feature channel. The feature map represents the vertical spatial pixel height. The width of the feature map in pixels horizontally.

[0076] S303, based on the global pose matrix of the cooperating vehicle, solve for the relative pose affine transformation matrix between the target vehicle and the cooperating vehicle, and use the relative pose affine transformation matrix to transform the high-risk blind zone polygon. The vertex coordinates are mapped to the pixel coordinate system of the second BEV feature map to obtain the projected polygon.

[0077] S304, based on the projected polygon, the second BEV feature map is cropped, retaining the effective feature region inside the projected polygon.

[0078] Among them, based on projected polygons In the second BEV feature map Generate a binary two-dimensional spatial mask matrix in the spatial dimension, and use this mask matrix to... Perform a bitwise clipping operation to remove redundant environmental features and retain only the valid feature regions inside the polygon.

[0079] S305 uses a multi-channel fusion approach to selectively converge effective feature regions and obtains local implicit feature vectors through quantization and dimensionality reduction.

[0080] A multi-channel fusion approach is used to selectively aggregate feature regions, and then quantization and dimensionality reduction are employed to obtain simplified and representative local implicit feature vectors. Vehicle-to-everything (V2X) communication technology is then used to transmit the relevant data to the vehicle terminal.

[0081] To fundamentally solve the severe ghosting problem caused by misalignment of the absolute coordinate systems of the two vehicles in traditional target-level fusion, the cooperating vehicle does not identify obstacles in its local coordinate system and then transform its coordinates. Instead, it directly maps the requested blind zone by calculating the relative affine transformation matrix between the two vehicles. In one embodiment, based on the global pose matrix of the cooperating vehicle, the relative pose affine transformation matrix between the target vehicle and the cooperating vehicle is solved. The vertex coordinates of the high-risk blind zone polygon are then mapped to the pixel coordinate system of the second BEV feature map using the relative pose affine transformation matrix to obtain the projected polygon. This includes the following steps:

[0082] S401, the product of the inverse of the global pose matrix of the cooperative vehicle and the global pose matrix of the target vehicle is determined as the relative pose affine transformation matrix.

[0083] Among them, the phase pose affine transformation matrix The solution formula is: in, For the global pose matrix of the cooperative vehicle The inverse matrix is ​​used to transform world coordinates to the cooperative vehicle's local coordinate system; This is the global pose matrix for transforming the target vehicle's local coordinate system to the world coordinate system. Phase pose affine transformation matrix. It accurately integrates the relative rotation and translation relationship between the two vehicles in three-dimensional space.

[0084] S402, the fixed-scale scaling and translation intrinsic parameter matrix between the cooperative vehicle local physical coordinate system and the feature map pixel coordinate system, the relative pose affine transformation matrix, and the vertex coordinates of the high-risk blind zone polygon are determined as the product of these three factors to map the vertex coordinates of the high-risk blind zone polygon to the pixel coordinates of the second BEV feature map; based on the mapped pixel coordinates, a projection polygon is constructed.

[0085] Pixel coordinates mapped to the second BEV feature map The calculation formula is: in, Homogeneous coordinate column vector of the vertices of the blind spot polygon sent to the target vehicle , This is a fixed-scale scaling and translation intrinsic parameter matrix for transforming the local physical coordinate system of the collaborative vehicle into the pixel coordinate system of the feature map. In other words, it is a fixed-scale scaling and translation intrinsic parameter matrix for transforming the continuous coordinates of the physical space into the discrete pixel grid coordinates of the network feature map inside the collaborative vehicle.

[0086] All of By connecting them in sequence, the precisely mapped projected polygon can be drawn in the second BEV feature map. .

[0087] In one embodiment, the second BEV feature map is cropped based on the projected polygon, retaining the effective feature region inside the projected polygon, including the following steps: S501, Generate a two-dimensional spatial mask matrix with the same size as the spatial resolution of the second BEV feature map. .

[0088] S502, iterate through all pixel coordinates in the two-dimensional spatial mask matrix. If the point is located within the projected polygon If the pixel is inside or on the boundary, then the element value corresponding to that pixel coordinate point is assigned the value 1, that is... Otherwise, assign a value of 0, that is... ; S503, perform a Hadamard product operation on the two-dimensional spatial mask matrix and the second BEV feature map to obtain the clipped feature tensor. That is, the effective feature region.

[0089] in, This represents a pixel-wise element-wise multiplication operation in the spatial dimension, making All tensor elements in the outer region of the polygon are set to zero, thereby eliminating the unnecessary occupation of communication bandwidth.

[0090] In this embodiment, the cooperative vehicle is based on a projected polygon. The geometric boundary is used to generate a two-dimensional binary space mask matrix with the same spatial resolution as its second BEV feature map using a polygon filling algorithm. The generation rule for this matrix is ​​as follows: if any feature map pixel coordinate is located inside the polygon or on its boundary line, the mask value is assigned a value of 1; if it is located in an irrelevant external environment region, the value is assigned a value of 0. Subsequently, a Hadamard product operation is performed on the original feature map tensor and this binary mask matrix. This operation instantly forces irrelevant background features outside the polygon to zero, which greatly reduces the interference of environmental background noise on information extraction.

[0091] In one embodiment, a multi-channel fusion approach is used to selectively converge effective feature regions, and local implicit feature vectors are obtained through quantization and dimensionality reduction, including the following steps: S601, Extract the cropped feature tensor Subtensors within the minimum bounding box region composed of non-zero elements .

[0092] S602, using pre-trained 1x1 convolutional kernels paired with tensors Dimensionality reduction is performed along the channel dimension to obtain FP32 floating-point feature data; where This represents the channel compression ratio.

[0093] Optionally, subtensors Input is fed into the network dimensionality reduction module. A single-point convolutional kernel is used ( Convolution performs cross-channel dimension aggregation and dimensionality reduction operations, reducing the number of tensor channels from the original... Extremely compressed to ,in The preset channel compression ratio parameter for the system.

[0094] S603 maps FP32 floating-point feature data to INT8 integer representation in [-128, 127] through quantization mapping and encoding conversion, thereby achieving data compression and obtaining a target tensor sequence with local implicit characteristics, i.e., local implicit feature vector.

[0095] The post-training quantization technique is used to map and transform the single-precision floating-point (FP32) feature data output by the network dimensionality reduction module to a discrete low-bit-width integer (INT8) data space, generating an extremely lightweight local implicit feature tensor payload, i.e., a local implicit feature vector, which is then rapidly transmitted back to the target vehicle that issued the request through the underlying communication network.

[0096] In this embodiment, the cropped effective feature regions (non-zero feature regions) are first subjected to cross-channel dimensionality reduction, and then further optimized using a quantization strategy. The dataset compressed in this stage, after fusing the target vehicle's attitude parameters and the current timestamp, becomes a local implicit feature tensor payload module. Relying on the low-latency, high-priority communication methods in the V2X protocol stack, this payload module directly transmits information to the target vehicle using unicast or multicast, bypassing the scheduling limitations caused by the business layer message queue, thereby achieving efficient information exchange.

[0097] S105, the local implicit feature tensor is fused with the BEV feature map from the first bird's-eye view to obtain global features. The state of potential obstacles in high-risk blind spots is predicted using global features to determine the control commands for the target vehicle.

[0098] First, a time synchronization compensation algorithm is employed to ensure the spatiotemporal consistency of the obtained local implicit feature tensors, followed by adjustment. The processed data is then correctly placed in its corresponding region. After multi-source information fusion and analysis, the predicted state of potential obstacles in high-risk blind spots and corresponding control commands are obtained. By using time synchronization to uniformly process spatiotemporal data, the most suitable operating mode is obtained to improve decision-making accuracy and ensure driving safety.

[0099] The target vehicle's communication module is in a continuous listening state. Once it receives the local implicit feature tensor payload triggered and transmitted back by surrounding cooperating vehicles based on request data packets, it immediately unpacks it and sends it to the fusion module. Due to the physical delay of the wireless communication medium and the computing power consumption of the edge computing nodes of the two vehicles, there is an inevitable time loss between the feature generation time and the current fusion time. The target vehicle must perform feature-level time synchronization compensation to prevent ghost braking caused by time delay. That is, before fusing the local implicit feature tensor with the first bird's-eye view BEV feature map, a time synchronization correction algorithm is used to perform time-series matching on the obtained local implicit feature tensor. The specific operation is as follows:

[0100] S701, Extract the generation timestamp carried by the local implicit feature tensor. .

[0101] The feature generation timestamp parsed from the received local implicit feature tensor data packet is: .

[0102] S702, the target vehicle's current system timestamp With the generation timestamp Time difference between This refers to communication delay.

[0103] S703, if time difference Greater than the preset maximum communication delay threshold If the local implicit feature tensor fails, it is determined to be invalid, discarded, and the target vehicle's speed within the blind zone is reduced. Specifically, the target vehicle is controlled to travel at a preset speed within the blind zone, which is less than a first speed threshold. The preset speed is an absolutely safe speed less than the first speed threshold.

[0104] If Exceeding the system's maximum communication latency threshold If the local implicit feature tensor is discarded, the system degenerates to a conservative single-vehicle driving strategy, which forces the underlying control system to adopt a deceleration and backup braking strategy to ensure absolute safety.

[0105] S704, if time difference Less than or equal to the maximum communication delay threshold Then, based on the obstacle-level prior velocity vector sent by the cooperative vehicle... Spatial translation compensation based on the optical flow field assumption is performed on the feature pixels in the local implicit feature tensor; the spatially translated local implicit feature tensor is used for feature fusion.

[0106] If time difference Less than or equal to the maximum communication delay threshold If so, the data is considered to be within its valid lifecycle.

[0107] Specifically, based on the obstacle-level prior velocity vector sent by the cooperative vehicle, spatial translation compensation based on the optical flow field assumption is performed on the feature pixels in the local implicit feature tensor, including: according to the obstacle-level prior velocity vector. and time difference Determine the translational displacement vector By translating the displacement vector By spatially translating each feature pixel in the local implicit feature tensor, a spatially translated compensated local implicit feature tensor is obtained, thereby completely eliminating the feature space position deviation caused by communication and computation delays. Optionally, this can be achieved by translating the local implicit feature tensor in two-dimensional space according to the translation vector. Resampling with reverse or forward translation is performed to eliminate the physical location offset in the feature space caused by communication delay.

[0108] This embodiment uses a time synchronization correction algorithm to perform time-series matching on the obtained local implicit feature tensors, thereby aligning the feature-level time delay errors.

[0109] Before physically stitching the time-compensated local implicit feature tensor to the corresponding missing region of the first BEV feature map, considering the differences in sensor aging, distance, and communication quality among different cooperative vehicles, a dynamic weight is introduced. The local implicit feature tensor is then weighted by this dynamic weight and fused with the first BEV feature map. In one embodiment, fusing the local implicit feature tensor with the first bird's-eye view BEV feature map to obtain global features includes the following steps:

[0110] S801, based on the time difference between the current system timestamp and the generated timestamp of the target vehicle. Determine the time decay term ,in, The time decay sensitivity coefficient is used to apply a non-linear penalty to data with long delays, and determines the curvature of the confidence level as the delay increases.

[0111] S802, based on the physical straight-line distance from the perception sensor of the cooperative vehicle to the geometric center of the high-risk blind spot polygon. and the maximum effective detection range of collaborative vehicle perception sensors Determine the spatial attenuation term .

[0112] S803, weights the time decay term and the spatial decay term to obtain the dynamic weights corresponding to the local implicit feature tensor.

[0113] Dynamic weights The calculation formula is: in, and These are empirical hyperparameters pre-trained using a large-scale dataset, used to adjust the relative importance of temporal and spatial decay, and satisfying normalization conditions. , ; The base of the natural logarithm is denoted by . This weighting formula ensures that the lower the data transmission latency, the closer the cooperative vehicle is to the blind spot, and the better the observation angle, the stronger the dominance of its features in the fusion network.

[0114] S804 defines the product of dynamic weights and local implicit feature tensors as the weighted local implicit feature tensor.

[0115] The local implicit feature tensor after spatial translation compensation and the dynamic weights are determined as the weighted local implicit feature tensor.

[0116] S805 aligns the weighted local implicit feature tensor in the spatial dimension and fills it into the coordinate pixel missing region of the corresponding high-risk blind zone polygon in the first BEV feature map to obtain the global feature.

[0117] The global features are the complete BEV feature map. The global features are input to the joint decoder (such as a Transformer-based detection head) through a cross-attention mechanism for decoding and inference. The output is a high-precision blind spot hidden obstacle prediction state. Based on this, active safety commands such as AEB (Autonomous Emergency Braking) or steering avoidance are sent to the chassis drive-by-wire module via the vehicle Ethernet.

[0118] The blind spot hidden obstacle prediction status includes the 3D bounding box coordinates, category probability, and future trajectory prediction polynomial of the hidden obstacles within the blind spot. If the predicted trajectory interferes with the expected trajectory of the target vehicle and triggers a collision warning, the domain controller immediately sends an AEB braking signal or a steer-by-wire avoidance signal to the chassis drive-by-wire module via the vehicle's local area network.

[0119] In one embodiment, there are multiple cooperating vehicles corresponding to the target vehicle. After calculating the weighted local implicit feature tensor of each cooperating vehicle, the weighted local implicit feature tensor is aligned in the spatial dimension and filled into the coordinate pixel missing area of ​​the corresponding high-risk blind zone polygon in the first BEV feature map. If there is an overlapping area among the multiple weighted local implicit feature tensors, for any overlapping point, the largest weighted local implicit feature tensor value corresponding to the overlapping point is used as the target feature value to fill into the corresponding coordinate of the coordinate pixel missing area of ​​the corresponding high-risk blind zone polygon in the first BEV feature map.

[0120] In one embodiment, such as Figure 4 As shown, Figure 4This invention presents an end-to-end network structure that integrates various coordinate types and includes spatial mask clipping and local feature fusion. This framework investigates how high-dimensional tensor data is transferred between different layers within the neural network during the migration from target detection to a vehicle platform, and explores the accompanying mathematical rules and physical meanings.

[0121] In the collaborative vehicle-side network, heterogeneous sensor data is first fed into the BEV backbone network. The backbone network utilizes a Transformer module with spatial self-attention or a deep 3D sparse convolution module to fuse and flatten the multimodal data in 3D voxel space, outputting a dense second BEV feature map. .

[0122] Next, we move on to the core module, "2D Spatial Mask Matrix Generation and Cropping." The collaborative vehicle utilizes the Scan-line Polygon Fill Algorithm or RayCasting Method from computer graphics, based on the mapped projected polygons... The geometric boundaries are dynamically instantiated in system memory into a two-dimensional binary mask matrix that has the same spatial dimensional resolution as the second BEV feature map. .

[0123] The pixel-level assignment logic of this mask matrix follows a strict inside / outside discrimination criterion: traversal All coordinate points If it is determined that the current pixel coordinates are inside the projected polygon or exactly on its geometric boundary line, then the value of the mask matrix position is assigned as... If the location is determined to be outside the polygon, then the value will be forcibly set to [value to be specified]. .

[0124] After the mask matrix is ​​generated, the second BEV feature map output by the backbone network is processed. With binary mask matrix Perform the Hadamard Product operation. The mathematical essence of the Hadamard Product is a pixel-by-pixel element-wise multiplication operation performed on a two-dimensional plane. Since the mask matrix has all zero values ​​outside the polygon, this product operation instantly forces all network activation values ​​corresponding to the vast irrelevant environmental features outside the blind zone of the second BEV feature map to zero (zeroing out), perfectly preserving only the effective high-dimensional feature regions inside the polygon. Extract the smallest sub-tens set containing non-zero elements within the bounding box. .

[0125] Subsequently, the tensor flow is transferred to the channel dimensionality reduction and INT8 quantization compression module. To further overcome the limitations of communication bandwidth, an adaptive weighted... Single-point convolution kernel paired tensors Perform cross-channel dimension aggregation and dimensionality reduction operations. This operation does not change the spatial resolution of the features, but reduces their number of channels from the initial value. Dimensions compressed rapidly to Dimension, among which This is a channel compression ratio parameter that the system dynamically adjusts based on the current channel congestion state (CBR, Channel Busy Ratio).

[0126] After dimensionality reduction, the data type in the feature tensor remains a single-precision floating-point number (FP32) occupying 4 bytes of GPU memory. To reduce the byte length, post-training quantization (PTQ) is performed. Utilizing the minimum and maximum values ​​of the statistically collected feature distribution, a symmetric or asymmetric linear quantization algorithm is used to map the floating-point features to a discrete 8-bit low-precision integer (INT8) numerical space. The quantization formula is expressed as follows: ,in For input floating-point values, This is the quantization scaling factor. This represents the zero-point offset. After undergoing three rigorous compression processes—mask space truncation, channel-level single-point convolution dimensionality reduction, and data type quantization—an extremely lightweight local implicit feature tensor is generated and transmitted via the V2X wireless communication channel.

[0127] In vehicular local area networks (VLANs), forward occlusion can lead to insufficient local information in the angle estimated by the higher-level BEVI (Before Image Variable Interference). To overcome this problem, the decompressed local feature data stream, obtained through a spatiotemporal alignment and feature fusion module (Concat), can fill in the missing information and improve edge continuity. The corrected feature map is then fed into the joint decoding unit. Employing a multilayer perceptron (MLP) or cross-attention layer mechanism, this module effectively suppresses abrupt changes in pixel values ​​at stitching boundaries and enables better fusion of deep features. After further refinement by the output head, the system can obtain a high-precision collaborative prediction target of 3D potential obstacles.

[0128] When applying the vehicle-to-vehicle cooperative perception method based on blind spot features and on-demand requests provided in this manual, it is not necessary to... Figure 1The steps shown are executed in sequence. The specific execution order of each step can be determined as needed, and this manual does not impose any restrictions on it.

[0129] The above describes a vehicle-to-vehicle cooperative perception method based on blind spot feature on-demand request, provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding vehicle-to-vehicle cooperative perception device based on blind spot feature on-demand request, which includes: The generation module is used to generate a first bird's-eye view BEV feature map based on the environmental data collected from the target vehicle, and to identify the location of physical obstructions of the target vehicle based on the first BEV feature map. The determination module is used to predict the expected driving trajectory of the target vehicle in a future time window based on the current kinematic state of the target vehicle, and to determine the high-risk blind zone polygon based on the expected driving trajectory and the position of physical obstructions; the high-risk blind zone polygon represents the area where the expected driving trajectory of the target vehicle overlaps with the blind zone behind the physical obstruction. The sending module is used to extract the vertex coordinates of the high-risk blind zone polygon, the global positioning information of the target vehicle, and the timestamp, form a collaborative perception request data packet, and send it to the collaborative vehicle within the communication radius of the target vehicle. The receiving module is used to receive the local implicit feature tensor returned by the cooperative vehicle; the local implicit feature tensor is a lightweight implicit feature representation that contains only the deep semantic information of the high-risk blind zone polygon region, as determined by the cooperative vehicle based on the cooperative perception request data packet. The prediction module is used to fuse the local implicit feature tensor with the BEV feature map from the first bird's-eye view to obtain global features. The global features are then used to predict the state of potential obstacles in high-risk blind spots and determine the control commands for the target vehicle.

[0130] Specific limitations regarding the vehicle-to-vehicle cooperative perception device based on blind spot feature-based on-demand requests can be found in the limitations of the vehicle-to-vehicle cooperative perception method based on blind spot feature-based on-demand requests mentioned above, and will not be repeated here. Each module in the aforementioned vehicle-to-vehicle cooperative perception device based on blind spot feature-based on-demand requests can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0131] This specification also provides a vehicle-to-vehicle cooperative perception system based on blind spot characteristics and on-demand requests, which is a device with blind spot characteristics consisting of multiple target vehicle devices and multiple cooperative vehicle devices arranged in a cooperative perception network.

[0132] The vehicle-mounted device is equipped with a computing component, and its main functional modules consist of the following parts.

[0133] This module mainly integrates various heterogeneous data in the vehicle environment, laying the groundwork for subsequent data analysis and processing.

[0134] The first-level BEV feature extraction module mainly analyzes the raw sensor data and constructs a first-view BEV based on the vehicle as a reference frame. Then, based on this, it identifies and locates targets that may be obscured.

[0135] By using ray tracing technology and trajectory intersection analysis, a 3D blind spot evaluation algorithm can be used to accurately build geometric models of dangerous and high-risk blind spots in complex scenes.

[0136] This module is mainly responsible for creating vehicle attitude information with polygon coordinates of high-risk blind spots and broadcasting it in a directional manner using the on-board unit (OBU).

[0137] The main function of this module is to integrate the local implicit feature tensor data of the cooperative vehicles. It uses time alignment and spatial geometric transformation algorithms to effectively fuse the feature information, obtaining a dynamic state description of the target in the blind spot. Finally, based on the results, it generates the corresponding action command sequence for the chassis control system.

[0138] This system embeds a computing platform into the collaborative vehicle equipment, and its core components are as follows.

[0139] The request listening and parsing module mainly captures the collaborative perception request messages sent by the vehicle broadcasting system and then separates the important information elements from them.

[0140] This module primarily collects real-time data on the environment surrounding the cooperative vehicle to obtain a second-view BEV feature map as the basis for analysis.

[0141] This module uses a relative pose transformation matrix to calibrate and reconstruct the position of high-risk blind zone polygons in the feature space of the target's second bird's-eye view.

[0142] This module mainly creates a two-dimensional spatial mask matrix, using the overlapping parts to select important intermediate layer feature information of the network.

[0143] The main task of this module is to perform dimensionality reduction and quantization on the selected good features, and return the generated local implicit feature tensor to the smart terminal device that requested the operation.

[0144] This system adopts a distributed architecture design, primarily using multiple in-vehicle terminal devices with communication capabilities and collaborative target vehicle devices. The intelligent connected vehicle hardware meets ASILD-level functional safety standards, therefore requiring the installation of an autonomous driving domain controller, high-performance solid-state storage, perception units composed of various sensors, and in-vehicle communication terminals using the C-V2X protocol. In the software layer, each in-vehicle unit can monitor the current driving environment in real time based on built-in algorithms or libraries, and autonomously determine its role according to specific traffic scenario requirements. When visibility is obstructed or path conflicts occur, it immediately switches to the role of "request initiator"; conversely, when visibility is good and an external request signal is received, it becomes an "information provider," thereby promoting closer cooperative perception among multiple vehicles and significantly improving situational awareness and decision-making in complex road conditions.

[0145] This invention designs an electronic device specifically for intelligent connected vehicles, the main purpose of which is to centralize the main functions of the Autonomous Driving Domain Controller (ADCU) to achieve the aforementioned... Figure 1 This invention provides a vehicle-to-vehicle cooperative perception method based on blind spot feature-based on-demand requests. The device employs a modular design, primarily consisting of a main control unit, non-volatile random access memory (NVM RAM), and read-only memory (ROM), using an internal high-speed communication bus for data exchange and transmission. The NVM RAM stores various tensor data and map feature information generated during runtime, while the ROM contains various algorithm frameworks for use by other programs. The main control unit interprets sensor data transmitted from the vehicle network using its built-in program instruction set, then determines the current task type based on the environmental conditions, such as whether visibility is obstructed or obstacles are encountered at intersections. When visibility is obstructed, the main control unit initiates the corresponding subroutine of the target vehicle's perception module, executing emergency avoidance behaviors according to pre-set rules. If good visibility is available and a high-priority blind spot distress signal is received, the cooperative interaction logic branch is invoked via interrupt, retrieving external communication protocols and corresponding algorithm components to achieve cross-node resource sharing and collaborative navigation.

[0146] This specification also provides a computer-readable storage medium, the main function of which is to encrypt and store a sequence of instruction codes. This storage medium stores a computer program, which can be used to execute the above-mentioned... Figure 1 The provided method is a vehicle-to-vehicle cooperative perception method based on blind spot features and on-demand requests.

[0147] The code stored in the computer-readable storage medium is specifically designed for the ECUs and central computing domain controllers of intelligent connected vehicles. The read program is written into the running memory, and driven by the computing core, it supports vehicle-to-vehicle cooperative perception schemes guided by blind spot features in various scenarios of the vehicle control system. This method can significantly reduce the high bandwidth resources required by the in-vehicle communication system, greatly improve the accuracy of spatial positioning and the speed of identification of attacked targets, and provide important technical support for improving the safety of autonomous driving. This technology can adapt well to dynamic target detection in complex environments, has wide applicability, and can also provide strong support for the improvement of intelligent decision-making systems in various intelligent transportation applications.

[0148] This instruction manual also provides Figure 5 The schematic diagram of the computer device shown is as follows: Figure 5 As shown, at the hardware level, this computer device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it to achieve the above. Figure 1 The provided method is a vehicle-to-vehicle cooperative perception method based on blind spot features and on-demand requests.

[0149] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0150] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A vehicle-to-vehicle cooperative perception method based on blind spot features and on-demand requests, characterized in that, include: Based on the environmental data collected from the target vehicle, a first bird's-eye view BEV feature map is generated, and the location of physical obstructions of the target vehicle is identified based on the first BEV feature map. Based on the current kinematic state of the target vehicle, predict the expected driving trajectory of the target vehicle in the future time window, and determine the high-risk blind zone polygon based on the expected driving trajectory and the position of physical obstructions; the high-risk blind zone polygon represents the area where the expected driving trajectory of the target vehicle overlaps with the blind zone behind the physical obstruction. Extract the vertex coordinates of the high-risk blind zone polygon, the global positioning information of the target vehicle, and the timestamp to form a collaborative perception request data packet, and send it to the collaborative vehicle within the communication radius of the target vehicle. Receive the local implicit feature tensor returned by the cooperative vehicle; the local implicit feature tensor is a lightweight implicit feature representation that contains only the deep semantic information of the high-risk blind zone polygon region, as determined by the cooperative vehicle based on the cooperative perception request data packet. The local implicit feature tensor is fused with the BEV feature map from the first bird's-eye view to obtain global features. The state of potential obstacles in high-risk blind spots is predicted using global features to determine the control commands for the target vehicle.

2. The method according to claim 1, characterized in that, Based on the expected driving trajectory and the location of physical obstructions, high-risk blind spot polygons are determined, including: Based on the location of the physical occluder, extract the set of visible edge points of the physical occluder in the first BEV feature map; Using the optical center of the target vehicle's sensing sensor as the origin of the light projection, a cluster of rays is emitted into the surrounding unknown space through the set of visible edge points; Based on the geometric intersection data of the ray cluster and the passable road plane, an initial view frustum occlusion space model is established; the initial view frustum occlusion space model represents the three-dimensional view frustum blind zone behind the physical occlusion object; By processing the expected driving trajectory envelope and the initial visual cone occlusion space model through Boolean intersection operations, the overlapping sub-regions of the two are identified as actual high-risk blind spots; The actual high-risk blind spot is projected vertically downwards onto the horizontal ground plane where the target vehicle is located. The set of boundary vertices of the projected area is extracted to generate a closed high-risk blind spot polygon.

3. The method according to claim 1, characterized in that, The structure of the data fields inside the collaborative awareness request data packet includes: Frame header identifier field; The unique identification code field for the target vehicle; The target vehicle's absolute pose field includes the longitude, latitude, altitude, and three-dimensional heading angle matrix output by the fusion of the global navigation satellite system and the inertial measurement unit, i.e., global positioning information; The time synchronization field records the absolute timestamp of the target vehicle when the high-risk blind spot polygon was generated. The polygon vertex sequence field contains the coordinates of multiple two-dimensional vertices of the high-risk blind spot polygon in the target vehicle's local coordinate system. The target feature hierarchy parameters include parameters such as which target information the collaborative vehicle needs to extract from the neural network structure and the depth position of each layer.

4. The method according to claim 1, characterized in that, Before fusing the local implicit feature tensor with the first bird's-eye view BEV feature map, the method further includes: Extract the generation timestamp carried by the local implicit feature tensor; The time difference between the current system timestamp and the generated timestamp of the target vehicle; If the time difference is greater than the preset maximum communication delay threshold, the local implicit feature tensor is discarded, and the target vehicle is controlled to drive in the blind area at a preset speed, which is less than the first speed threshold. If the time difference is less than or equal to the maximum communication delay threshold, then based on the obstacle-level prior velocity vector sent by the cooperative vehicle, spatial translation compensation based on the optical flow field assumption is performed on the feature pixels in the local implicit feature tensor; the spatially translated local implicit feature tensor is used for feature fusion.

5. The method according to claim 4, characterized in that, Based on the obstacle-level prior velocity vector sent by the cooperative vehicle, spatial translation compensation based on the optical flow field assumption is performed on the feature pixels in the local implicit feature tensor, including: The translational displacement vector is determined based on the obstacle-level prior velocity vector and the time difference. By spatially translating the feature pixels in the local implicit feature tensor using a translation vector, we obtain the spatially translated and compensated local implicit feature tensor.

6. The method according to claim 1, characterized in that, The local implicit feature tensor is fused with the BEV feature map from the first bird's-eye view to obtain global features, including: The time decay term is determined based on the time difference between the current system timestamp and the generated timestamp of the target vehicle; The spatial attenuation term is determined based on the physical straight-line distance between the perception sensor of the cooperative vehicle and the geometric center of the high-risk blind zone polygon and the maximum effective detection range of the perception sensor of the cooperative vehicle; By weighting the time decay term and the spatial decay term, the dynamic weights corresponding to the local implicit feature tensor are obtained; The product of the dynamic weights and the local implicit feature tensor is determined as the weighted local implicit feature tensor; The weighted local implicit feature tensor is aligned in the spatial dimension and filled into the coordinate pixel missing region of the corresponding high-risk blind zone polygon in the first BEV feature map to obtain the global feature.

7. The method according to claim 1, characterized in that, The cooperative vehicle performs the following steps when generating local implicit feature tensors: Receive the cooperative perception request data packet broadcast or multicast from the target vehicle, and perform protocol parsing on the cooperative perception request data packet to extract the target vehicle's global pose matrix, the target vehicle's absolute timestamp, and the vertex coordinate set of the high-risk blind zone polygon. The sensor environment perception data stream of the cooperative vehicle is acquired, and multi-scale features of the sensor environment perception data stream are extracted through a pre-deployed BEV backbone network to generate a second BEV feature map in the local coordinate system of the cooperative vehicle. Based on the global pose matrix of the cooperative vehicle, the relative pose affine transformation matrix between the target vehicle and the cooperative vehicle is solved. The vertex coordinates of the high-risk blind zone polygon are mapped to the pixel coordinate system of the second BEV feature map using the relative pose affine transformation matrix to obtain the projected polygon. Based on the projected polygon, the second BEV feature map is cropped to retain the effective feature region inside the projected polygon. A multi-channel fusion approach is used to selectively converge effective feature regions, and local implicit feature vectors are obtained through quantization and dimensionality reduction.

8. The method according to claim 7, characterized in that, Based on the global pose matrix of the cooperative vehicle, the relative pose affine transformation matrix between the target vehicle and the cooperative vehicle is solved. Using this matrix, the vertex coordinates of the high-risk blind zone polygon are mapped to the pixel coordinate system of the second BEV feature map, resulting in the projected polygon, which includes: The product of the inverse of the global pose matrix of the cooperative vehicle and the global pose matrix of the target vehicle is determined as the relative pose affine transformation matrix. The fixed-scale scaling from the local physical coordinate system of the cooperative vehicle to the pixel coordinate system of the feature map, the product of the translation intrinsic parameter matrix, the relative pose affine transformation matrix, and the vertex coordinates of the high-risk blind zone polygon, are used to map the vertex coordinates of the high-risk blind zone polygon to the pixel coordinates of the second BEV feature map. Construct a projection polygon based on the mapped pixel coordinates.

9. The method according to claim 7, characterized in that, Based on the projected polygon, the second BEV feature map is cropped to retain the effective feature regions within the projected polygon, including: Generate a two-dimensional spatial mask matrix with the same size as the spatial resolution of the second BEV feature map; Iterate through all pixel coordinates in the two-dimensional spatial mask matrix. If the point is located inside or on the boundary of the projected polygon, assign the element value corresponding to the pixel coordinates to 1; otherwise, assign the value to 0. Perform a Hadamard product operation between the two-dimensional spatial mask matrix and the second BEV feature map to obtain the clipped feature tensor, which is the effective feature region.

10. The method according to claim 9, characterized in that, A multi-channel fusion approach is used to selectively converge effective feature regions, and local implicit feature vectors are obtained through quantization and dimensionality reduction, including: Extract the subtensor within the smallest bounding rectangle region composed of the non-zero elements of the clipped feature tensor; Dimensionality reduction of the sub-tensor by channel dimension is performed using a pre-trained 1x1 convolutional kernel to obtain floating-point feature data; By quantization mapping and encoding conversion, floating-point feature data is mapped to INT8 integer representation to obtain local implicit feature vectors.