Internet of vehicles cooperative sensing method based on space confidence map driving

By generating a global spatial confidence set and optimizing communication resource allocation, the problem of balancing communication load and perception accuracy in vehicle-to-everything (V2X) collaborative perception is solved, achieving efficient and accurate perception results and resource utilization, and adapting to the real-time requirements of dynamic scenarios.

CN121509928APending Publication Date: 2026-02-10NINGBO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511612138.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing vehicle-to-everything (V2X) collaborative perception solutions struggle to balance limited communication bandwidth with the need for high-precision perception. Furthermore, intense competition for resources among multiple vehicles leads to wasted communication resources and frequent transmission conflicts, failing to meet real-time and reliability requirements.

Method used

A vehicle-to-everything (V2X) collaborative perception method driven by spatial confidence maps is adopted. A global spatial confidence set is generated by the base station to select high-value perception areas and optimize the allocation of communication resources to maximize the blind spot coverage effect and achieve on-demand transmission and efficient resource utilization.

Benefits of technology

While reducing communication load, it improves target detection recall and average accuracy in key scenarios, coordinates multi-vehicle resource competition, ensures the system's real-time adaptability to dynamic scenarios, and meets the high efficiency, accuracy and reliability requirements of vehicle-to-everything (V2X) collaborative perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509928A_ABST
    Figure CN121509928A_ABST
Patent Text Reader

Abstract

The invention relates to an Internet of Vehicles cooperative sensing method based on space confidence map driving, and the method comprises the steps: screening a high-value sensing region from a source through a local space confidence map, and achieving the reduction of the order of magnitude of a communication load through combining with the transmission of sensing feature data of a to-be-transmitted grid set, and fundamentally relieving the bandwidth constraint. Secondly, global optimization scheduling based on a perception missing score ensures that limited communication resources are accurately delivered to a link which is most effective for compensating a perception blind area, so that the target detection recall rate and the average precision in a key scene are improved on the contrary while great load reduction is carried out, and unification of load reduction and precision increase is realized; finally, according to the scheme, through periodic global optimization, multi-vehicle resource competition is effectively coordinated, transmission conflicts are avoided, the system is endowed with excellent real-time adaptability to dynamic scenes such as vehicle movement and shielding changes, and the technical requirements of cooperative sensing of the Internet of Vehicles for high efficiency, accuracy and reliability are comprehensively met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle-to-everything (V2X) communication technology, and more specifically, to a V2X collaborative perception method driven by spatial confidence maps. Background Technology

[0002] With the development of intelligent connected vehicles, multi-entity collaborative perception through vehicle-to-everything (V2X) has become a key technology to overcome the limitations of single-vehicle perception. However, existing collaborative perception solutions face a core bottleneck when deployed on a large scale: the contradiction between limited communication bandwidth and the need for high-precision perception.

[0003] Existing technologies suffer from the following main problems: First, transmission schemes based on raw data or full features incur huge communication overhead, making it difficult to meet real-time requirements. Second, existing methods generally lack consideration for the "spatial heterogeneity" of perceived information, assuming equal transmission of information across all spatial regions. This results in a significant waste of valuable communication resources in non-critical or background areas that the vehicle itself can clearly perceive. Finally, in multi-vehicle concurrent scenarios, the lack of an effective global coordination mechanism leads to intense competition for communication resources, frequent transmission conflicts, and an inability to guarantee the real-time and reliable interaction of critical perceived information.

[0004] Therefore, there is an urgent need in this field for a collaborative sensing strategy that can accurately guarantee the perception accuracy of key areas while strictly controlling the communication load, in order to solve the problem of the "accuracy-bandwidth" trade-off and the problem of multi-vehicle resource competition. Summary of the Invention

[0005] The technical problem to be solved by this invention is how to solve the key technical problems of balancing communication load and perception accuracy, prominent competition for resources among multiple vehicles, and insufficient real-time performance in dynamic scenarios in multi-subject collaborative perception of vehicle networks.

[0006] This invention provides a vehicle-to-everything (V2X) cooperative perception method driven by spatial confidence maps, comprising: Step 1, the base station signal covers the grid area Each of the vehicles participating in collaborative perception is equipped with a dual-RF link with a PC5 interface and a Uu interface. Step 2: Each vehicle converts the collected environmental perception data into a bird's-eye view feature map, generates a local spatial confidence map based on the bird's-eye view feature map, and sends the local spatial confidence map to the base station through the Uu interface; Step 3: The base station receives the local spatial confidence maps of all vehicles to form a global spatial confidence set. Taking any vehicle as the receiving vehicle, the base station calculates a perception loss score map based on the global spatial confidence set to characterize the receiving vehicle's own perception of the lack of that grid position in the signal coverage grid area and the degree of blind spot filling by other cooperating vehicles. Step 4: With the goal of maximizing the blind spot filling utility between the receiving vehicle and the transmitting vehicle, the base station selects the target transmission pair and the set of grids to be transmitted through an optimization algorithm under the constraint of communication resources. Step 5: The base station allocates communication resources to each target transmission pair and issues scheduling instructions to the vehicles in the target transmission pair through the Uu interface. The sending vehicle extracts the perception feature data corresponding to the transmission grid set from its own bird's-eye view feature map according to the scheduling instructions, and sends it to the receiving vehicle through the PC5 interface on the allocated communication resources. Step 6: The receiving vehicle fuses the received perception feature data with its own bird's-eye view feature map, performs target detection on the fused feature map, and outputs the environmental perception result.

[0007] Compared with existing technologies, this application has the following advantages: First, by screening high-value perception areas from the source through local spatial confidence maps, and then combining this with the transmission of perception feature data of only the set of grids to be transmitted, the communication load is reduced by an order of magnitude, fundamentally alleviating bandwidth constraints. Second, based on global optimization scheduling of perception missing score maps, it ensures that limited communication resources are accurately delivered to the links most effective in filling perception blind spots, so that while significantly reducing the load, the target detection recall rate and average accuracy in key scenarios are improved, achieving a unity of load reduction and accuracy enhancement. Finally, by maximizing the optimization of blind spot filling effectiveness, it effectively coordinates the competition for resources among multiple vehicles, avoids transmission conflicts, and endows the system with excellent real-time adaptability to dynamic scenarios such as vehicle movement and occlusion changes, fully meeting the technical requirements of efficient, accurate, and reliable collaborative perception in the Internet of Vehicles.

[0008] In one possible implementation, the collaborative sensing information includes vehicle position, speed, heading angle, and point cloud-related metadata.

[0009] In one possible implementation, step 2 specifically includes: Step 201: Each vehicle collects environmental perception data using its onboard sensors. ; Step 202: The environmental perception data is converted into a bird's-eye view feature map using an encoder. The expression is: ; In the formula, This represents a bird's-eye view feature map. Indicates encoder, , These represent the height and width of the feature map in the bird's-eye view, respectively. Indicates the number of feature channels; Step 203, employing a detection decoder and a binarization function. Based on bird's-eye view feature map The expression for generating a local spatial confidence map is: ; In the formula, Indicates the detection decoder, This represents a local spatial confidence map, used to quantify the vehicle's position in each grid cell within the base station signal coverage area. Perceived confidence level, ,in, This indicates the grid area covered by the base station signal. ; Step 204: A local spatial confidence map will be generated for each vehicle. The data is sent to the base station via the Uu interface.

[0010] Compared to existing technologies, the above-mentioned technical solution can convert environmental perception data into a structured local spatial confidence map from a bird's-eye view through encoders and detector decoders. Its direct technical principle lies in discretizing the continuous perceived environment into a rasterized probabilistic representation, thereby achieving quantitative modeling of spatial perception uncertainty. The direct technical effect is that each vehicle can efficiently represent its perception quality of the surrounding environment in the form of a low-dimensional local spatial confidence map, providing a computable input for subsequent value-based communication decisions. This setup transforms the complex question of "what features to transmit" into the question of "which spatial locations are worth transmitting," laying the foundation for selecting high-value perceived information and suppressing redundant transmission of background areas from the outset.

[0011] In one possible implementation, step 3 specifically includes: Step 301: The base station receives the local spatial confidence map of all vehicles. This forms a global spatial confidence set. ,in, This represents the collection of all vehicles in the vehicle-to-everything (V2X) network. ; Step 302: Use any vehicle as the receiving vehicle. Other cooperating vehicles are the transmitting vehicles. The base station uses a maximization operation to fuse the local spatial confidence maps of the other cooperating vehicles and introduces a distance weighting function. Calculate the receiving vehicle Weighted fusion confidence graph The expression is: ; In the formula, Indicates vehicle Within the signal coverage grid area, the grid position Aerial view features, Indicates grid position To receive the vehicle physical distance, This represents the distance weighting function. ; Step 303, based on the receiving vehicle Calculation of local spatial confidence map of receiving vehicle Perception loss score map at each grid location within the signal coverage grid area. The expression is: ; In the formula, Indicates receiving vehicle Grid in bird's-eye view The local spatial confidence map, the perception missing score map The receiving vehicles were quantified. At grid position The degree of perception gaps and the extent to which other collaborating vehicles fill in the gaps.

[0012] Compared to existing technologies, the above-mentioned technical solution enables accurate identification of blind spot filling needs through the fusion and missing data analysis of the global confidence map by the base station. Its direct technical principle lies in: first, maximizing operation and distance weighting to filter out the most effective blind spot filling sources for the receiving vehicle from a global perspective; then, quantifying the intersection of the receiving vehicle's own perception gaps and the blind spot filling capabilities of collaborating vehicles through a perception missing data score map. The direct technical effect is the generation of a demand map guiding "on-demand transmission," transforming qualitative collaborative needs into spatially fine-grained, sortable numerical scores. This setup ensures that the base station's scheduling decisions are no longer blind or based on simple heuristic rules, but rather a data-driven optimization process aimed at maximizing blind spot filling benefits, providing a core basis for achieving efficient and accurate allocation of communication resources.

[0013] In one possible implementation, step 4 aims to maximize the blind spot compensation utility between the receiving vehicle and the transmitting vehicle. Under communication resource constraints, the base station uses an optimization algorithm to select the target transmission pair and the set of grid cells to be transmitted, expressed as follows: ; Communication resource constraints: ; ; ; ; In the formula, Indicates the vehicle being sent With receiving vehicle Constructed transmission edge The effect of filling the blind spot Filling the gaps effect The calculation formula is: ; in, Indicates grid position The weights; Indicates based on the receiving vehicle Perception deficit score map and sending vehicles Local spatial confidence map Filter out the set of grid cells to be transmitted , , The preset threshold for perceptual deficit score; Describe the binary decision variables of the transmission edge. ; Represents the set of transmission edges. ; Indicates the number of sub-channels in the communication resources. This indicates the maximum latency limit for transmission on a V2V link; The delay in transmitting sensing feature data at the transmission edge is represented by the following formula: ; in, Indicates the vehicle being sent to receiving vehicle Total number of bits of transmitted sensing feature data , , This represents the number of bits in the feature map of each dimension of the bird's-eye view. Represents a grid Number of bits for coordinate index, , Indicates fixed metadata overhead; Indicates the vehicle being sent To and receiving vehicles The signal-to-noise ratio between the two is calculated using Shannon's formula for the sending vehicles. To and receiving vehicles Link transmission rate , In the formula, Indicates the bandwidth of the link. Indicates the vehicle being sent To receive the vehicle The signal-to-noise ratio of the link between them.

[0014] In one possible implementation, step 6 specifically includes: Step 601, the receiving vehicle A feature fusion network is used to receive the incoming vehicle data. Perception feature data and receiving vehicle The bird's-eye view feature maps are adaptively fused to obtain enhanced fused features, expressed as follows: ; In the formula, Indicating feature fusion networks, Indicates the vehicle being sent Perceptual feature data; Indicates receiving vehicle Corresponding bird's-eye view feature map, Indicates enhanced fusion features; Step 602: Use a detection decoder to process the fused feature map. Perform object detection and output the environment awareness result, expressed as: ; In the formula, Indicates the detection decoder, This indicates the output of the environmental perception result.

[0015] In one possible implementation, the environment perception results include the target category and the size of the three-dimensional bounding box. Center position coordinates and orientation angle A sequence of rotated frames. Attached Figure Description

[0016] Figure 1 This is a diagram of a vehicle-to-everything (V2X) scenario at an intersection, as described in this application. Figure 2 This is a curve comparing the average perceptual accuracy of the experimental rotated box sequence and the real labeled box when the intersection-union ratio is 0.5. Figure 3 This is a curve comparing the average perception accuracy of the experimental rotated box sequence and the actual labeled box when the intersection-union ratio is 0.7. Detailed Implementation

[0017] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0018] In the description of the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the embodiments of this application based on the specific circumstances.

[0019] In the embodiments of this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0020] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0021] See Figures 1-3 As shown in the figure, this application discloses a vehicle-to-everything (V2X) cooperative perception method driven by spatial confidence maps, including: Step 1, the base station signal covers the grid area The system comprises 10 vehicles participating in collaborative perception, each equipped with a dual-radio link consisting of a PC5 interface and a Uu interface. This embodiment is applied to a typical crossroads vehicle-to-everything (V2X) scenario, where one base station (eNB / gNB) and its associated edge computing server are deployed. The base station's signal covers the entire crossroads area. Intelligent connected vehicles (ICVs) equipped with LiDAR and cameras drive on roads within the city and can perform collaborative perception; defining express The V2X in this embodiment adopts 5G NR V2X Mode 1, with centralized scheduling performed by the base station, which uniformly allocates radio resources for V2I and V2V links. The PC interface and Uu interface operate in Out-of-band mode, occupying different frequency bands, which allows vehicle terminals with dual radio frequency links to complete V2I and V2V transmission and reception operations in parallel within the same time slot. Based on the above V2X communication mechanism, vehicles can not only exchange information but also conduct necessary cooperative sensing, thereby improving the sensing depth and accuracy of individual vehicles; and a cooperative sensing period is set. .

[0022] Step 2: Each vehicle converts the collected environmental perception data into a bird's-eye view feature map, generates a local spatial confidence map based on the bird's-eye view feature map, and sends the local spatial confidence map to the base station via the Uu interface; specifically including: Step 201: Each vehicle collects environmental perception data using its onboard sensors. The environmental perception data This includes metadata related to vehicle position, speed, heading angle, and point cloud data. Step 202: The environmental perception data is converted into a bird's-eye view feature map using an encoder. To avoid the complexity of coordinate transformation caused by differences in the perspectives of multiple vehicles, this embodiment uses the bird's-eye view (BEV) as a unified feature representation, expressed as: ; In the formula, This represents a bird's-eye view feature map. Indicates encoder, , These represent the height and width of the feature map in the bird's-eye view, respectively. Indicates the number of feature channels; Step 203, employing a detection decoder and a binarization function. Based on bird's-eye view feature map The expression for generating a local spatial confidence map is: ; In the formula, Indicates the detection decoder, This represents a local spatial confidence map, used to quantify the vehicle's position in each grid cell within the base station signal coverage area. Perceived confidence level, ,in, This indicates the grid area covered by the base station signal. ; Step 204: A local spatial confidence map will be generated for each vehicle. Sending data to the base station via the Uu interface, only transmitting non-zero grid indexes and confidence levels, with a transmission latency of [missing information]. .

[0023] Furthermore, this application also employs a differentiated design through a bird's-eye view conversion strategy based on data from different sensors: For RGB image data, use After the encoder extracts the front and rear view features, it then uses a differentiable warping transformation function based on the camera intrinsic and extrinsic parameter matrices to map the features from the image coordinate system to a unified global bird's-eye view (BEV) coordinate system, thereby achieving spatial alignment of multiple vehicle views. The 3D LiDAR point cloud data is discretized into a bird's-eye view spatial raster structure through voxelization, and then... The encoder extracts feature maps from the bird's-eye view; it further integrates semantic information such as target contours and color attributes from the RGB image to enhance feature representation capabilities.

[0024] Step 3: The base station receives the local spatial confidence maps of all vehicles to form a global spatial confidence set. Taking any vehicle as the receiving vehicle, the base station calculates a perception loss score map based on the global spatial confidence set, which characterizes the receiving vehicle's own perception of the lack of a specific grid location in the signal coverage grid area and the degree of blind spot filling by other cooperating vehicles; specifically including: Step 301: The base station receives the local spatial confidence map of all vehicles. This forms a global spatial confidence set. ,in, This represents the collection of all vehicles in the vehicle-to-everything (V2X) network. ; Step 302: Use any vehicle as the receiving vehicle. Other cooperating vehicles are sending vehicles, and the base station excludes receiving vehicles. The bird's-eye view feature map only considers other cooperating vehicles, i.e., the sending vehicles. To enhance its blind spot coverage capability, the base station employs a maximization operation to fuse the local spatial confidence maps of other cooperating vehicles and introduces a distance weighting function. Calculate the receiving vehicle Weighted fusion confidence graph To encourage blind spot coverage in medium-distance areas, the expression is: ; In the formula, Indicates vehicle Within the signal coverage grid area, the grid position Aerial view features, Indicates grid position To receive the vehicle physical distance, This represents the distance weighting function. ; Step 303, based on the receiving vehicle Calculation of local spatial confidence map of receiving vehicle Perception loss score map at each grid location within the signal coverage grid area. The expression is: ; In the formula, Indicates receiving vehicle Grid in bird's-eye view The local spatial confidence map, the perception missing score map The receiving vehicles were quantified. At grid position The degree of perception gaps and the extent to which other collaborating vehicles fill in the gaps.

[0025] Step 4: With the goal of maximizing the blind spot filling utility between the receiving vehicle and the transmitting vehicle, the base station selects the target transmission pair and the set of grids to be transmitted through an optimization algorithm under the constraint of communication resources. Step 5: The base station allocates communication resources to each target transmission pair and issues scheduling instructions to the vehicles in the target transmission pair through the Uu interface. The sending vehicle extracts the sensing feature data corresponding to the transmission grid set from its own bird's-eye view feature map according to the scheduling instructions, and sends it to the receiving vehicle through the PC5 interface on the allocated communication resources; the expression is: ; Communication resource constraints: ; ; ; ; In the formula, Indicates the vehicle being sent With receiving vehicle Constructed transmission edge The effect of filling the blind spot Filling the gaps effect The calculation formula is: ; in, Indicates grid position The weights; Indicates based on the receiving vehicle Perception deficit score map and sending vehicles Local spatial confidence map Filter out the set of grid cells to be transmitted , , The preset threshold for perceptual deficit score; Describe the binary decision variables of the transmission edge. ; Represents the set of transmission edges. ; Indicates the number of sub-channels in the communication resources. This indicates the maximum latency limit for transmission on a V2V link; The delay in transmitting sensing feature data at the transmission edge is represented by the following formula: ; in, Indicates the vehicle being sent to receiving vehicle Total number of bits of transmitted sensing feature data , , This represents the number of bits in the feature map of each dimension of the bird's-eye view. Represents a grid Number of bits for coordinate index, , Indicates fixed metadata overhead; Indicates the vehicle being sent To and receiving vehicles The signal-to-noise ratio between the two is calculated using Shannon's formula for the sending vehicles. To and receiving vehicles Link transmission rate , In the formula, Indicates the bandwidth of the link. Indicates the vehicle being sent To receive the vehicle The signal-to-noise ratio of the link between them.

[0026] Step 6: The receiving vehicle fuses the received perception feature data with its own bird's-eye view feature map, performs target detection on the fused feature map, and outputs the environmental perception result, specifically including: Step 601, the receiving vehicle A feature fusion network is used to receive the incoming vehicle data. Perception feature data and receiving vehicle The bird's-eye view feature maps are adaptively fused to obtain enhanced fused features, expressed as follows: ; In the formula, Indicating feature fusion networks, Indicates the vehicle being sent Perceptual feature data; Indicates receiving vehicle Corresponding bird's-eye view feature map Indicates enhanced fusion features; Step 602: Use a detection decoder to process the fused feature map. Perform object detection and output the environment awareness result, expressed as: ; In the formula, Indicates the detection decoder, This represents the output environment perception result; the environment perception result includes the target category and the size of the 3D bounding box. Center position coordinates and orientation angle A sequence of rotated frames.

[0027] Compared to schemes that transmit full-scale bird's-eye view feature maps or raw collaborative perception information, this application uses locally generated local spatial confidence scores, which are extremely lightweight data used to quantify perception value. Subsequently, the vehicle only transmits perception feature data corresponding to the set of grid cells to be transmitted as specified by the base station scheduling instructions. This achieves a transmission mechanism that filters data on demand from the source, fundamentally avoiding the waste of communication resources on background areas or low-value information that the vehicle can already clearly perceive. Thus, while ensuring perception performance, the amount of data transmitted is extremely compressed.

[0028] While reducing communication load, this application accurately locates the intersection of the receiving vehicle's perception blind spot and the blind spot filling capabilities of other cooperating vehicles by calculating the perception missing score map; and adopts the priority allocation of communication resources to the link with the highest blind spot filling value, so that the receiving vehicle can fuse the most effective and critical complementary information for filling its own blind spot in subsequent steps after global optimization and screening. This not only maintains the baseline accuracy of single-vehicle perception, but also significantly improves perception performance in key scenarios (such as distant targets and severely occluded targets), so that limited communication resources are transformed into the highest perception accuracy benefit.

[0029] Furthermore, this application quantifies the blind spot compensation relationships between vehicles into weighted candidate transmission pairs, employing global optimization scheduling aimed at maximizing transmission utility. The scheduling considers vehicle half-duplex constraints, sub-channel constraints, and transmission latency constraints, avoiding disordered and conflicting V2V communication between vehicles. It orderly arranges concurrent perception data streams on limited PC5 sub-channel resources, ensuring the reliability of high-value link transmission. Simultaneously, periodic confidence map reporting and scheduling (e.g., every 100ms) enable the system to quickly respond to dynamic scenarios such as vehicle position changes and the appearance of new obstacles, adjusting cooperation strategies in real time to meet the stringent requirements of autonomous driving for low latency and high reliability.

[0030] In the description of the embodiments of this application, it should be noted that the terms "inner" and "outer" and other terms indicating direction or positional relationship are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this application.

[0031] experiment This experiment uses the OPV2V public dataset to construct a multi-intelligent connected vehicle cooperative perception scenario at an intersection. The simulation framework is built on the OpenCOOD open-source platform. The PointPillars model is used to extract BEV features, and the Transformer attention mechanism is used to achieve multi-source feature fusion, ensuring the reproducibility and technical fairness of the experiment.

[0032] The simulation process uses a collaborative sensing cycle. The process involves a fixed 20ms time for basic steps such as vehicle-side local sensing processing, base station data calculation, and Uu interface uplink and downlink transmission (not included in the optimization scope). This invention only optimizes the V2V link sensing feature transmission within the remaining 80ms time window, focusing on resource scheduling, data filtering, and latency control. The specific implementation process is as follows: S1: Load LiDAR point cloud, RGB image and vehicle dynamic metadata (position, speed, heading angle), and initialize vehicle perception-communication module and base station resource scheduling module; S2: The vehicle generates a BEV feature map and a spatial confidence map (SCM), and reports the SCM to the roadside base station via the Uu interface; S3: The base station aggregates multiple vehicle SCMs to generate a global perception blind spot filling demand matrix and quantifies the missing score of each vehicle in the BEV grid; S4: The base station uses an optimization algorithm to select high-value grid sets to be transmitted and target transmission pairs, generates a communication mask and resource scheduling strategy that maximizes utility while satisfying "half-duplex, communication delay and bandwidth constraints", and sends them to the vehicle.

[0033] S5: Vehicles transmit / fuse perceived features according to dispatch instructions, outputting the target category and the size of the 3D bounding box. Center position coordinates and orientation angle A sequence of rotated frames; S6: Frame iteration and termination judgment, and calculate the average sensing accuracy (AP) of the bottle by calculating the intersection-union ratio of the rotated box sequence and the real labeled box, and statistically analyze core indicators such as AP@0.5, AP@0.7, V2V average communication bits and time delay.

[0034] To comprehensively evaluate the performance of the scheme, four types of comparison benchmarks are set: ① Single-Cav (no collaborative blind spot filling); ② Full-Cavs (no communication overhead constraints); ③ Proportional Fairness (PF) scheme based on fixed threshold; ④ The technical solution of this application.

[0035] Get as Figure 2 and Figure 3 The performance comparison curves shown demonstrate that, under a 10MHz bandwidth constraint, AP@0.5 and AP@0.7 significantly outperform Single-Cav and approach the performance limit of Full-Cavs, while significantly reducing communication overhead. Furthermore, by dynamically filtering high-value transmission content, it comprehensively surpasses the fixed-threshold PF scheme, exhibiting outstanding advantages in the "perception accuracy - communication efficiency" dimension. This provides a practical and efficient technical path for the engineering implementation of vehicle-to-everything (V2X) collaborative perception, effectively improving the perception accuracy and decision-making reliability of autonomous driving in complex intersection scenarios.

[0036] In the description of the embodiments of this application, it should be noted that the terms "inner" and "outer" and other terms indicating direction or positional relationship are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this application.

[0037] In the description of this application, the references to terms such as "an embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0038] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A vehicle-to-everything (V2X) cooperative perception method driven by spatial confidence maps, characterized in that, include: Step 1, the base station signal covers the grid area Each of the vehicles participating in collaborative perception is equipped with a dual-RF link with a PC5 interface and a Uu interface. Step 2: Each vehicle converts the collected environmental perception data into a bird's-eye view feature map, generates a local spatial confidence map based on the bird's-eye view feature map, and sends the local spatial confidence map to the base station through the Uu interface; Step 3: The base station receives the local spatial confidence maps of all vehicles to form a global spatial confidence set. Taking any vehicle as the receiving vehicle, the base station calculates a perception loss score map based on the global spatial confidence set to characterize the receiving vehicle's own perception of the lack of that grid position in the signal coverage grid area and the degree of blind spot filling by other cooperating vehicles. Step 4: With the goal of maximizing the blind spot filling utility between the receiving vehicle and the transmitting vehicle, the base station selects the target transmission pair and the set of grids to be transmitted through an optimization algorithm under the constraint of communication resources. Step 5: The base station allocates communication resources to each target transmission pair and issues scheduling instructions to the vehicles in the target transmission pair through the Uu interface. The sending vehicle extracts the perception feature data corresponding to the transmission grid set from its own bird's-eye view feature map according to the scheduling instructions, and sends it to the receiving vehicle through the PC5 interface on the allocated communication resources. Step 6: The receiving vehicle fuses the received perception feature data with its own bird's-eye view feature map, performs target detection on the fused feature map, and outputs the environmental perception result.

2. The vehicle-to-everything (V2X) cooperative perception method based on spatial confidence graph driving according to claim 1, characterized in that, The environmental data includes vehicle position, speed, heading angle, and point cloud-related metadata.

3. The vehicle-to-everything (V2X) cooperative perception method based on spatial confidence graph driving according to claim 2, characterized in that, Step 2 specifically includes: Step 201: Each vehicle collects environmental perception data using its onboard sensors. ; Step 202: The environmental perception data is converted into a bird's-eye view feature map using an encoder. The expression is: ; In the formula, This represents a bird's-eye view feature map. Indicates encoder, , These represent the height and width of the feature map in the bird's-eye view, respectively. Indicates the number of feature channels; Step 203, employing a detection decoder and a binarization function. Based on bird's-eye view feature map The expression for generating a local spatial confidence map is: ; In the formula, Indicates the detection decoder, This represents a local spatial confidence map, used to quantify the vehicle's position in each grid cell within the base station signal coverage area. Perceived confidence level, ,in, Indicates the grid area covered by the base station signal. ; Step 204: A local spatial confidence map will be generated for each vehicle. The data is sent to the base station via the Uu interface.

4. The vehicle-to-everything (V2X) cooperative perception method based on spatial confidence graph driving according to claim 3, characterized in that, Step 3 specifically includes: Step 301: The base station receives the local spatial confidence map of all vehicles. This forms a global spatial confidence set. ,in, This represents the collection of all vehicles in the vehicle-to-everything (V2X) network. ; Step 302: Use any vehicle as the receiving vehicle. Other cooperating vehicles are the transmitting vehicles. The base station uses a maximization operation to fuse the local spatial confidence maps of the other cooperating vehicles and introduces a distance weighting function. Calculate the receiving vehicle Weighted fusion confidence graph The expression is: ; In the formula, Indicates vehicle In the signal coverage grid area, the grid position Aerial view features, Indicates grid position To receive the vehicle physical distance, This represents the distance weighting function. ; Step 303, based on the receiving vehicle Calculate the local spatial confidence map of the receiving vehicle Perceptual loss score map at each grid location within the signal coverage grid area. The expression is: ; In the formula, Indicates receiving vehicle Grid in bird's-eye view The local spatial confidence map, the perception missing score map The receiving vehicles were quantified. At grid position The degree of perception gaps and the extent to which other collaborating vehicles fill in the gaps.

5. The vehicle-to-everything (V2X) cooperative perception method based on spatial confidence graph driving according to claim 4, characterized in that, Step 4 aims to maximize the blind spot compensation utility between the receiving and transmitting vehicles. Under communication resource constraints, the base station uses an optimization algorithm to select the target transmission pair and the set of grid cells to be transmitted, expressed as follows: ; Communication resource constraints: ; ; ; ; In the formula, Indicates the vehicle being sent With receiving vehicle Constructed transmission edge The effect of filling the blind spot Filling the gaps effect The calculation formula is: ; in, Indicates grid position The weights; Indicates based on the receiving vehicle Perception deficit score map and sending vehicles Local spatial confidence map Filter out the set of grid cells to be transmitted , , The preset threshold for perceptual deficit score; Describe the binary decision variables of the transmission edge. ; Represents the set of transmission edges. ; Indicates the number of sub-channels in the communication resources. This indicates the maximum latency limit for transmission on a V2V link; The delay in transmitting sensing feature data at the transmission edge is represented by the following formula: ; in, Indicates the vehicle being sent to receiving vehicle Total number of bits of transmitted sensing feature data , , This represents the number of bits in the feature map of each dimension of the bird's-eye view. Represents grid Number of bits for coordinate index, , Indicates fixed metadata overhead; Indicates the vehicle being sent To and receiving vehicles The signal-to-noise ratio between the two is calculated using Shannon's formula for the sending vehicles. To and receiving vehicles Link transmission rate , In the formula, Indicates the bandwidth of the link. Indicates the vehicle being sent To receive the vehicle The signal-to-noise ratio of the link between vehicles is related to the distance between them.

6. The vehicle-to-everything (V2X) cooperative perception method based on spatial confidence graph driving according to claim 5, characterized in that, Step 6 specifically includes: Step 601, the receiving vehicle A feature fusion network is used to receive the incoming vehicle data. Perception feature data and receiving vehicle The bird's-eye view feature maps are adaptively fused to obtain enhanced fused features, expressed as follows: ; In the formula, Indicating feature fusion networks, Indicates the vehicle being sent Perceptual feature data; Indicates receiving vehicle Corresponding bird's-eye view feature map, Indicates enhanced fusion features; Step 602: Use a detection decoder to process the fused feature map. Perform object detection and output the environment awareness result, expressed as: ; In the formula, Indicates the detection decoder, This indicates the output of the environmental perception result.

7. The vehicle-to-everything (V2X) cooperative perception method based on spatial confidence graph driving according to claim 6, characterized in that, The environmental perception results include the category of the detected target and the size of the target's 3D bounding box. The center coordinates of the target and orientation angle A sequence of rotated frames.