Intelligent collaborative awareness-oriented instance-level semantic simplification interaction method

By adopting instance-level semantic streamlining interaction method and feature fusion method based on timing information in the multi-agent collaborative perception system, the problem of ignoring instance-level semantic information and timing data in the prior art is solved, and more efficient feature interaction and fusion is achieved, and the system's perception ability and robustness are improved.

CN120047914AActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510218615.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-27
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing multi-agent collaborative perception technology ignores the value of instance-level semantic information and the impact of time series data on feature fusion, resulting in poor perception effect in complex scenarios and insufficient system robustness.

Method used

An instance-level semantic streamlined interaction method is adopted, and features are decoupled into private and public features through the decoupling module, only private features are transmitted, and a feature fusion method based on timing information is used in the fusion module to enhance the fusion effect.

Benefits of technology

It reduces communication resource overhead, expands the perception range of the agent, and improves perception capabilities and system robustness in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047914A_ABST
    Figure CN120047914A_ABST
Patent Text Reader

Abstract

The invention provides an instance-level semantic simplification interaction method for intelligent collaborative awareness, and is used for the technical field of automatic driving. According to the method, a sensing model, a decoupling module, a fusion module and a detection module are arranged on each agent of an application scene; sensor data collected by each agent is converted into an object-level mask and a middle-term feature through a sensing model, the middle-term feature is decoupled into a private feature and a public feature by a decoupling module, and the agent only processes the private feature by using the object-level mask and then transmits the processed private feature; and the intelligent agent fuses own historical features, own public features and private features and object-level private features of the cooperative intelligent agent through the fusion module, and the fused features are input into the detection module for target detection. According to the method, instance-level feature interaction is realized, the communication resource overhead is reduced, the perceptual view of the intelligent agent is expanded, the fusion effect is enhanced by using the time sequence data, and efficient feature fusion in a complex scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of multi-agent cooperation technology and autonomous driving technology, and particularly relates to an instance-level semantic reduction interaction method for intelligent cooperative perception. Background Art

[0002] Perception ability is crucial for intelligent vehicles as it directly relates to the vehicle's ability to drive safely and effectively in complex road environments. Intelligent vehicles use a series of sensors, such as cameras, radars, lidar (LiDAR), ultrasonic sensors, etc., to obtain detailed information about the surrounding environment, including but not limited to the positions and speeds of other vehicles, pedestrians, traffic signal states, road signs, and obstacles. However, the real road conditions are very complex, with a large number of occlusions and blind spots. It is difficult for a single agent to perceive these areas through its own data, which requires introducing mutual communication among agents for collaborative perception. Compared with traditional single-vehicle perception, multi-agent collaborative perception can expand the perception range, improve the accuracy of environmental understanding, and enhance the system robustness, providing a good opportunity for the development of intelligent vehicles and autonomous driving.

[0003] Multi-agent collaborative perception has been proven to be an effective way to improve the perception ability of intelligent systems. Given the communication condition limitations in actual application scenarios, the information exchange between multi-agent systems urgently needs a more intelligent and efficient design strategy. Most of the current related technologies focus on compressing global representations to reduce data volume and improve transmission efficiency, but such methods usually ignore the value and role of instance-level semantic information. In addition, efficient feature fusion in complex scenarios is also an issue that cannot be ignored in the collaborative perception system, which is related to the safe operation of autonomous vehicles.

[0004] The Chinese invention patent application with the application number CN202411109995.5 disclosed a collaborative confidence fusion perception method for autonomous driving on November 26, 2024. In this solution, the environmental perception ability of agents is evaluated using the 3D object detection results of different agents after time alignment, and the detection results are corrected. Using the corrected detection results, confidence levels, offline maps, and the coordinates of intelligent perception bodies, a bird's-eye view BEV space confidence map is constructed. During the multi-vehicle collaboration process, this solution does not adopt mid-term feature fusion but instead uses late fusion, which is weaker in performance and more vulnerable to interference. To achieve multi-vehicle collaborative perception based on late fusion, this solution designs a correction mechanism for the detection results of a single agent, performing local correction and global confidence correction on the confidence levels of the detection results. Although this design can reduce the communication volume during the collaboration process, in practical applications, only correcting and merging the detection results are vulnerable to the interference of agent positioning errors, which is not conducive to the robustness of the system. At the same time, due to over-reliance on the detection results of a single agent, this technical solution is not conducive to the detection of difficult targets, has poor performance in complex scenarios, and poses certain safety hazards in autonomous driving.

[0005] The Chinese invention patent application with the application number CN202411155073.8 disclosed a vehicle-road collaborative joint perception method and system under non-ideal conditions on December 6, 2024. A V2X graph is constructed around the selected ego node to share metadata. Visual feature maps are extracted and shared at each node. The visual feature maps are gradually compressed along the channel dimension using a series of 1×1 convolutions. At the ego node, V2X attention fusion and delay compensation are performed on the received feature maps, and the aggregated corrected feature maps are used for detection. The drawback of this technical solution is that it does not consider how to handle semantic information at the instance level, which is very important for understanding complex scenarios. The lack of this part may limit the application scope of the system in complex environments. In addition, although this method takes into account communication constraints and gradually compresses the mid-term features along the channel dimension using a series of 1×1 convolutions, this compression method may lose some important spatial information, affecting the accuracy of the final detection and classification.

[0006] The Chinese invention patent application with the application number CN202111495732.9 disclosed a communication-sensitive multi-agent collaboration method on April 12, 2022. The vehicle agent encodes the locally observed message into a hidden state vector, scores the message before sending it, the edge node sends confirmation information to the agents with the top K message values, the agents send messages after receiving the confirmation information, the edge node receives all messages, extracts the valid information and summarizes and distributes it. The disadvantage of this solution is that it does not consider the confidence differences between the models of the agents, and directly sorts and filters the messages generated by all agents uniformly. In addition, since the confirmation information is only sent to the agents with the top K message values, it may cause the problem of uneven resource allocation, especially when some agents frequently obtain the sending permission while other agents are in a waiting state for a long time.

[0007] The Chinese invention patent application with the application number CN202410771385.5 disclosed a multi-modal large model-based autonomous driving collaborative perception method and device on November 1, 2024. The multi-modal large model processes the point cloud data of the host vehicle to obtain text information, extracts text features from the text information, extracts image features from the image data of the host vehicle, and extracts depth map features from the depth map corresponding to the point cloud data; fuses the depth map features and the image features according to the text features to obtain a first fusion feature; fuses the first fusion feature and the features of the object to be detected sent by the target end to obtain a second fusion feature; the target end includes at least one of the collaborative end of the host vehicle and the road end; performs a multi-end collaborative perception vision task based on the second fusion feature. The disadvantage of this solution is that it does not consider the constraints of communication conditions and in-vehicle unit computing capabilities in actual applications. The vehicle needs to process various modal data such as text, images, and point clouds, and generates the final perception result by collecting the multi-modal fusion features of itself and other vehicles, which poses certain requirements on the computing power and communication capabilities of the vehicle that initiates the collaborative driving request, resulting in a higher cost of the vehicle. Summary of the Invention

[0008] There are such defects in the existing multi-vehicle agent collaborative perception technology: it ignores the value and role of instance-level semantic information, and does not consider the influence of temporal data on feature fusion; by compressing the global representation to reduce the data volume and improve the transmission efficiency, although this compression strategy can reduce the communication volume, it will lose some foreground information, which is not conducive to the final perception effect. In view of these deficiencies, the present invention provides an instance-level semantic reduction and interaction method for intelligent collaborative perception, which realizes instance-level feature interaction, reduces the communication resource overhead, expands the perception field of the agent, and also uses temporal data to enhance the fusion effect to achieve efficient feature fusion in complex scenarios.

[0009] An instance-level semantic reduction interaction method for intelligent collaborative perception provided by the present invention sets a perception model, a decoupling module, a fusion module, and a detection module on each intelligent agent in the application scenario of autonomous driving. The present invention includes the following steps:

[0010] Step 1: Each intelligent agent inputs the data collected by its own sensors into the perception model, and the perception model converts the input data into an object-level mask and intermediate features; the perception model consists of a feature extractor and a detection head. The feature extractor converts the input data into a bird's-eye view (BEV) feature with multi-scale feature representation. The detection head aggregates the BEV features using an attention mechanism, generates candidate regions and confidence scores for the existence of detected objects through a convolutional network, and filters them according to the confidence threshold and IoU threshold to generate an object-level mask. The intermediate features are the bird's-eye view BEV features with multi-scale feature representation. Each intelligent agent processes its private features using the object-level mask and then transmits them, realizing distributed instance-level feature interaction among intelligent agents.

[0011] Step 2: The decoupling module of the intelligent agent decouples its intermediate features into two parts: private features and public features. The intelligent agent only processes its own private features using the object-level mask and then transmits them. Each intelligent agent decouples the private features from the intermediate features and only transmits the private features, realizing a data reduction representation driven by collaborative tasks among intelligent agents. The intelligent agents first transmit information about their current positions and perception ranges to each other to determine cooperation partners and the overall cooperation perception range, and then the cooperating intelligent agents transmit object-level private features to each other.

[0012] Step 3: After receiving the object-level private features transmitted by the cooperating intelligent agent, the intelligent agent inputs its own historical features, its own public and private features, and the object-level private features of the cooperating intelligent agent into the fusion module. The fusion module uses a feature fusion method based on temporal information to obtain the fusion features and inputs them into the detection module. The detection module sends the fusion features into a multi-layer convolutional network for object detection.

[0013] The fusion module uses a feature fusion method based on temporal information to obtain the fusion features, including: first, inputting the object-level private features of the cooperating intelligent agent and the private features of the current intelligent agent into the attention module to obtain the fused private features, then fusing the public features of the current intelligent agent to obtain the preliminary fusion features, splicing the preliminary fusion features with the fused private features of the previous frame of the current intelligent agent in the channel dimension, and then sending them into the self-attention module for processing to obtain the fusion features at the current moment.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0015] (1) The prior art ignores the value and role of instance-level semantic information and does not consider the impact of temporal data on feature fusion. The method of the present invention can maintain the model perception effect and enhance the object perception ability of the intelligent collaborative perception model in complex scenarios by means of collaborative task-driven data reduction representation, distributed instance-level feature interaction, and feature fusion based on temporal information, while reducing communication overhead.

[0016] (2) The method of the present invention compresses features in a decoupled manner to achieve collaborative task-driven data reduction representation and only transmits the private features of the current agent, thereby reducing communication overhead while realizing feature complementarity among multiple agents. The collaborative task-driven data reduction representation method of the present invention overcomes the drawback that traditional compression methods rely too much on single-task driving to a certain extent and can reduce communication overhead without affecting the overall performance of the model.

[0017] (3) The method of the present invention adopts a distributed instance-level feature interaction method to overcome the system risks brought by centralized processing. At the same time, it performs instance-level masking processing on private features and only transmits instance-level information to reduce communication overhead. The distributed instance-level feature interaction method of the present invention can also reduce the impact of background information on downstream tasks.

[0018] (4) The method of the present invention also designs a novel feature fusion method that takes temporal information into account, can improve the feature fusion effect of multiple agents, and also improves the system robustness to a certain extent. Brief Description of the Drawings

[0019] Figure 1 is the overall implementation block diagram of the instance-level semantic reduction interaction method for intelligent collaborative perception of the present invention;

[0020] Figure 2 is the implementation flowchart of the distributed instance-level feature interaction of the embodiment of the present invention;

[0021] Figure 3 is the implementation schematic diagram of the data reduction representation of the embodiment of the present invention;

[0022] Figure 4 is the implementation schematic diagram of the feature fusion based on temporal information of the embodiment of the present invention. Detailed Description of the Invention

[0023] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0024] The present invention designs an instance-level semantic reduction interaction method for intelligent collaborative perception. The designed vehicle has the ability of autonomous decision-making, can process single-vehicle perception data, and transmit foreground information, so as to achieve instance-level feature interaction and reduce communication resource overhead. In addition, a data fusion module based on time series is designed to enhance the fusion effect by using time series data, realize efficient feature fusion in complex scenarios, and thus improve the overall perception ability of the vehicle collaborative perception system. The method of the present invention considers the following three aspects of problems and provides corresponding solutions: (1) Design a data reduction representation method driven by collaborative tasks; (2) Design a distributed instance-level feature interaction method; (3) Design a feature fusion method based on time series information.

[0025] First, most of the existing research on feature compression is carried out according to two dimensions of real space and channels. Spatial confidence selection mainly focuses on how to reduce the spatial resolution of the feature map while maintaining important information. This method usually evaluates the importance or confidence of the features at each position for the task, such as by calculating the influence degree of different regions on the final prediction result. Regions with high confidence mean that the features in these parts are more critical for the task, so they should be retained; while regions with low confidence can be compressed or ignored. On the other hand, channel feature downsampling focuses on reducing the number of channels to reduce the model parameter quantity and computational cost. It generally adopts a learnable way, such as using convolution, fully connected or attention mechanism, to let the network adaptively learn which channels are more important and automatically adjust resource allocation. These algorithms compress the data volume transmitted during the collaboration process, but still have two disadvantages: one is that it is difficult to accurately measure the impact on the overall performance of the model when compressing data, which may lead to the loss of key information. The other is that it may be too dependent on the tasks of the supervised compression method, that is, the spatial confidence generation and channel selection algorithms driven by single-vehicle tasks cannot take into account the needs of the entire collaborative perception task. To solve these problems, the method of the present invention designs a data reduction representation method driven by collaborative tasks and compresses the features in a decoupled manner. Specifically, through a decoupling module, the features that the agent needs to transmit are decoupled into private and public parts, and only the private features of the current agent are transmitted, so as to reduce communication overhead while realizing feature complementarity among multiple agents. Since the decoupling module measures the feature relationship between agents, it overcomes the drawback of traditional compression methods being too dependent on single-vehicle task driving to a certain extent. In addition, by distinguishing public features and private features and reducing the transmission of redundant features, the communication overhead can be reduced without affecting the overall performance of the model.

[0026] Second aspect, existing multi-agent feature interaction methods generally directly transmit scene-level features. Existing scene-level feature interaction methods can usually be divided into two categories: message-passing based and centralized processing based. In message-passing based methods, each agent extracts scene-level features based on its local observations and exchanges these features with other agents through a communication protocol. However, one challenge faced by this approach is that excessive message passing may lead to communication bottlenecks, especially when the number of agents is large or the communication bandwidth is limited. In centralized processing based methods, the local observations of all agents are collected at a central node, which processes them uniformly to generate global scene-level features and then distributes them back to each agent. The advantage of this method is that it can optimize the decision-making process from a global perspective. However, the overall risk of the system increases because the performance of the entire system depends on the processing capacity of the central node. In addition, directly transmitting scene-level features causes background features to affect the learning of downstream tasks. Considering the limitations of communication bandwidth and the overall robustness of the system in practical applications, the method of the present invention designs a distributed instance-level feature interaction method. Specifically, through distributed information transmission, the system risks brought by centralized processing are overcome, and at the same time, instance-level masking processing is performed on the features to transmit only instance-level information, reducing communication overhead and also reducing the impact of background information on downstream tasks.

[0027] Third aspect, most existing multi-agent mid-term feature fusion methods directly fuse features. Among the non-learnable fusion methods, such as direct concatenation, weighted average, etc., the feature vectors of different agents are simply processed. This method is easy to implement and does not require additional learning parameters, but it does not consider the complex dependency relationships between different features. Neural network-based methods introduce adaptive weights to dynamically adjust the importance of the features of each agent. By learning the mutual influence between the features of different agents, this type of method can focus on the most relevant feature parts and improve the fusion effect, but these methods do not consider the improvement of the feature fusion effect by temporal information. The method of the present invention designs a feature fusion method based on temporal information, which can utilize temporal information to improve the fusion effect of multi-agent features and also improve the system robustness to a certain extent.

[0028] As Figure 1 shown, in the scenario where the instance-level semantic simplification and interaction method for intelligent collaborative perception according to the embodiment of the present invention is applied, there are two types of agents: intelligent vehicles and roadside devices. Perception models, decoupling modules, fusion modules, and detection modules are all set on the vehicles and roadside devices.

[0029] The vehicle perception model is responsible for converting the data obtained by vehicle sensors into object-level masks and mid-term features. The model consists of a feature extractor and a detection head. Specifically, the feature extractor converts the input data into bird's-eye view (BEV) features with multi-scale feature representations through multiple convolutional layers and fully connected layers, which are highly refined representations of the original data. The detection head uses an attention mechanism to aggregate these BEV features and generates detection objects through a convolutional network. These detection objects are filtered (confidence score threshold and IoU threshold), and object-level masks are generated based on their physical information. The mid-term features in the embodiments of the present invention are the bird's-eye view BEV features with multi-scale feature representations.

[0030] The roadside perception model has a similar structure to the vehicle perception model and is responsible for converting the data obtained by the sensors of roadside devices into object-level masks and mid-term features.

[0031] The decoupling module decouples the mid-term features that the agent needs to transmit into two parts: private features and public features. Each agent transmits the private features. At this time, the private features are still scene-level features, and object-level masks are needed to process the private features to further reduce the communication volume and at the same time reduce the impact of background features on downstream tasks.

[0032] After receiving the object-level private features transmitted by other agents, the fusion module takes the historical features (temporal information), its own public features and private features, and the private features of neighboring agents and sends them into the fusion module together, and finally obtains fusion features with rich semantic information.

[0033] The detection module sends the fusion features into a multi-layer convolutional network to obtain the final detection result.

[0034] For the application scenario under study, the method of the present invention models multi-agent collaborative perception as an optimization problem as follows:

[0035]

[0036] where R represents the optimization objective, B represents the communication bandwidth constraint, N represents the number of agents, is the feature observed by agent i at time t, represents the feature transmitted from agent j to agent i at time t, and Decoder() represents the encoder. This optimization problem can be stated as finding the optimal model under a given communication bandwidth to achieve the best collaborative perception effect among multi-agents. Therefore, the core of this optimization problem is to perform a concise representation of the features required to be transmitted without affecting the model accuracy.

[0037] In the method of the present invention, each agent first inputs the data collected by its own sensors into the perception model, and the perception model converts the input data into an object-level mask and intermediate features; the decoupling module decouples the intermediate features of the agent into two parts: private features and public features; each agent only transmits its own private features, and processes the private features with the object-level mask before transmitting them. This involves an improved collaborative task-driven data reduction representation method and a distributed instance-level feature interaction method of the present invention.

[0038] Considering the limited communication bandwidth and the strict requirements for the overall system robustness in the actual application scenario, the method of the present invention proposes a distributed instance-level feature interaction method, aiming to optimize the data transmission efficiency and enhance the reliability of the system. In today's complex and changing data processing environment, especially for those application scenarios that require real-time response and high concurrent processing capabilities, such as intelligent traffic management, industrial automation, and large-scale Internet of Things deployment, the effective utilization of communication bandwidth and system stability have become key considerations. Traditionally, many data processing tasks rely on a centralized architecture, where all data is collected at a central node for unified processing. However, this mode not only increases the risk of single-point failure but also may lead to network bottlenecks. Especially when faced with a massive data stream, it may significantly reduce the system's response speed and processing efficiency. In addition, the centralized processing method often requires the transmission of a large amount of redundant background information, which further exacerbates the communication burden. Moreover, directly transmitting scene-level features will introduce a large number of background features, and excessive background features will affect the learning of downstream tasks. To solve the above problems, the method of the present invention overcomes the system risks brought by centralized processing by dispersing the computing tasks to multiple nodes, and reduces the interference of background features on downstream tasks through object-level instance interaction. Specifically, the present invention provides a distributed instance-level feature interaction method, introducing an instance-level mask processing technology, and only transmitting the feature information related to a specific instance, rather than all attributes of the entire data set or object. For example, in a video surveillance application, the system can only transmit the key features of the detected pedestrians or vehicles, rather than the complete image frame, thus significantly reducing the communication overhead. This process ensures that even under poor network conditions, efficient data transmission can be maintained, and the impact of background information (such as irrelevant environmental details) on downstream tasks is minimized. In this way, the model can focus more on the important characteristics of the target object, improving the accuracy and efficiency of task execution. As Figure 2As shown in the figure, the process of the distributed instance-level feature interaction method is divided into five parts: (1) First, the BEV features are sent to the detection head to generate a large number of candidate regions and their corresponding confidence scores; (2) Then, according to the confidence threshold and the IoU threshold, the candidate regions are screened and de-duplicated by NMS (Non-Maximum Suppression), and an object-level mask that can act on the features is generated based on the spatial position information of the screened candidate regions. The IoU threshold (Intersection over Union threshold) is used to measure the degree of overlap between candidate regions. In the embodiment of the present invention, the spatial feature values corresponding to these candidate regions after screening are marked as 1, and the spatial feature values of other regions are marked as 0 to obtain the object-level mask. The object-level mask changes according to the observed data at each moment. The object-level mask is a 0-1 matrix, and the size of the matrix is the same as the size of the private features. (3) In order to further reduce the communication volume, the intermediate features can be divided into public features and private features, and the private features of the current agent are masked to obtain the object-level private features. (4) Communication starts between multiple agents. First, simple information is sent, that is, information such as the position and perception range of the agent. (5) After determining the cooperation partners and the overall cooperation perception range accordingly, the object-level private features are selectively sent to other agents.

[0039] The method of the present invention innovatively implements a collaborative task-driven data reduction and representation method, aiming to overcome the common limitations in traditional data compression technologies. In traditional data processing and transmission schemes, compression algorithms are often designed with the needs of a single agent as the core. Although this single-agent task-driven method can meet the requirements in specific scenarios, in the face of complex scenarios, it may lead to serious data redundancy or information loss, thus affecting the real-time performance and effect of the system. To address these challenges, the method of the present invention realizes more efficient data compression and transmission by introducing the concept of collaborative task-driven, not only considering the needs of a single agent, but also particularly focusing on the complementarity and redundancy of features between multiple agents, so as to provide more optimized data processing capabilities in complex multi-task environments. Specifically, the collaborative task-driven data reduction and representation method of the present invention distinguishes between public features and private features in the agent's observed data. Public features refer to the background information observed by the agent and the redundant part of the public object features that can be observed by multiple agents. Private features refer to the features observed by the agent that are different from other agents, such as the blind spots of other agents or the special information of a public object from the perspective of this agent. By clearly dividing these two types of features, the present invention can intelligently reduce the transmission of redundant information. For example Figure 3As shown, during the data reduction and representation process, the private feature encoder and public feature encoder parameters are shared among agents. Their purpose is to compress features at the channel level, thereby reducing the communication volume among agents. To capture deeper features hidden in the intermediate features, the method of the present invention uses a combination of convolutional layers and fully connected layers to implement an encoder for further feature extraction. In terms of design, the private feature encoder and public feature encoder can have similar architectures, but their parameters are independently trained. Agents A and B use their own public feature encoders and private feature encoders to decouple the intermediate features obtained from their respective perception models into public features and private features. After obtaining the private features and public features, a consistency constraint needs to be added to the public features among multiple agents to ensure that the public feature information is consistent. To ensure that the private features and public features can retain the original feature information as losslessly as possible, we need to design a decoder to implement reconstruction loss supervision so that the decoded features are as consistent as possible with the original features. Mark this decoder as decoder A. Specifically, after sending the private features and public features into the decoder A with shared parameters, the reconstructed features are output, and reconstruction loss supervision is added between the input features (i.e., intermediate features) and the reconstructed features, so that while the decoupling module compresses the features, key information is not lost as much as possible. In addition, decoder A only serves for the reconstruction loss during the training process and can be removed from the decoupling module during the inference process to speed up the inference speed.

[0040] In the method of the present invention, after the cooperating agents transmit object-level private features to each other, the agents need to perform feature fusion. In a multi-agent system, feature fusion is a key link for achieving efficient cooperation and decision-making. Existing multi-agent intermediate feature fusion methods usually adopt direct fusion strategies, such as non-learnable fusion methods like direct concatenation and weighted average. Although these methods are easy to implement and do not require additional learning parameters, they do not fully consider the complex dependencies between different features, resulting in the fused information may not accurately reflect the actual situation. The neural network-based method dynamically adjusts the importance of each agent's features by introducing adaptive weights. This method can learn the mutual influence between different agents' features, enabling the system to focus more on the most relevant feature parts, thereby improving the fusion effect. However, such methods mainly focus on the fusion of static features and ignore the important role of temporal information in the feature fusion effect. To make up for the deficiencies of existing methods, the method of the present invention proposes a feature fusion method based on temporal information, which not only considers the complex dependencies between features of different agents but also particularly emphasizes the role of temporal information. This method not only improves the effect of multi-agent feature fusion but also enhances the robustness and adaptability of the system to a certain extent, enabling it to cope with more complex dynamic environments. Such as Figure 4As shown, after the corresponding multi-agent feature fusion module receives the object-level private features transmitted from other agents, it sends its own private features and the received object-level private features to the attention module. Here, its own private features are used as the query feature Q, and the object-level private features of other vehicles are used as the key K and value V. The attention module outputs the fused private features, which contain the key information of the current agent and the collaborative agents. The attention weight F calculated in this process is expressed as follows:

[0041]

[0042] where, W Q , W K and W V are the weight matrices of the linear transformations applied to Q, K, and V respectively, d k is the scaling factor in the attention, usually equal to the dimension of K, and the superscript T represents the transpose. At this time, the fused private features still lack background information and redundant information of some common objects. The fused private features are fused with the public features of the current agent through the decoder B to obtain the preliminary fusion features. After obtaining the preliminary fusion features, it is necessary to fuse the key information of the fusion features of the previous frame of the current agent to achieve a better perception effect. To fuse the temporal information, the method of the present invention first concatenates the preliminary fusion features of the current frame of the current agent and the fused private features of the previous frame in the channel dimension, and then sends them to the self-attention module for processing. After that, the fusion features at the current moment are obtained, and sending them to the detection head of the final detection module can obtain a more comprehensive and accurate perception result.

[0043] Except for the technical features described in the specification, they are all known technologies to those skilled in the art. The present invention omits the description of well-known components and well-known technologies to avoid redundancy and unnecessary limitation of the present invention. The embodiments described in the above examples do not represent all embodiments consistent with the present application. Based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.

Claims

1. An instance-level semantically simplified interaction method for intelligent collaborative perception, characterized in that: A perception model, a decoupling module, a fusion module and a detection module are set on each intelligent agent in the autonomous driving application scenario; the method comprises the following steps: Step 1: Each agent inputs the data collected by its own sensors into the perception model, and the perception model converts the input data into object-level masks and mid-term features; The perception model consists of a feature extractor and a detection head. The feature extractor converts input data into bird's-eye view BEV features represented by multi-scale features. The detection head aggregates BEV features using an attention mechanism, generates candidate regions and confidence scores for the existence of detection objects through a convolutional network, and screens them according to confidence thresholds and IoU thresholds to generate object-level masks. The mid-term features are BEV features represented by multi-scale features; Step 2: The decoupling module of the agent decouples its own mid-term features into private features and public features. The agent only transmits its own private features, and uses the object-level mask to process the private features to obtain object-level private features, and then transmits them out; the agents first transmit information about their current positions and perception ranges to each other, determine the collaborative partners and the overall collaborative perception range, and then the collaborative agents transmit object-level private features to each other; Step 3: After receiving the object-level private features transmitted by the collaborative agent, the agent inputs its own historical features, its own public features and private features, and the object-level private features of the collaborative agent into the fusion module. The fusion module uses a feature fusion method based on temporal information to obtain fused features and inputs them into the detection module. The detection module sends the fused features to the multi-layer convolutional network for target detection. The fusion module uses a feature fusion method based on time series information to obtain fusion features, including: first, inputting the object-level private features of the collaborative agent and the private features of the current agent into the attention module to obtain the fused private features, then fusing the public features of the current agent to obtain the preliminary fusion features, splicing the preliminary fusion features with the fused private features of the previous frame of the current agent on the channel, and then sending them to the self-attention module for processing to obtain the fusion features at the current moment.

2. The method according to claim 1, characterized in that In step 1, each agent processes the private features using the object-level mask and then transmits them out, thereby realizing distributed instance-level feature interaction between agents; The generation of object-level masks includes: the detection head of the perception model generates candidate regions for the detection object, and then screens the candidate regions according to the confidence threshold and IoU threshold, and performs non-maximum suppression (NMS) to deduplicate the candidate regions. The values ​​of the final screened candidate regions are marked as 1, and the values ​​of other regions are marked as 0 to generate an object-level mask. The matrix size of the mask is consistent with the size of the private feature.

3. The method according to claim 1, characterized in that In the step 2, each agent decouples the mid-term features from the private features and transmits only the private features, thereby realizing the data simplification driven by the collaborative task between the agents; in the decoupling module, a private feature encoder and a public feature encoder are designed by combining a convolutional layer and a fully connected layer, and the mid-term features are respectively input into the private feature encoder and the public feature encoder to obtain private features and public features; Add consistency constraints to the common features between agents to ensure the consistency of the information of the common features; share the parameters of the private feature encoder and the public feature encoder between agents; During training, the parameters of the private feature encoder and the public feature encoder are trained independently, and a decoder with shared parameters among the agents is set. The agent sends its own private features and public features to the decoder to output the reconstructed features. Reconstruction loss supervision is added between the input mid-term features and the reconstructed features, and the parameters in the decoupling module are trained to make the reconstructed features and the input mid-term features as consistent as possible. After the training is completed, the private feature encoder and the public feature encoder with trained parameters are used to form a decoupling module.

4. The method according to claim 1, characterized in that: In step 3, the private features of the current agent are used as query features Q, and the object-level private features received from other collaborative agents are used as keys K and values ​​V, which are fused through the attention module.

Citation Information

Patent Citations

  • A communication-sensitive multi-agent collaboration method

    CN114327935B

  • Automatic driving collaborative perception method and device based on multi-modal large model

    CN118887632A

  • Collaborative confidence fusion perception method applied to automatic driving

    CN119027934A

  • Vehicle-road cooperative joint sensing method and system under non-ideal condition

    CN119091152A

  • Heterogeneous graph network-based multi-modal cooperative detection method and system

    CN115512319A