Peer-to-peer network federal collaborative decision-making method and device

Through the peer-to-peer network federated collaborative decision-making method, the BEV bird's-eye feature conversion and federal integrated knowledge distillation algorithm are used to solve the problems of equipment heterogeneity, low communication efficiency and insufficient modal complementarity in P2P collaborative decision-making, and efficient and real-time multi-view compensation fusion feature interaction and collaborative decision-making are achieved, improving decision-making accuracy.

CN120499189AActive Publication Date: 2025-08-15YUNNAN NORMAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510909791.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-15
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing P2P collaborative decision-making methods cannot adapt to device hardware heterogeneity, inefficient communication efficiency, missing target depth semantics and insufficient modal complementarity, resulting in insufficient real-time and decision-making accuracy.

Method used

The peer-to-peer network federated collaborative decision-making method is adopted, and multi-modal fusion is performed by initializing the deep network model, using BEV bird's-eye feature conversion technology, combining Top-k sorting dynamic mask selection and federal integrated knowledge distillation algorithm to realize the interaction and collaborative decision-making of multi-view compensation fusion features.

Benefits of technology

It improves the information communication efficiency between heterogeneous devices, meets real-time requirements, enhances deep semantic information perception, and uses multimodal complementarity to improve decision-making accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499189A_ABST
    Figure CN120499189A_ABST
Patent Text Reader

Abstract

The invention relates to a peer-to-peer network federation collaborative decision-making method and device, and belongs to the field of peer-to-peer network edge computing. The method comprises the following steps: firstly, initializing a deep network of each participant, and extracting two modal features from RGB modal data and radar point cloud modal data by using a feature extraction network; then, the BEV aerial view feature conversion technology is adopted to convert the two modal features into a BEV space for alignment to obtain a multi-modal fusion feature, the multi-modal fusion feature is compressed, a difference value is calculated to generate a sparse compensation feature, and interaction with other participants is carried out through a Top-k dynamic mask selection strategy to obtain a multi-view compensation fusion feature; and finally, inputting the compensation fusion features into a prediction network, training the prediction network by using a federal integrated knowledge distillation algorithm, and interacting for many times to realize peer-to-peer network federal collaborative decision-making. According to the method, the complementarity of different modal views is fully utilized, the problems of semantic missing, heterogeneity, real-time performance and the like of peer-to-peer network participants in decision making are solved, and the method is suitable for multi-participant collaborative decision making with different parameters and hardware heterogeneity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a peer-to-peer network federation collaborative decision-making method and device, and belongs to the field of peer-to-peer network edge computing. Background Art

[0002] With the continuous advancement of network technology and computing power, peer-to-peer (P2P) networks, due to their advantages such as decentralization, self-organization, high robustness, and good scalability, have been widely adopted in modern complex intelligent systems, especially in multi-device collaboration such as the Internet of Things (IoT), multi-drone systems, vehicle-to-everything (V2X), blockchain, and edge computing. By enabling communication and interaction between nodes, P2P networks greatly improve system flexibility and responsiveness, becoming a critical technical foundation for decentralized intelligent perception, collaborative decision-making, and real-time data transmission.

[0003] Current P2P collaborative decision-making methods typically rely on the cloud to interact with each participant (peer), leveraging the cloud's powerful computing capabilities to centrally collect data, process data, make decisions, and then distribute the results to each participant. However, while this approach fully utilizes the cloud's computing advantages, it still has the following shortcomings in achieving fast and convenient collaborative decision-making:

[0004] (1) Heterogeneity of hardware among P2P participants. Current P2P collaborative decision-making usually adopts a centralized federated learning architecture, which requires all participants to share the same model structure with the cloud. However, the heterogeneity of hardware configuration among P2P participants limits the application of this method in actual deployment.

[0005] (2) Real-time P2P communication. The current cloud-to-P2P communication requires the participants to upload data first, which is then aggregated and processed by the cloud before the calculation results are sent down. This model is time-consuming due to the time-consuming data transmission and centralized processing, resulting in low communication efficiency and difficulty in meeting real-time requirements.

[0006] (3) Lack of P2P deep semantic information. Current P2P collaborative decision-making methods do not fully consider the impact of target deep semantics, resulting in the loss of semantic information under different views, which affects the accuracy and effectiveness of collaborative decision-making;

[0007] (4) Complementarity of P2P multimodal information. Current P2P collaborative decision-making methods mainly rely on a single modality for analysis and decision-making, which does not fully utilize the complementary advantages between different modalities and views, limiting the model's perception ability and leading to a decline in decision-making performance.

[0008] In response to the above problems, the present invention discloses a peer-to-peer network federated collaborative decision-making method and device, which solves the problems in related technologies such as the inability of P2P decision-making methods to adapt to hardware heterogeneity, low communication efficiency, and missing target depth semantics. It not only improves the efficiency of data transmission and meets real-time requirements, but also fully considers the complementarity of information from different modal views, and can be applied to P2P collaborative decision-making tasks with different parameters and hardware heterogeneity. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a peer-to-peer network federated collaborative decision-making method and device thereof, which solves the problems in related technologies such as the inability of P2P collaborative decision-making methods to adapt to device heterogeneity, low communication efficiency, lack of target depth semantics and modal complementarity, ensures data transmission efficiency and real-time requirements, and improves the accuracy of collaborative decision-making in P2P networks.

[0010] The technical solution of the present invention is: a peer-to-peer network federation collaborative decision-making method, the specific steps of which are:

[0011] S1: Initialize the deep network models involved in decision-making in the peer-to-peer (P2P) network, and use the feature extraction network to encode the RGB modal data and radar point cloud modal data to obtain the two modal features;

[0012] S2: Based on the two modal features, the BEV bird's-eye view feature conversion technology is used to convert them into the BEV space, and then alignment is performed to obtain the BEV multimodal fusion features;

[0013] S3: compressing the BEV multimodal fusion features and calculating the difference values to obtain sparse compensation features, communicating with other participants in the peer-to-peer network, and obtaining multi-view compensation fusion features through a Top-k sorting dynamic mask selection strategy;

[0014] S4: The multi-view compensation fusion features are input into the carried prediction network and trained using the federated integrated knowledge distillation algorithm, so as to continuously interact and realize peer-to-peer network federated collaborative decision-making.

[0015] Optionally, the S1 initializes each deep network model involved in decision-making in a peer-to-peer (P2P) network, and uses a feature extraction network to encode the RGB modal data and the radar point cloud modal data to obtain two modal features, the steps including:

[0016] S1.1: Initialize the deep network model of each participant in the peer-to-peer network ,in represents the RGB modality feature extraction network, represents the radar point cloud feature extraction network, Represents different task prediction networks;

[0017] S1.2: Encode the RGB modality data using the RGB modality feature extraction network to obtain semantic structure features;

[0018] S1.3: Use the radar point cloud feature extraction network to encode the radar point cloud modal data to obtain geometric structure features.

[0019] Optionally, the S2 is based on the two modal features, converts them into BEV space using BEV bird's-eye view feature conversion technology, and aligns them to obtain BEV multimodal fusion features, and the steps include:

[0020] S2.1: For the obtained semantic structure features, use RGB space perception target depth semantics to calculate the semantic structure features, and use BEV bird's eye view feature conversion technology to convert them into BEV space;

[0021] S2.2: For the obtained geometric structure features, perform the z-axis flattening operation and convert them into BEV space using the BEV bird's-eye view feature conversion technology;

[0022] S2.3: Align the converted features in the BEV space to form BEV multimodal fusion features.

[0023] Optionally, the step S3 compresses the BEV multimodal fusion features and calculates difference values to obtain sparse compensation features, and communicates with other participants in the peer-to-peer network to obtain multi-view compensation fusion features through a Top-k sorting dynamic mask selection strategy. The steps include:

[0024] S3.1: The obtained BEV multimodal fusion features are compressed using 1×1 convolution and passed through a The sliding window of different sizes acts on the compressed features, and the difference between the features at each position and the central features is calculated to obtain the sparse compensation features;

[0025] S3.2: Receive the sparse compensation features sent by other participants, interact with the sparse compensation features obtained by the current participant, calculate the information interaction compensation score and sort it into Top-k, and form multi-view compensation fusion features through the Top-k dynamic mask selection strategy.

[0026] Optionally, the step S4 inputs the multi-view compensation fusion feature into the carried prediction network and trains it using a federated integrated knowledge distillation algorithm, so as to continuously interact and realize peer-to-peer network federated collaborative decision-making. The steps include:

[0027] S4.1: Input the obtained multi-view compensation fusion features into the prediction network, and use the federated integrated knowledge distillation algorithm to train and update the deep network model parameter;

[0028] S4.2: Based on multiple information interactions between participants, repeat the above steps to achieve peer-to-peer network federation collaborative decision-making.

[0029] In addition, the present application also proposes a peer-to-peer network federation collaborative decision-making device, which includes:

[0030] Peer network participant initialization module, used to initialize the deep network model of each participant ;

[0031] BEV bird's-eye view feature fusion module, used to generate multimodal fusion features;

[0032] Interactive compensation feature fusion module, used to generate multi-view compensation fusion features;

[0033] The multi-participant collaborative decision-making module is used to input the obtained multi-view compensation fusion features into the carried decision-making model for collaborative decision-making.

[0034] In addition, the present application also proposes a peer-to-peer network federal collaborative decision-making device, including a memory, a processor, and a peer-to-peer network federal collaborative decision-making method program stored in the memory and runnable on the processor, wherein the processor executes the peer-to-peer network federal collaborative decision-making method program to implement the steps of the peer-to-peer network federal collaborative decision-making method described above.

[0035] In addition, the present application also proposes a computer-readable storage medium, on which a peer-to-peer network federal collaborative decision-making method program is stored. When the peer-to-peer network federal collaborative decision-making method program is executed by a processor, the steps of the peer-to-peer network federal collaborative decision-making method described above are implemented.

[0036] The beneficial effects of the present invention are:

[0037] (1) It supports collaborative decision-making among multiple heterogeneous participants and is more powerful and universal than existing methods;

[0038] (2) Without the need for cloud model participation, the obtained BEV multimodal fusion features are compressed and the difference values are calculated to obtain sparse compensation features, which improves the efficiency of information communication between participants;

[0039] (3) The multi-view compensation fusion feature is formed through the Top-k sorting dynamic mask selection strategy, which can better perceive the deep semantic information;

[0040] (4) It takes advantage of the complementary advantages of information between different modalities and views of multiple participants in the peer-to-peer network. By using the federated integrated knowledge distillation algorithm, it better transfers information between different modalities and improves the accuracy of model decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of the steps of the present invention;

[0042] Figure 2 This is an overall schematic diagram of peer-to-peer network federation collaborative decision-making according to Example 1 of the present invention;

[0043] Figure 3 This is a schematic diagram of a specific process flow for a single participant to make a decision according to Example 1 of the present invention;

[0044] Figure 4 This is a schematic diagram of depth semantic calculation of RGB space perception target according to Example 1 of the present invention;

[0045] Figure 5 Schematic diagram of a BEV feature generation model integrating deep semantics according to Example 1 of the present invention;

[0046] Figure 6 This is a schematic diagram of multi-view compensation fusion feature calculation in Example 1 of the present invention;

[0047] Figure 7 Schematic diagram of the federated integrated knowledge distillation algorithm according to Example 1 of the present invention;

[0048] Figure 8 This is a schematic diagram of the hardware structure involved in Example 1 of the present invention. DETAILED DESCRIPTION

[0049] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0050] Example 1: Figure 1 The figure shows an overall schematic diagram of a peer-to-peer network federated collaborative decision-making method. The implementation process mainly includes four steps: S1: initialization of peer network participants; S2: BEV (bird's eye view) bird's eye view feature fusion; S3: interactive compensation feature fusion; S4: multi-participant collaborative decision-making. Figure 3 The figure shows a schematic diagram of the specific decision-making process of a single participant. The following describes each step in detail:

[0051] S1: Specific implementation steps for initialization of peer network participants.

[0052] Initialize the deep network models involved in decision-making in the peer-to-peer (P2P) network, and use the feature extraction network to encode the RGB modal data and the radar point cloud modal data to obtain two modal features, specifically:

[0053] S1.1: Initialize the deep network model of each participant in the peer-to-peer network ,in represents the RGB modality feature extraction network, represents the radar point cloud feature extraction network, Represents different task prediction networks;

[0054] S1.2: Encode the RGB modality data using the RGB modality feature extraction network to obtain semantic structure features. Specifically:

[0055] set up Indicates the Participants in the sample exist RGB modality data at the moment, through the RGB modality feature extraction network , we can get the RGB semantic structure features ,in , and Represent the height, width and number of channels of the feature map respectively.

[0056] S1.3: Use the radar point cloud feature extraction network to encode the radar point cloud modal data to obtain geometric structure features. Specifically:

[0057] set up Indicates the Participants in the sample exist The radar point cloud modal data at each moment is extracted through the radar point cloud feature extraction network , we can get the geometric structure characteristics of the radar point cloud ,in , and Represents the X-axis, Y-axis, and Z-axis coordinate numbers respectively.

[0058] S2: Specific implementation steps of BEV bird’s-eye view feature fusion.

[0059] Based on the two modal features, the BEV bird's-eye view feature conversion technology is used to convert them into the BEV space, and alignment is performed to obtain the BEV multimodal fusion features, specifically:

[0060] S2.1: For the semantic structure features obtained in S1.2, the semantic structure features are calculated using the RGB space perception target depth semantics, and converted to the BEV space using the BEV bird's-eye view feature conversion technology. Specifically:

[0061] RGB semantic structure features obtained based on S1.2 ,In order to better describe the target depth semantic information, the RGB space perception target depth semantic calculation method is used, such as Figure 4 The figure shows the schematic diagram of RGB space perception target depth semantic calculation, where Indicates the data perception height, Represents the estimated target height, using the perceptron parameter matrix , the rotation matrix and translation vectors , for each pixel coordinate , get the length from the target to the ground The length from the camera to the target It can be expressed as:

[0062] (1)

[0063] in, Representation matrix The coordinate of the second component corresponding to the Y axis; Representation matrix The element in the second row and first column of Representation matrix The element in the 2nd row and 2nd column of ; Representation matrix The element of the 2nd row and 3rd column. For each pixel coordinate , estimate the depth of the object , expressed as:

[0064] (2)

[0065] By utilizing the camera parameter network Estimated target height , so the estimated height can be Input Obtaining deep semantic information through calculation , and then through the feature map Perform multiplication to convert BEV bird's-eye view features to obtain target semantic structure features .

[0066] S2.2: Apply z-axis flattening to the geometric structure features of the radar point cloud obtained in S1.3 and convert them into BEV space using the BEV bird's-eye view feature conversion technology. Specifically:

[0067] Given radar point cloud data , by following The axis is flattened to perform BEV bird's-eye view feature conversion to obtain the geometric structure features of the target .

[0068] S2.3: Align the features converted from S2.1 and S2.2 in the BEV space to form BEV multimodal fusion features. Specifically:

[0069] like Figure 5Schematic diagram of the BEV feature generation model for integrating deep semantics, given the full convolutional network operation , the target semantic structure features obtained in S2.1 and the geometrical structural features obtained by S2.2 Input into the convolutional network to obtain BEV multimodal fusion features , expressed as:

[0070] (3)

[0071] in, Indicates that the target semantic structure features and geometric features The two tensors are concatenated and merged into one tensor along the column direction.

[0072] S3: Specific implementation steps of interactive compensation feature fusion.

[0073] The BEV multimodal fusion features are compressed and the difference values are calculated to obtain sparse compensation features, which are then communicated with other participants in the peer-to-peer network. The multi-view compensation fusion features are obtained through the Top-k sorting dynamic mask selection strategy, specifically:

[0074] S3.1: Compress the obtained BEV multimodal fusion features and calculate the difference value to obtain sparse compensation features. Specifically:

[0075] Based on S2.2 BEV multimodal fusion features obtained from participants To reduce the information transmission overhead between participants, we first use 1×1 convolution to Compress to obtain compressed features , then through a The sliding window of size acts on the compressed features, aiming to calculate the difference between the features at each position and the central features, and then pass a sigmoid function Normalize the calculated differences to interval, taking the average difference of the entire window as the sparse compensation feature , used for information interaction between participants, expressed as:

[0076] (4)

[0077] in, , 、 Representation feature map No. Row, No. The pixel value of the column; 、 Represents a sliding window The subscript value to sum.

[0078] S3.2: The sparse compensation features obtained in S3.1 are exchanged with other participants to form multi-view compensation fusion features through the Top-k sorting dynamic mask selection strategy. Specifically:

[0079] like Figure 6 This is a schematic diagram of multi-view compensation fusion feature calculation. Each participant receives the Sparse compensation features sent by participants When , the information interaction compensation score can be calculated ,in and Respectively represent and The sparse compensation features of participants, symbol Represents the tensor and Calculate by multiplying the corresponding elements, and then Perform Top-k sorting to obtain the feature interaction mask, which is expressed as , and finally through the feature interaction mask Get from Multi-view compensation fusion features of participants , For the BEV multimodal fusion features obtained from 10 participants.

[0080] S4: Specific implementation steps for multi-participant collaborative decision-making.

[0081] The multi-view compensation fusion features are input into the prediction network and trained using the federated integrated knowledge distillation algorithm. This allows for continuous interaction to achieve peer-to-peer network federated collaborative decision-making, specifically:

[0082] S4.1: Input the multi-view compensation fusion features obtained in S3.2 into the prediction network carried out, and use the federated integrated knowledge distillation algorithm to train and update the deep network model Parameters, specifically:

[0083] like Figure 7 This is a schematic diagram of the federated integrated knowledge distillation algorithm. participants, given the participant's deep network model , the results from S3.2 are respectively from Participants and Participants input the prediction network Different prediction outputs are obtained, which are expressed as and ,in, For the To better utilize the complementary advantages of information between different modalities and views of multiple participants in a peer-to-peer network, this paper proposes a federated integrated knowledge distillation algorithm to better transfer information between different modalities. The formula is as follows:

[0084] (5)

[0085] in, is the total loss, Represents the true label of specific task data, function Represents the cross entropy loss, the first term The KL divergence is used to measure the consistency of knowledge from different participants for interactive consistency knowledge distillation loss. The measurement comes from The local knowledge supervision loss of each participant, the third The measurement comes from The local knowledge supervision loss of each participant.

[0086] S4.2: Based on multiple information interactions between participants, repeat the above steps to achieve peer-to-peer network federation collaborative decision-making.

[0087] This application also proposes a peer-to-peer network federation collaborative decision-making device, such as Figure 8 The figure shows a schematic diagram of the structure of a device for running a program of a peer-to-peer network federated collaborative decision-making method in a hardware operating environment involved in an embodiment of the present application.

[0088] like Figure 8 As shown, a peer-to-peer network federation collaborative decision-making device includes: a processor 1001, such as a central processing unit CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between the various component modules, the user interface 1003 may include a display screen, an input unit such as a keyboard, etc., and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed random access memory RAM, or a stable non-volatile memory NVM, such as a disk storage. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.

[0089] Those skilled in the art will understand that Figure 8The structure shown in does not constitute a limitation on a peer-to-peer network federated collaborative decision-making device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0090] Optionally, the memory 1005 is connected to the processor 1001, and the processor 1001 can be used to control the operation of the memory 1005 and read the data in the memory 1005 to implement a peer-to-peer network federation collaborative decision-making method.

[0091] Alternatively, as Figure 8 As shown, the memory 1005 as a storage medium may include an operating system, a data storage module, a network communication module, a user interface module and a multi-participant collaborative decision-making method program.

[0092] Optionally, in Figure 8 In the peer-to-peer network federated collaborative decision-making device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the peer-to-peer network federated collaborative decision-making device of the present application can be set in a peer-to-peer network federated collaborative decision-making device device.

[0093] like Figure 8 As shown, the peer-to-peer network federated collaborative decision-making device calls the multi-participant collaborative decision-making method program stored in the memory 1005 through the processor 1001, and executes the following Figure 2 As shown, the four relevant steps of a peer-to-peer network federated collaborative decision-making method provided in the first embodiment of the present application are: S1: initialization of peer network participants; S2: BEV bird's-eye view feature fusion; S3: participant interaction compensation feature fusion; S4: multi-participant collaborative decision-making.

[0094] In addition, the present application also proposes a computer-readable storage medium, on which a peer-to-peer network federated collaborative decision-making method program is stored. When the multi-participant collaborative decision-making method program is executed by a processor, the steps of the peer-to-peer network federated collaborative decision-making method as described above are implemented.

[0095] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0096] It should be noted that in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present application may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0097] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0098] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A peer-to-peer network federation collaborative decision-making method, characterized in that: The federal collaborative decision-making method includes: S1: Initialize the deep network models involved in decision-making in the peer-to-peer network, and use the feature extraction network to encode the RGB modal data and the radar point cloud modal data to obtain the two modal features; S2: Based on the two modal features, the BEV bird's-eye view feature conversion technology is used to convert them into the BEV space, and then alignment is performed to obtain the BEV multimodal fusion features; S3: compressing the BEV multimodal fusion features and calculating the difference values to obtain sparse compensation features, and interactively communicating with other participants in the peer-to-peer network to obtain multi-view compensation fusion features through a Top-k sorting dynamic mask selection strategy; S4: Input the multi-view compensation fusion features into the carried prediction network and train it using the federated integrated knowledge distillation algorithm, so as to continuously interact to achieve peer-to-peer network federated collaborative decision-making.

2. A peer-to-peer network federation collaborative decision-making method according to claim 1, characterized in that: The S1 initializes each deep network model involved in decision-making in the peer-to-peer network, and uses a feature extraction network to encode the RGB modal data and the radar point cloud modal data to obtain two modal features. The steps include: S1.1: Initialize the deep network model of each participant in the peer-to-peer network ,in represents the RGB modality feature extraction network, represents the radar point cloud feature extraction network, Represents different task prediction networks; S1.2: Encode the RGB modality data using the RGB modality feature extraction network to obtain semantic structure features; S1.3: Use the radar point cloud feature extraction network to encode the radar point cloud modal data to obtain geometric structure features.

3. A peer-to-peer network federation collaborative decision-making method according to claim 1, characterized in that: The S2 is based on the two modal features, uses the BEV bird's-eye view feature conversion technology to convert to the BEV space, and aligns to obtain the BEV multimodal fusion features, the steps including: S2.1: For the obtained semantic structure features, use RGB space perception target depth semantics to calculate the semantic structure features, and use BEV bird's eye view feature conversion technology to convert them into BEV space; S2.2: For the obtained geometric structure features, perform the z-axis flattening operation and convert them into BEV space using the BEV bird's-eye view feature conversion technology; S2.3: Align the converted features in the BEV space to form BEV multimodal fusion features.

4. A peer-to-peer network federation collaborative decision-making method according to claim 1, characterized in that: The S3 compresses the BEV multimodal fusion features and calculates the difference value to obtain sparse compensation features, and interacts with other participants in the peer network to obtain multi-view compensation fusion features through a Top-k sorting dynamic mask selection strategy. The steps include: S3.1: The obtained BEV multimodal fusion features are compressed using 1×1 convolution and passed through a The sliding window of different sizes acts on the compressed features, and the difference between the features at each position and the central features is calculated to obtain the sparse compensation features; S3.2: Receive the sparse compensation features sent by other participants, interact with the sparse compensation features obtained by the current participant, calculate the information interaction compensation score and sort it into Top-k, and form multi-view compensation fusion features through the Top-k dynamic mask selection strategy.

5. A peer-to-peer network federation collaborative decision-making method according to claim 1, characterized in that: The S4 inputs the multi-view compensation fusion features into the carried prediction network and trains it using the federated integrated knowledge distillation algorithm. This continuously interacts to achieve peer-to-peer network federated collaborative decision-making. The steps include: S4.1: Input the obtained multi-view compensation fusion features into the prediction network, and use the federated integrated knowledge distillation algorithm to train and update the deep network model parameter; S4.2: Based on multiple information interactions between participants, repeat the above steps to achieve peer-to-peer network federation collaborative decision-making.

6. A device for implementing the peer-to-peer network federation collaborative decision-making method according to claim 1, characterized in that: The device comprises: Peer network participant initialization module, used to initialize the deep network model of each participant ; BEV bird's-eye view feature fusion module, used to generate multimodal fusion features; Interactive compensation feature fusion module, used to generate multi-view compensation fusion features; The multi-participant collaborative decision-making module is used to input the obtained multi-view compensation fusion features into the carried decision-making model for collaborative decision-making.

7. A peer-to-peer network federation collaborative decision-making device, characterized in that: The invention comprises a memory, a processor and a peer-to-peer network federation collaborative decision-making method program stored in the memory and executable on the processor, wherein the processor executes the peer-to-peer network federation collaborative decision-making method program to implement the steps of the peer-to-peer network federation collaborative decision-making method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a peer-to-peer network federation collaborative decision-making method program, which, when executed by a processor, implements the steps of a peer-to-peer network federation collaborative decision-making method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-granularity representation map reconstruction method

    CN119785315A

  • Vector map construction method based on autoregression depth estimation and related device

    CN120047640A

  • Electronic device and method with birds-eye-view image processing

    US20250086953A1