Method and device for updating swarm intelligence knowledge in resource-constrained scenarios
By constructing a dual-channel knowledge interaction architecture, the problems of inconsistent knowledge sharing and unstable model updates among swarm intelligence nodes in resource-constrained scenarios are solved, achieving stable and consistent prototype representation and robust cross-node parameter sharing, while reducing communication volume.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN MSU-BIT UNIVERSITY
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-24
AI Technical Summary
In resource-constrained scenarios, the lack of effective knowledge sharing among swarm intelligence nodes, inconsistent prototype semantics, and unstable model updates result in high communication burdens and difficulty in learning effective representations on long-tail or scarce data.
A dual-channel knowledge interaction architecture is constructed, including a prototype collaborative sharing channel and a parameter collaborative sharing channel. A global prototype is generated through global semantic alignment and feature parameters are constrained and corrected to achieve stable knowledge sharing and robust parameter transmission across nodes.
With limited data, we achieved stable and consistent prototype representation and robust cross-node parameter sharing, reducing communication volume and improving model stability and knowledge transfer capabilities.
Smart Images

Figure CN121351982B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of swarm intelligence knowledge updating technology, and in particular to a method and apparatus for swarm intelligence knowledge updating in resource-constrained scenarios. Background Technology
[0002] In swarm intelligence or federated learning frameworks, agents typically rely on local data sources for model training and periodically exchange information with the central node or other nodes to achieve the overall evolution of collective knowledge.
[0003] Among related technologies, federated learning based on parameter averaging can be used. This approach focuses on directly aggregating model weights. However, when the data source is extremely unbalanced or limited, the local model differences between different nodes are significant, leading to unstable convergence of the aggregation results. The model has difficulty learning effective representations on long-tail or scarce data, and the amount of parameters transmitted across nodes is large, resulting in a high communication burden.
[0004] Another learning method is prototype-based learning. Prototype learning can improve model robustness and greatly reduce communication costs by using prototypes as communication carriers. However, prototypes are difficult to unify semantics across different nodes, are prone to deviation when data is scarce, cannot represent the true distribution, lack cross-node adversarial verification, and prototype quality is not guaranteed. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and computer-readable storage medium for updating swarm intelligence knowledge in resource-constrained scenarios, which solves problems such as lack of effective knowledge sharing among swarm intelligence nodes, inconsistent prototype semantics, and unstable model updates under resource-constrained conditions.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] Firstly, a method for updating swarm intelligence knowledge in resource-constrained scenarios is provided, including:
[0008] A dual-channel knowledge interaction architecture is constructed for each intelligent node, the dual-channel knowledge interaction architecture including a prototype collaborative sharing channel and a parameter collaborative sharing channel;
[0009] Based on the prototype collaborative sharing channel, the interaction of local prototypes among the intelligent nodes is supported, and a global semantic alignment operation is performed on the same type of prototypes of the intelligent nodes to generate a global prototype. The feature parameters of the intelligent nodes are constrained and corrected through the global prototype.
[0010] Based on the parameter collaborative sharing channel, the interaction of feature parameters of each intelligent node after prototype correction is supported;
[0011] Each of the intelligent nodes has a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes.
[0012] Secondly, a swarm intelligence knowledge update device for resource-constrained scenarios is provided, comprising:
[0013] The building module is used to build a dual-channel knowledge interaction architecture applied to each intelligent node. The dual-channel knowledge interaction architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel.
[0014] The prototype collaboration and sharing module is used to support the interaction of local prototypes among the intelligent nodes based on the prototype collaboration and sharing channel, perform global semantic alignment operation on the same type of prototypes of the intelligent nodes to generate a global prototype, and constrain and correct the feature parameters of the intelligent nodes through the global prototype.
[0015] The parameter collaboration and sharing module is used to support the interaction of feature parameters of each intelligent node after prototype correction based on the parameter collaboration and sharing channel.
[0016] Each of the intelligent nodes has a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes.
[0017] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the swarm intelligence knowledge update method in a resource-constrained scenario as described in any one of the first aspects above.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the swarm intelligence knowledge update method in a resource-constrained scenario as described in any one of the first aspects above.
[0019] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the swarm intelligence knowledge update method in a resource-constrained scenario as described in any of the first aspects.
[0020] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0021] In this embodiment, a dual-channel knowledge interaction architecture is first constructed for each intelligent node. This architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel. Then, based on the prototype collaborative sharing channel, the interaction of local prototypes among intelligent nodes is supported. Global semantic alignment is performed on similar prototypes of each intelligent node to generate a global prototype. The feature parameters of each intelligent node are constrained and corrected using the global prototype. Finally, based on the parameter collaborative sharing channel, the interaction of feature parameters of each intelligent node after prototype correction is supported. This enables the swarm intelligence system to obtain stable and consistent prototype representations even with limited data and the inability to directly share original data. It features a robust cross-node parameter sharing mechanism, robust knowledge transfer capabilities against verification, and an efficient collaboration mode that significantly reduces communication volume. This addresses problems such as lack of effective knowledge sharing among swarm intelligence nodes, inconsistent prototype semantics, and unstable model updates under resource-constrained conditions.
[0022] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0024] Figure 1 This is a flowchart illustrating a swarm intelligence knowledge update method for resource-constrained scenarios provided in this application embodiment;
[0025] Figure 2 This is a system architecture diagram of the swarm intelligence knowledge update mechanism with dual shared channels provided in the embodiments of this application;
[0026] Figure 3 This is a structural block diagram of the swarm intelligence knowledge update device for resource-constrained scenarios provided in the embodiments of this application;
[0027] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] The embodiments of the technical solutions of this application will now be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this application, and are therefore merely examples and should not be used to limit the scope of protection of this application. When the following description relates to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.
[0029] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0030] It should be noted that with the increasing uncertainty caused by renewable energy generation and system failures, traditional deep learning methods cannot effectively handle multimodal data from power systems. Furthermore, large language models lack knowledge of the power system domain, making it difficult for traditional data-driven methods to comprehensively, effectively, and reliably assess power system stability. In addition, current autonomous power grid operations lack effective deep model selection techniques, requiring manual judgment of the current system operating conditions and selection of the corresponding model, resulting in low efficiency, insufficient automation, and low accuracy. This invention proposes a swarm intelligence knowledge update method for resource-constrained scenarios, applicable to adaptive analysis of power transient stability under various operating conditions. It accurately identifies the current operating condition and automatically selects a deep learning model to assess system stability. Based on power system domain knowledge-guided learning, this method can achieve interpretable decision result analysis while ensuring real-time computational efficiency, enabling operators to accurately understand the current operating conditions and system stability, providing reliable and accurate decision-making basis for subsequent operations. A power system domain knowledge graph is constructed based on power system knowledge, transient stability decision knowledge, and deep learning model knowledge. By constructing a data-knowledge hybrid driven large model based on a large language model and knowledge graph, we can achieve the identification of different operating conditions of the power system and the automatic model selection.
[0031] It should be noted that the execution subject of the swarm intelligence knowledge update method in the resource-constrained scenario of this embodiment can be a swarm intelligence knowledge update device in the resource-constrained scenario, hereinafter referred to as "device". The device can be configured in any type of electronic device, and this application embodiment does not limit it.
[0032] See Figure 1 This is a flowchart illustrating the swarm intelligence knowledge update method for resource-constrained scenarios provided in this application embodiment. Figure 1 As shown, the swarm intelligence knowledge update method in resource-constrained scenarios may include the following steps:
[0033] Step 101: Construct a dual-channel knowledge interaction architecture for each intelligent node. The dual-channel knowledge interaction architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel.
[0034] Each intelligent node has a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes.
[0035] Among them, intelligent nodes can be basic units in a swarm intelligence system that have independent data processing, model training and information interaction capabilities, or entities with computing and storage functions such as edge computing devices, robots, drones, federated learning clients, etc., which are the local execution entities of the knowledge update process.
[0036] Among them, the dual-channel knowledge interaction architecture can be a two-way interaction framework built to achieve efficient knowledge sharing among collective intelligence nodes. It includes two parallel and complementary interaction paths: the prototype collaborative sharing channel and the parameter collaborative sharing channel. Its core function is to replace the traditional full model parameter transmission, reduce communication overhead, and ensure the consistency of knowledge sharing.
[0037] Among them, the prototype collaborative sharing channel can be one of the channels in the dual-channel knowledge interaction architecture. It is used to carry the transmission and interaction of local prototypes and global prototypes between various intelligent nodes and is the core data transmission carrier for realizing cross-node prototype semantic alignment.
[0038] Among them, the parameter collaboration and sharing channel can be one of the channels in the dual-channel knowledge interaction architecture. It is used to carry the transmission and interaction of feature parameters of each intelligent node after prototype correction, and to provide data transmission support for subsequent adversarial loss verification and cross-node parameter collaborative optimization.
[0039] Local data sources refer to the original data sets owned locally by the intelligent nodes for model training and prototype generation. These may include images, text, sensor-collected data, behavior logs, etc., and are subject to resource constraints such as scarce data, heterogeneous distribution, or inability to be directly shared.
[0040] The local model can be a machine learning / deep learning model deployed locally on each intelligent node, including a feature encoder and a classification head, used to realize feature extraction, category discrimination or behavior prediction of local data, and is the specific carrier and execution carrier of group knowledge locally.
[0041] Among them, the local feature encoder can be the core component of the local model, which has the function of transforming the original input data (from the local data source) into high-dimensional feature vectors (i.e. feature parameters). Its output feature parameters are the core data foundation for subsequent prototype generation, parameter sharing and adversarial training.
[0042] Among them, the local category or behavior prototype set can be a set of prototypes generated by intelligent nodes based on local data sources to represent various categories (such as "target category") or behavior patterns (such as "obstacle avoidance behavior") in local data. Each prototype corresponds to the core semantic representation of a type of data and is a condensed embodiment of local knowledge.
[0043] Specifically, a swarm intelligence system containing N intelligent nodes (denoted as A1, A2, ..., AN) can be constructed. Each intelligent node is configured with a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes. The local model consists of a feature encoder and a classification head (using a 1-2 layer multilayer perceptron). The feature encoder adopts the ResNet50 architecture to transform the raw data into 2048-dimensional feature parameters. A dual-channel knowledge interaction architecture is constructed, including a prototype collaborative sharing channel and a parameter collaborative sharing channel. The channels use the TCP / IP protocol to realize data transmission.
[0044] More specifically, a swarm intelligence system containing N intelligent nodes (N=2 in this embodiment, denoted as A1 and A2) can be constructed. Each intelligent node is configured with an independent local data source, local model, local feature encoder, and local category or behavior prototype set, as follows:
[0045] Local data sources: The local data source for A1 is image data collected from indoor robot obstacle avoidance scenarios (a total of 500 images, including two categories: "obstacle avoidance successful" and "obstacle avoidance failed", of which 10 images are related to "obstacle avoidance behavior"); the local data source for A2 is image data collected from outdoor robot obstacle avoidance scenarios (a total of 400 images, including two categories of labels, with 8 images related to "obstacle avoidance behavior"). The data from the two nodes exhibit scene heterogeneity, which is consistent with the characteristic of "heterogeneous data distribution" in resource-constrained scenarios.
[0046] The local model adopts an integrated architecture of "feature encoder + classifier head". The classifier head is a 2-layer multilayer perceptron (MLP) with an input dimension of 2048 and an output dimension of 2 (corresponding to two classes of labels). The number of hidden layer neurons is 1024, and the activation function is ReLU.
[0047] Local Feature Encoder: A ResNet50 pre-trained model is selected and fine-tuned. The input is a 3-channel 64×64 pixel image, and the output is a 2048-dimensional high-dimensional feature parameter, which is used to transform the original image data into a feature representation that can be used for prototype generation and parameter sharing.
[0048] Dual-channel knowledge interaction architecture: Based on the TCP / IP protocol, a prototype collaborative sharing channel and a parameter collaborative sharing channel are constructed. The prototype collaborative sharing channel is used to transmit prototype data with a dimension of 2048; the parameter collaborative sharing channel is used to transmit corrected 2048-dimensional feature parameters. Both channels use the AES encryption algorithm to ensure data transmission security.
[0049] Global Collaboration Node: Configure an edge server to receive local prototypes from each smart node, perform semantic alignment and fusion operations, and train the discriminator.
[0050] Step 102: Based on the prototype collaborative sharing channel, support the interaction of local prototypes among intelligent nodes, perform global semantic alignment operation on similar prototypes of each intelligent node to generate a global prototype, and constrain and correct the feature parameters of each intelligent node through the global prototype.
[0051] Here, the prototype is defined as the statistical center of the feature representation in the sample set corresponding to a certain type of semantic / behavior, used to represent the low-dimensional semantic anchor of that type of semantic. The calculation formula can be as follows:
[0052]
[0053] in Indicates the client-side feature encoder i. Indicate category The sample set.
[0054] Optionally, based on the prototype collaborative sharing channel, the interaction of local prototypes between intelligent nodes is supported, and a global semantic alignment operation is performed on the same type of prototypes of each intelligent node to generate a global prototype. The feature parameters of each intelligent node are constrained and corrected through the global prototype.
[0055] Among them, the global semantic alignment operation can solve the problem of inconsistent local prototype semantics of different intelligent nodes, and perform a unified processing on the same type of local prototypes across nodes, eliminating prototype semantic deviations caused by data heterogeneity and model differences among nodes.
[0056] Among them, the global prototype can be a unified semantic representation unit corresponding to a certain data category or behavior pattern, generated through global semantic alignment operation. It is the result of alignment and fusion of local prototypes of the same type across nodes, and can be used as a unified benchmark for the correction of feature parameters of each intelligent node.
[0057] Among them, the feature parameters can be high-dimensional feature vectors output by the local feature encoder after extracting features from the original input data. They are abstract feature representations of the original data and core data objects for prototype generation, parameter sharing, adversarial training, and feature correction.
[0058] Among them, feature parameter constraint correction can be a process of adjusting and optimizing the feature parameters of each intelligent node based on the global prototype. Specifically, it is achieved by directly accumulating the perturbation of the global prototype onto the feature parameters. The core purpose is to align the feature parameters of each node to the global semantic space and reduce cross-node feature differences.
[0059] Optionally, the step of performing a global semantic alignment operation on similar prototypes of each intelligent node to generate a global prototype specifically includes:
[0060] Through the prototype collaboration and sharing channel, the local prototypes of each intelligent node are sent to the global collaboration node. The global collaboration node classifies and collects the received local prototypes according to data category or behavior pattern to obtain a cross-node set of similar prototypes.
[0061] Perform alignment and fusion operations on each type of prototype set to generate a global prototype for the corresponding category;
[0062] The generated global prototypes are distributed to each intelligent node through the prototype collaboration and sharing channel.
[0063] Among them, the global collaboration node is a central node or collaboration node with the ability to collect, process and distribute cross-node data. Its core responsibilities include receiving local prototypes sent by each intelligent node, performing classification and collection of similar prototypes, generating global prototypes and distributing global prototypes in reverse to each intelligent node. It is the core execution entity for global semantic alignment operations.
[0064] Among them, the alignment and fusion operation is the core operation of the global collaborative node to uniformly process the set of similar prototypes across nodes to generate a global prototype. Specifically, it adapts different processing logic according to the generation method of local prototypes (aggregation operation / cluster analysis) to ensure that the generated global prototype can accurately integrate the core semantics of the local prototypes of each node.
[0065] Optionally, perform an alignment and fusion operation for each type of prototype set to generate a global prototype for the corresponding category, including:
[0066] If the local prototype is generated through aggregation operations, then the summation and averaging of all local prototypes in the same prototype set are performed to obtain the global prototype of the same prototype set.
[0067] or,
[0068] If the local prototype is generated through cluster analysis, then the clustering operation is performed again on all local prototypes in the same prototype set, and the cluster center obtained by clustering is used as the global prototype of the same prototype set.
[0069] Among them, aggregation operation is a data processing method for generating local prototypes or fusing global prototypes. Specifically, it is an operation that sums up the features of similar samples or similar local prototypes and then takes the average. The core features of similar data are extracted by statistical averaging to form the corresponding prototype.
[0070] Cluster analysis is another data processing method for generating local prototypes or fusing global prototypes. It clusters similar features or prototypes into a class by calculating the distance between features, and uses the cluster center obtained by clustering as the corresponding prototype. It is suitable for scenarios with complex data distribution.
[0071] Among them, feature difference is the difference in feature vectors between different local prototypes. It is usually calculated by metrics such as Euclidean distance and cosine similarity. It is used to evaluate the semantic consistency of local prototypes across nodes and is an important basis for identifying abnormal local prototypes and judging the effect of global semantic alignment.
[0072] Optionally, before performing the alignment and fusion operation for each set of prototypes to generate a global prototype for the corresponding category, the following steps are also included:
[0073] Calculate the feature differences between local prototypes within the same prototype set, and identify and remove abnormal local prototypes whose differences exceed the target value.
[0074] Optionally, the local feature encoder of the smart node extracts features of similar samples from the local data source, and performs at least one operation, such as aggregation or clustering analysis, on the extracted features of similar samples to obtain the feature statistical center of the corresponding data category or behavior pattern as a local prototype.
[0075] Among them, abnormal local prototypes are local prototypes in the cross-node prototype set whose features differ from other local prototypes by a preset target value. Their generation is usually affected by local data noise, abnormal data distribution, or model bias. Removing such prototypes can improve the reliability and accuracy of the global prototype.
[0076] Among them, the features of similar samples are the set of feature parameters obtained by the local feature encoder after the samples in the local data source belong to the same data category or behavior pattern. They are the direct data basis for the intelligent node to generate local prototypes, and their quality directly affects the accuracy of the local prototypes.
[0077] The target value is a threshold used to determine whether a local prototype is an abnormal prototype. It is preset by the user or the system according to the specific application scenario, data characteristics and accuracy requirements. When the feature difference between two local prototypes is greater than this value, one or both of them are determined to be an abnormal local prototype.
[0078] In cluster analysis, the cluster center is the central location vector of all feature vectors (or local prototypes) within the same cluster. It is the core feature representation of the data within the cluster and can be used as a local prototype or a global prototype.
[0079] Specifically, each intelligent node extracts 2048-dimensional features of similar samples from the local data source through a local feature encoder. For the "obstacle avoidance behavior" category, intelligent node A1 extracts 10 similar sample features and uses an aggregation operation of summation and averaging to obtain the local prototype P1 of this category; intelligent node A2 extracts 8 similar sample features and uses K-Means clustering analysis (Euclidean distance is the distance metric) to obtain the local prototype P2 of this category.
[0080] In this process, A1 and A2 send P1 and P2 to the global collaboration node through the prototype collaboration sharing channel. The global collaboration node groups P1 and P2 according to the "obstacle avoidance behavior" category to obtain a prototype set {P1, P2} of the same type. The Euclidean distance between P1 and P2 is calculated to be 0.3, and the preset target value is 0.5. Since 0.3 < 0.5, it is determined that there are no abnormal local prototypes. P1 is generated through aggregation operation, and P2 is generated through cluster analysis. Using the fusion strategy of "re-clustering", K-Means clustering is performed on {P1, P2} to obtain the cluster center P_global as the global prototype of the "obstacle avoidance behavior" category. P_global has a dimension of 2048. The global collaboration node distributes P_global back to A1 and A2 through the prototype collaboration sharing channel.
[0081] More specifically, A1 sends P1 to the global collaboration node through the prototype collaboration sharing channel. During transmission, a data compression algorithm is used to compress the data volume. A2 sends P2 to the global collaboration node in the same way. After receiving P1 and P2, the global collaboration node parses the category labels of each prototype and, according to the "data category / behavioral pattern" grouping rule, groups P1 and P2 into the same prototype set, denoted as S={P1,P2}. The global collaboration node uses Euclidean distance to calculate the feature difference between P1 and P2 within the same prototype set S, for example, if the calculation result is 0.3. Considering the obstacle avoidance scenario requirements of this embodiment, the preset anomaly judgment target value is 0.5; since 0.3 < 0.5, P1 and P2 are determined to be normal local prototypes and do not need to be removed. The prototype set S={P1,P2} is retained for subsequent fusion operations.
[0082] Since P1 in the prototype set S is generated through aggregation and P2 is generated through cluster analysis, this embodiment adopts a "re-clustering" alignment and fusion strategy. Specifically, it uses the same K-Means algorithm as the local prototype generation for A2, with K=1 clusters, Euclidean distance as the distance metric, 50 iterations, and a convergence threshold of 1e-6 (iteration stops when the change in cluster centers between two iterations is less than 1e-6). P1 and P2 are input into the K-Means clustering model, and convergence is achieved after 32 iterations, yielding the cluster centers, which are the global prototype P_global for the "obstacle avoidance behavior" category. P_global is a 2048-dimensional vector, consistent with the dimensions of the local prototypes and feature parameters. The Euclidean distances between P_global and P1 and P2 are calculated to be 0.12 and 0.15, respectively, both less than the distance of 0.3 between P1 and P2 before fusion. This demonstrates that the fused global prototype achieves cross-node semantic alignment, reducing the semantic deviation between nodes. The global collaborative node distributes P_global to A1 and A2 respectively through the prototype collaborative sharing channel. After receiving P_global, each intelligent node decompresses and verifies it (the verification method is to calculate the hash value of the received data and the sent data to ensure the integrity of data transmission). After the verification is passed, it is stored locally as a unified semantic benchmark for subsequent feature parameter constraint correction.
[0083] Step 103: Based on the parameter collaboration and sharing channel, support the interaction of feature parameters of each intelligent node after prototype correction.
[0084] Optionally, local perturbations can be added to the feature parameters after prototype correction. The local perturbation adopts random Gaussian perturbation. The feature parameters of each intelligent node after local perturbation are collected. The real client ID corresponding to each feature parameter is used as the label. KL divergence is used as the loss function to train the discriminator. Then, the trained discriminator is distributed to each intelligent node. The discriminator parameters are frozen during the local model training of each intelligent node. Then, the real feature parameters of each intelligent node are input into the discriminator with frozen parameters to obtain the predicted client source. The KL divergence between the predicted client source and the uniform distribution vector is calculated as the adversarial loss.
[0085] Optionally, the global prototype can be included in the perturbation process of the feature parameters of each intelligent node. In this case, the feature parameter correction is achieved by directly accumulating the global prototype onto the feature parameters. The dimensions of the global prototype and the feature parameters are kept consistent.
[0086] In this process, A1 and A2 directly add P_global to the 2048-dimensional feature parameters output by their respective local feature encoders to achieve constraint correction of the feature parameters, resulting in the corrected feature parameters F1 and F2.
[0087] Specifically, random Gaussian perturbations (variance set to 0.1) can be added to F1 and F2 to obtain perturbed feature parameters F1' and F2'. Global collaborative nodes collect F1' and F2', using the real client IDs of A1 and A2 as labels, and KL divergence as the loss function to train the discriminator D. The trained D is distributed to A1 and A2, and its parameters are frozen. A1 inputs the real feature parameters F1 into D to obtain the predicted client source, calculates the KL divergence between this predicted source and a uniformly distributed vector (dimensional 2) as the adversarial loss, and uses it to train the local feature encoder, forcing F1 to generate client-independent information. A2 performs the same operation. A1 and A2 achieve cross-node parameter collaboration by sharing the corrected feature parameters F1 and F2 through parameter collaboration channels.
[0088] Specifically, the device also includes a group consistency assessment module for evaluating global state drift, actively triggering corrections, and learning global knowledge. The main evaluation metric for global state drift is the prototype consistency metric, i.e., the prototype differences across nodes:
[0089]
[0090] if If the value exceeds the target threshold, retraining of both the global classifier and discriminator is triggered, and the next round of local model training begins. The local model's performance can be evaluated using top-1 classification accuracy and the infoNCE loss between the generated features and the global prototype.
[0091] In this embodiment, a dual-channel knowledge interaction architecture is first constructed for each intelligent node. This architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel. Then, based on the prototype collaborative sharing channel, the interaction of local prototypes between intelligent nodes is supported. Global semantic alignment is performed on similar prototypes of each intelligent node to generate a global prototype. The feature parameters of each intelligent node are constrained and corrected using the global prototype. Subsequently, based on the parameter collaborative sharing channel, the interaction of feature parameters of each intelligent node after prototype correction is supported. Each intelligent node possesses a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes. This enables the swarm intelligence system to obtain stable and consistent prototype representations even with limited data and the inability to directly share original data. It features a robust cross-node parameter sharing mechanism, robust knowledge transfer capabilities against verification, and an efficient collaboration mode that significantly reduces communication volume. This addresses problems such as lack of effective knowledge sharing among swarm intelligence nodes, inconsistent prototype semantics, and unstable model updates under resource-constrained conditions.
[0092] Figure 2This diagram illustrates the interaction pattern between agent A and agent B, which is primarily composed of two parts: a "cooperation channel" and "internal modules." The diagram shows two independent shared channels responsible for transmitting different types of information: a prototype cooperation sharing channel and a feature cooperation sharing channel. Each agent (A and B) contains two modules: a consensus module and an adversarial module.
[0093] The beneficial effects of the embodiments of this application are as follows:
[0094] Adversarial collaboration makes model updates between nodes less susceptible to noise and data imbalance. Compared to sending complete model parameters, it significantly reduces the number of communication parameters, with communication costs depending only on the feature dimension and data sample size. The number of local model parameters depends on the sum of the parameters of the local model's feature encoder and classifier. It features high prototype consistency, unified semantics across nodes, and the ability to form a unified knowledge structure for behaviors, classes, and parameters in swarm intelligence. It boasts strong scalability and is suitable for various scenarios such as robot swarms, drone swarms, edge task collaboration, and federated large models.
[0095] Corresponding to the swarm intelligence knowledge update method in resource-constrained scenarios described in the above embodiments, Figure 3 This is a structural block diagram of the swarm intelligence knowledge update device for resource-constrained scenarios provided in the embodiments of this application.
[0096] Reference Figure 3 The swarm intelligence knowledge update device 200 for resource-constrained scenarios includes:
[0097] Module 210 is used to build a dual-channel knowledge interaction architecture applied to each intelligent node. The dual-channel knowledge interaction architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel.
[0098] The prototype collaboration and sharing module 220 is used to support the interaction of local prototypes among the intelligent nodes based on the prototype collaboration and sharing channel, perform global semantic alignment operation on the same type of prototypes of the intelligent nodes to generate a global prototype, and constrain and correct the feature parameters of the intelligent nodes through the global prototype.
[0099] The parameter collaboration and sharing module 230 is used to support the interaction of feature parameters of each intelligent node after prototype correction based on the parameter collaboration and sharing channel.
[0100] Each of the intelligent nodes has a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes.
[0101] Optional, a prototype collaboration and sharing module, specifically used for:
[0102] Through the prototype collaboration sharing channel, the local prototypes of each intelligent node are sent to the global collaboration node. The global collaboration node classifies and collects the received local prototypes according to data category or behavior pattern to obtain a cross-node set of similar prototypes.
[0103] An alignment and fusion operation is performed on each set of prototypes to generate a global prototype for the corresponding category;
[0104] The generated global prototypes are distributed in reverse to each intelligent node through the prototype collaboration and sharing channel.
[0105] Optional, a prototype collaboration and sharing module, specifically used for:
[0106] If the local prototype is generated through aggregation operation, then the summation and averaging of all local prototypes in the same prototype set are performed to obtain the global prototype of the same prototype set.
[0107] or,
[0108] If the local prototype is generated through cluster analysis, then the clustering operation is performed again on all local prototypes in the same prototype set, and the cluster center obtained by clustering is used as the global prototype of the same prototype set.
[0109] Optional, a prototype collaboration and sharing module, specifically used for:
[0110] Calculate the feature differences between local prototypes within the same prototype set, and identify and remove abnormal local prototypes whose differences are greater than the target value.
[0111] Optional, a prototype collaboration and sharing module, specifically used for:
[0112] The local feature encoder of the intelligent node extracts features of similar samples from the local data source;
[0113] Perform at least one operation, such as aggregation or cluster analysis, on the extracted features of the same type of samples to obtain the feature statistical centers of the corresponding data category or behavior pattern, which serve as the local prototype.
[0114] Optional, parameter collaboration and sharing module, specifically used for:
[0115] Local perturbations are added to the feature parameters after prototype correction, and the local perturbations are random Gaussian perturbations;
[0116] Collect the feature parameters of each intelligent node after local perturbation, use the real client ID corresponding to each feature parameter as a label, and use KL divergence as a loss function to train a discriminator.
[0117] The trained discriminator is distributed to each of the intelligent nodes, and the discriminator parameters are frozen during the local model training process of each of the intelligent nodes.
[0118] The true feature parameters of each intelligent node are input into the discriminator of the frozen parameters to obtain the predicted client source;
[0119] The KL divergence between the predicted client source and the uniform distribution vector is calculated as the adversarial loss.
[0120] Optional, parameter collaboration and sharing module, specifically used for:
[0121] The global prototype is involved in the perturbation process of the feature parameters of each intelligent node, wherein the feature parameter correction is achieved by directly accumulating the global prototype onto the feature parameters; wherein the dimension of the global prototype is consistent with that of the feature parameters.
[0122] Optionally, the device is applied to resource-constrained scenarios, including scenarios where the amount of local data on intelligent nodes is scarce, scenarios where data distribution is heterogeneous, or scenarios where raw data cannot be directly shared.
[0123] The device can also be used in at least one of the following scenarios: robot swarm collaboration, drone swarm control, edge computing node task collaboration, and federated large model training.
[0124] In this embodiment, a dual-channel knowledge interaction architecture is first constructed for each intelligent node. This architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel. Then, based on the prototype collaborative sharing channel, the interaction of local prototypes between intelligent nodes is supported. Global semantic alignment is performed on similar prototypes of each intelligent node to generate a global prototype. The feature parameters of each intelligent node are constrained and corrected using the global prototype. Subsequently, based on the parameter collaborative sharing channel, the interaction of feature parameters of each intelligent node after prototype correction is supported. Each intelligent node possesses a local data source, a local model, a local feature encoder, and a local set of category or behavior prototypes. This enables the swarm intelligence system to obtain stable and consistent prototype representations even with limited data and the inability to directly share original data. It features a robust cross-node parameter sharing mechanism, robust knowledge transfer capabilities against verification, and an efficient collaboration mode that significantly reduces communication volume. This addresses problems such as lack of effective knowledge sharing among swarm intelligence nodes, inconsistent prototype semantics, and unstable model updates under resource-constrained conditions.
[0125] in addition, Figure 3The swarm intelligence knowledge update device shown in the resource-constrained scenario can be a software unit, hardware unit, or a combination of software and hardware built into an existing electronic device, or it can be integrated into the electronic device as an independent component, or it can exist as an independent electronic device.
[0126] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0127] Figure 4 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 4 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 4 Only one is shown in the diagram), memory 51, and computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, it implements the steps in the above embodiments of the swarm intelligence knowledge update method under any of the resource-constrained scenarios.
[0128] The electronic device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 5 and does not constitute a limitation on electronic device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0129] The processor 50 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0130] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may be an external storage device of the electronic device 5, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., equipped on the electronic device 5. Further, the memory 51 may include both internal storage units and external storage devices of the electronic device 5. The memory 51 is used to store operating systems, applications, boot loaders, data, and other programs, such as the program code of the computer program. The memory 51 can also be used to temporarily store data that has been output or will be output.
[0131] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.
[0132] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.
[0133] If the integrated unit is implemented as a software functional unit and used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0134] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0135] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0136] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0137] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0138] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for updating swarm intelligence knowledge in resource-constrained scenarios, characterized in that, It is applied to scenarios such as robot swarm collaboration, drone swarm control, edge computing node task collaboration, or federated large model training, including: A dual-channel knowledge interaction architecture is constructed for each intelligent node. The dual-channel knowledge interaction architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel. The intelligent node is an edge computing device, robot, or drone with independent data processing, model training, and information interaction capabilities. The dual-channel knowledge interaction architecture is built based on the TCP / IP protocol. Based on the prototype collaborative sharing channel, the interaction of local prototypes between the intelligent nodes is supported, and a global semantic alignment operation is performed on the same type of prototypes of the intelligent nodes to generate a global prototype. The feature parameters of the intelligent nodes are constrained and corrected by the global prototype. The feature parameters are high-dimensional feature vectors extracted from the original perceptual data by the local feature encoder of the intelligent node. The local feature encoder adopts the ResNet50 architecture. Based on the parameter collaborative sharing channel, the interaction of feature parameters of each intelligent node after prototype correction is supported; Each of the intelligent nodes has a local data source, a local model, a local feature encoder, and a local category or behavior prototype set. The local data source is the image, sensor data, or behavior log data collected by the intelligent node. The local model consists of a feature encoder and a classification head, and the classification head adopts a multilayer perceptron structure.
2. The method according to claim 1, characterized in that, The step of performing a global semantic alignment operation on the similar prototypes of each intelligent node to generate a global prototype specifically includes: Through the prototype collaboration sharing channel, the local prototypes of each intelligent node are sent to the global collaboration node. The global collaboration node classifies and collects the received local prototypes according to data category or behavior pattern to obtain a cross-node set of similar prototypes. An alignment and fusion operation is performed on each set of prototypes to generate a global prototype for the corresponding category; The generated global prototypes are distributed in reverse to each intelligent node through the prototype collaboration and sharing channel.
3. The method according to claim 2, characterized in that, The step of performing an alignment and fusion operation on each type of prototype set to generate a global prototype for the corresponding type includes: If the local prototype is generated through aggregation operation, then the summation and averaging of all local prototypes in the same prototype set are performed to obtain the global prototype of the same prototype set. or, If the local prototype is generated through cluster analysis, then the clustering operation is performed again on all local prototypes in the same prototype set, and the cluster center obtained by clustering is used as the global prototype of the same prototype set.
4. The method according to claim 3, characterized in that, Before performing the alignment and fusion operation for each type of prototype set to generate a global prototype for the corresponding category, the method further includes: Calculate the feature differences between local prototypes within the same prototype set, and identify and remove abnormal local prototypes whose differences are greater than the target value.
5. The method according to claim 2, characterized in that, Also includes: The local feature encoder of the intelligent node extracts features of similar samples from the local data source; Perform at least one operation, such as aggregation or cluster analysis, on the extracted features of the same type of samples to obtain the feature statistical centers of the corresponding data category or behavior pattern, which serve as the local prototype.
6. The method according to claim 1, characterized in that, Also includes: Local perturbations are added to the feature parameters after prototype correction, and the local perturbations are random Gaussian perturbations; Collect the feature parameters of each intelligent node after local perturbation, use the real client ID corresponding to each feature parameter as a label, and use KL divergence as a loss function to train a discriminator. The trained discriminator is distributed to each of the intelligent nodes, and the discriminator parameters are frozen during the local model training process of each of the intelligent nodes. The true feature parameters of each intelligent node are input into the discriminator of the frozen parameters to obtain the predicted client source; The KL divergence between the predicted client source and the uniform distribution vector is calculated as the adversarial loss.
7. The method according to claim 1, characterized in that, The step of constraining and correcting the feature parameters of each intelligent node using the global prototype specifically includes: The global prototype is involved in the perturbation process of the feature parameters of each intelligent node, wherein the feature parameter correction is achieved by directly accumulating the global prototype onto the feature parameters; wherein the dimension of the global prototype is consistent with that of the feature parameters.
8. The method according to any one of claims 1-7, characterized in that, The method is applied to resource-constrained scenarios, including scenarios where the amount of local data on intelligent nodes is scarce, scenarios where data distribution is heterogeneous, or scenarios where raw data cannot be directly shared. The application scenarios of the method also include at least one of robot swarm collaboration, drone swarm control, edge computing node task collaboration, and federated large model training.
9. A swarm intelligence knowledge update device for resource-constrained scenarios, characterized in that, It is applied to scenarios such as robot swarm collaboration, drone swarm control, edge computing node task collaboration, or federated large model training, including: The module is used to build a dual-channel knowledge interaction architecture for each intelligent node. The dual-channel knowledge interaction architecture includes a prototype collaborative sharing channel and a parameter collaborative sharing channel. The intelligent node is an edge computing device, robot or drone with independent data processing, model training and information interaction capabilities. The dual-channel knowledge interaction architecture is built based on the TCP / IP protocol. The prototype collaboration and sharing module is used to support the interaction of local prototypes among the intelligent nodes based on the prototype collaboration and sharing channel, and to perform global semantic alignment operation on the same type of prototypes of the intelligent nodes to generate a global prototype. The global prototype is used to constrain and correct the feature parameters of the intelligent nodes. The feature parameters are high-dimensional feature vectors extracted from the original perceptual data by the local feature encoder of the intelligent node. The local feature encoder adopts the ResNet50 architecture. The parameter collaboration and sharing module is used to support the interaction of feature parameters of each intelligent node after prototype correction based on the parameter collaboration and sharing channel. Each of the intelligent nodes has a local data source, a local model, a local feature encoder, and a local category or behavior prototype set. The local data source is the image, sensor data, or behavior log data collected by the intelligent node. The local model consists of a feature encoder and a classification head, and the classification head adopts a multilayer perceptron structure.
10. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Interactive intelligent question and answer method based on power grid practical training question and answer knowledge base
CN114417880A
Intelligent knowledge base management method and system based on AI
CN120671796A