Federal learning method for tactical edge intelligence
Through the federated learning method of building a data layer of multimodal local data in a tactical edge environment, the model layer of deploying available neural network models, and the application layer for diversified tasks, the problems of insufficient communication bandwidth, low terminal device reliability, and multi-tasking and multi-modal data sharing in tactical edge environments are solved, and efficient communication and data sharing are achieved, meeting the needs of tactical edge intelligence.
Patent Information
- Application Number
- CN202510069479.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-06
AI Technical Summary
In a tactical edge environment, federated learning faces the difficulties of insufficient communication bandwidth, low terminal device reliability, and multi-tasking and multi-modal data sharing, resulting in increased communication delays, deterioration in learning quality and low data utilization efficiency.
A federated learning method for tactical edge intelligence is proposed, and a three-layer architecture is formed by building a data layer of multimodal local data, a model layer for deploying available neural network models, and an application layer for diversified tasks. This method performs federated cluster division and data retention in the data layer, the model layer conducts local neural network model training and parameter aggregation, and the application layer adopts asynchronous hierarchical federated learning strategy, combining the Encoder-Decoder form and the weighted network hybrid structure to realize multimodal and multitasking learning.
Through local training and upper-level aggregation model, the number of communications is reduced, communication efficiency is improved, the problems of poor communication quality and network discontinuity are avoided, the security and efficiency of data sharing are enhanced, and the multi-tasking and multi-modal data sharing needs are met in tactical edge environments.
Smart Images

Figure CN120106180A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network communication and machine learning, and specifically relates to a federated learning method for tactical edge intelligence. Background Art
[0002] The tactical edge is a very dynamic environment, that is, the terminal position status is constantly changing, the frequency of equipment use changes frequently, the network connectivity is intermittent, the information flow bursts frequently, and the schedule and task plan are constantly adjusted. At the same time, with the continuous development of modern warfare technology, more efficiency, simplification, and flattening are required. In order to meet the needs of modern warfare informationization and tactical edge environment, how to efficiently provide agile, flexible, and intelligent information services for tactical forces is an important issue to be solved in the development of the tactical edge.
[0003] With the advent of the Internet of Things and the 5G era, the scale of the network edge device layer has expanded rapidly, and the amount of data has increased dramatically, which has put forward higher bandwidth requirements for data transmission. At the same time, application scenarios such as virtual reality, smart homes, and smart connected cars have also put forward higher requirements for the real-time performance of data processing. The centralized processing based on the traditional cloud computing model can no longer effectively solve the real-time requirements and insufficient bandwidth, energy consumption, data security and privacy issues. Therefore, edge computing came into being.
[0004] At the same time, smart terminal devices have gradually become an indispensable part of people's lives. In the development and evolution of devices, people have put forward further requirements for service quality. In this case, the processing mode represented by centralized cloud computing will not be able to efficiently process the data generated by edge devices, and most of the current artificial intelligence computing tasks are deployed on platforms with large-scale computing resources such as cloud computing centers, which greatly limits the convenience brought by artificial intelligence to people and cannot meet people's demand for service quality. Therefore, delegating intelligence to the level close to users has become an urgent problem to be solved, so the concept of edge intelligence is about to emerge.
[0005] Edge intelligence refers to terminal intelligence, an open platform that integrates the core capabilities of network, computing, storage, and applications, and provides edge intelligence services to meet the key needs of industry digitalization in terms of agile connection, real-time business, data optimization, application intelligence, security and privacy protection, etc. Deploying intelligence on edge devices can bring intelligence closer to users, provide users with intelligent services faster and better, and improve user experience.
[0006] Federated learning mainly performs shared machine learning and multi-party secure computing on financial data islands. However, there are still several problems with federated learning in tactical edge scenarios:
[0007] First, there is the problem of communication bandwidth. Although the edge device hardware used in the tactical edge is constantly upgraded, most mobile phones and other edge devices are gradually equipped with AI chips, and computing power is constantly increasing, the wireless communication bandwidth is seriously insufficient. Although 5G has improved communication conditions to a certain extent, the communication bandwidth is limited in most areas, and the problem of uneven distribution still exists. Limited communication bandwidth will lead to increased communication delay, which will directly increase the convergence time of federated learning.
[0008] Secondly, it is the reliability of terminal devices. Federated learning is an iterative process, and the terminal devices involved in the process must continue to communicate until the learning process is completed. However, in actual applications, due to various factors such as the complex tactical edge environment and bad weather, some terminals may exit the process, which results in the inability of federated learning to fully utilize data, resulting in the continuous deterioration of the quality of federated learning.
[0009] The tactical edge accumulates a large amount of multi-task data, and the demand for personalized data services continues to increase. The differences in high-quality data from different terminals are becoming more and more obvious. Therefore, under the premise of ensuring data sharing security, how to enable multiple participants in the tactical edge to achieve multi-task and multi-modal data sharing, give terminals or edge nodes dynamic autonomy, and form a win-win tactical edge ecosystem is a very important issue. Summary of the invention
[0010] The purpose of the present invention is to provide a federated learning method for tactical edge intelligence, which improves communication efficiency and effectively avoids the problems of poor communication quality and network intermittent.
[0011] The present invention is achieved through the following technical solutions:
[0012] A federated learning approach for tactical edge intelligence, building a three-layer architecture of tactical edge intelligence, including a data layer covering multimodal local data, a model layer of available neural network models deployed at the tactical edge, and a tactical edge application layer for diverse tasks;
[0013] The data layer conducts a "federation" union of tactical edge computing nodes, divides cluster partitions, strengthens data retention locally, and clusters and scores based on the amount of multimodal task data and the computing resources and network resources of tactical edge computing nodes to form a customized federation cluster partitioning strategy;
[0014] The model layer trains the local neural network model based on the multimodal data of tactical edge nodes in different partitions. is the local data training set, where is the number of training samples, x j is the jth input training sample, y jis the label of the corresponding sample, and the vector w is defined as the full model parameter vector, and f(x j ,y j , w) is a simplified representation of the loss function, which specifically means the loss function of the jth sample; the goal of the training process is to minimize the average loss F(w), which is specifically defined as follows:
[0015]
[0016] After the local model is trained, the model layer randomly samples the model parameters and customizes the local model parameters to be aggregated and uploaded to the upper-layer server, forming a distributed federated aggregation framework for differentiated data in different partitions until the model converges globally.
[0017] The application layer calls the asynchronous hierarchical federated learning strategy of the model layer, providing a global model upward and a customized and personalized neural network model for edge node adaptation downward. It focuses on adapting the close application of downstream tasks and local data to form customized downstream task auxiliary decision-making capabilities, and builds a complete tactical edge intelligence framework of local data partitioning-distributed model training aggregation-customized and personalized global applications.
[0018] Furthermore, the data layer performs clustering and scoring based on the amount of multimodal task data and the computing resources and network resources of the tactical edge computing nodes. For customized cluster i, the amount of multimodal data is vi, the computing resources are ci, and the network resources are ni. The cluster score is si=x*vi+y*ci+z*ni, where the coefficient weight satisfies x+y+z=1. It is dynamically adjusted according to the data, computing, and network resources of the tactical edge environment to meet the convergence requirements of downstream customized tasks and upstream general models.
[0019] Furthermore, the asynchronous hierarchical federated learning strategy is based on the cloud-edge-end hierarchical architecture of the tactical edge. In the asynchronous hierarchical federated learning strategy, data is iterated locally for a specified number of rounds from the bottom device, and then the preliminary model parameters are uploaded to the edge for aggregation. After the edge completes the model aggregation and performs complex data reinforcement training, the model is uploaded to the cloud. The cloud finally aggregates the data to form a global model, and broadcasts the global model within the network to ensure that each node can receive the model parameter push and perform local model updates.
[0020] With the time-series asynchronous federated learning algorithm as the core, the time-series decay contribution method is used to allocate parameter weights. That is, the model with a longer last update time has a larger update weight ratio than the model with a shorter update time, ensuring that a parameter model close to the real data change characteristics is fitted while taking into account both historical information and new information.
[0021] Furthermore, in the asynchronous hierarchical federated learning strategy, the Encoder-Decoder form and the weighted network hybrid structure are specifically used to achieve multi-modality and multi-task;
[0022] During the training process, the Epoch timing method is used to discretize the continuous time. The main process uses the K variable for counting. The threads use τ and t, which are the global transfer timestamp and local timestamp for synchronization respectively. Through the gradient descent method, F(w) in formula (1) gradually approaches the minimum value, and finally w is the optimal model parameter. k is the index of the number of updates, and η is used as the scale of the update gradient. The model update formula based on the gradient descent method is as follows:
[0023]
[0024] In the asynchronous hierarchical federated learning strategy, the data set is distributed on N computing terminals, and the data set of each terminal is and And the gradient update is divided into local update and aggregate update, and the functions of the two processes are F(w) and F i (w), the definitions of the two functions are as follows:
[0025]
[0026] In the layering of the asynchronous hierarchical federated learning strategy, it is assumed that there are L edge devices in the edge layer, and e is the device index. Each edge device has C l Terminal node, let k 1 Update the index for the terminal device layer and let κ 2 Update the index for the edge layer. If the total index k can be 1 If it is divisible by , then the local update is terminated, edge aggregation is started, and κ is reset 1 ; If the total index k can be 1 κ 2 If it is divisible by , then the edge gradient update is terminated, cloud aggregation is started, and κ is reset at the same time 1 and κ 2 , start counting again from 1;
[0027] The gradient of the device No. i belonging to the edge region No. l is specified as At this point, based on the above description, the following definitions are available:
[0028]
[0029] At this time, according to the asynchronous hierarchical federated learning strategy, the number of training rounds is used to define the time. Set t∈T, that is, the total number of training rounds is T, the current number of training rounds is t, the number of training rounds where the sample is submitted is τ, and the attenuation function is set to s(·), then α t ←α×s(t-τ), where α t is the decay rate at time t;
[0030] At this time, set w t is the existing aggregation model, w t-1 is the aggregation model of the previous Epoch, w new For newly uploaded models, there is w t ←(1-α t ) t-1 +α t w new ; Timestamp each sample data, and convey the time according to the data gradient during aggregation, as well as the overall training effect to achieve attenuation.
[0031] The main innovative features of the present invention include the following three points:
[0032] 1. Aiming at the characteristics of tactical edge network intermittent, high security requirements and limited communication resources, the present invention proposes a data layer covering multimodal local data, a model layer of available neural network models suitable for deployment at the tactical edge, and a tactical edge application layer for diversified tasks, forming a three-layer architecture of tactical edge intelligence data layer, model layer and application layer. The computing nodes are "federated" and divided into cluster partitions. The data layer strengthens the retention of data locally to adapt to network intermittent and security requirements. The model layer trains and schedules local neural network models for multimodal data of nodes in different partitions. The upper server is responsible for customized aggregation of locally uploaded model parameters to form a distributed federation aggregation framework for differentiated data in different partitions. The application layer provides a unified global model with strong security and low communication cost to the upper layer, and provides a customized personalized neural network model for edge node adaptation to the lower layer, highlighting the close application combination with local tasks, forming a customized downstream task auxiliary decision-making capability, and building a complete tactical edge intelligence framework of local data partitioning-distributed model training aggregation-customized personalized global application.
[0033] 2. Aiming at the multi-task and multimodal learning needs of the tactical edge, a customized federated cluster division strategy is proposed for clustering and scoring based on the amount of multimodal task data and the computing resources and network resources of the tactical edge computing nodes; in view of the intermittent characteristics of the tactical edge network and limited communication resources, an asynchronous mode is designed to form a training strategy for the local model and the global model of asynchronous hierarchical federated learning, and the Encoder-Decoder form and the weighted network hybrid structure are used to achieve multimodal and multi-task, thereby highlighting the downstream task customization capability of the local model and data of the tactical edge node, and the resource-saving performance of the global model, effectively enabling multi-dimensional perception of the tactical edge environment situation.
[0034] 3. In an environment where the tactical edge network is intermittent, security requirements are high, and communication resources are limited, the computing nodes are "federated". The multimodal and multi-task data of the nodes are retained locally, and the upper-level server is responsible for customized aggregation of the locally uploaded processing information. Ultimately, a usable neural network model with strong security and low communication cost is formed, which is suitable for deployment at the tactical edge. Combined with the application layer, it forms a tactical edge intelligent framework to assist in task decision-making.
[0035] Federated learning greatly reduces the number of communications and improves communication efficiency through the mode of local training + upper-level aggregation model + global push update, thus effectively avoiding the problems of poor communication quality and network interruption. At the same time, by moving the training tasks downward, the number of unnecessary data transfers in the system architecture is reduced, solving the problem of using big data in tactical edge environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Tactical Edge Federated Intelligence Framework
[0037] Figure 2 Tactical Edge System Deployment Architecture
[0038] Figure 3 Asynchronous hierarchical federated learning strategy
[0039] Figure 4 Schematic diagram of weighted hybrid network
[0040] Figure 5 Encoder-Decoder Structure
[0041] Figure 6 Asynchronous hierarchical federated learning strategy architecture example
[0042] Figure 7 Related case data
[0043] Figure 8 Illustration of the relationship between image elements and concepts
[0044] Fig. 9Colab virtual machine configuration diagram
[0045] Fig.10 CLIP API Diagram
[0046] Fig.11 Unplash Dataset scale diagram
[0047] Fig.12 Illustration of multimodal joint retrieval results DETAILED DESCRIPTION
[0048] The present invention is further described in detail below in conjunction with specific embodiments, which are intended to explain the present invention rather than to limit it.
[0049] In view of the intermittent characteristics of tactical edge networks, high security requirements, and limited communication resources, the present invention proposes a data layer covering multimodal local data, a model layer suitable for deploying available neural network models at the tactical edge, and a tactical edge application layer for diversified tasks, forming a three-layer architecture of tactical edge intelligence: data layer, model layer, and application layer.
[0050] (1) At the data layer, the tactical edge computing nodes are federated, divided into cluster partitions, and data is retained locally to adapt to network intermittency and security requirements. Clustering and scoring are performed based on the amount of multimodal task data and the computing resources and network resources of the tactical edge computing nodes to form a customized federated cluster partitioning strategy;
[0051] The scoring formula is designed as follows:
[0052] For customized cluster i, the amount of multimodal data is vi, the computing resources are ci, the network resources are ni, and the cluster score is
[0053] si=x*vi+y*ci+z*ni, where the coefficient weight satisfies x+y+z=1, and can be dynamically adjusted according to the tactical edge environment data, computing, and network resources to meet the downstream customized tasks and upstream general model convergence requirements.
[0054] (2) The model layer trains the local neural network model based on the multimodal data of tactical edge nodes in different partitions. is the local data training set, where is the number of training samples, x j is the jth input training sample, y j is the label of the corresponding sample, and the vector w is defined as the full model parameter vector, and f(x j ,y j, w) is a simplified representation of the loss function, which specifically means the loss function of the jth sample. The goal of the training process is to minimize the average loss F(w), which is specifically defined as follows:
[0055]
[0056] After the local model is trained, the model layer uses random sampling to customize the model parameters and aggregate the local uploaded model parameters to the upper-level server to form a distributed federated aggregation framework for differentiated data in different partitions until the model converges globally.
[0057] (3) The application layer calls the asynchronous hierarchical federated learning strategy of the model layer, providing a unified and universal global model with strong security and low communication cost, and providing a customized and personalized neural network model for edge node adaptation. It focuses on adapting the close application of downstream tasks and local data to form customized downstream task auxiliary decision-making capabilities, and builds a complete tactical edge intelligence framework of local data partitioning-distributed model training aggregation-customized and personalized global application.
[0058] The asynchronous hierarchical federated learning strategy can solve the problems of poor communication quality and frequent network disconnections at the tactical edge. At the model layer of the tactical edge intelligent framework, in order to better adapt to the information resource aggregation structure of the command level and realize hierarchical authority information control and aggregation, an asynchronous hierarchical federated learning strategy that takes into account both security and multi-dimensional information aggregation is designed to support multi-modal and multi-task joint retrieval applications at the application layer.
[0059] The asynchronous hierarchical federated learning strategy is aimed at the multi-task and multimodal learning needs of the tactical edge. It is based on a customized federated cluster division strategy at the data layer. In the federated clusters at the tactical edge, the data and downstream tasks of the tactical edge federated clusters are divided according to the learning task objectives and the types of perception features, so that the functions of each cluster are personalized. In some clusters, computing power and storage resources are concentrated to process a single type of task. Through downstream task drive and local data protection, two model levels, the downstream task layer and the global task layer, are designed. Aiming at the intermittent characteristics of the tactical edge network and limited communication resources, an asynchronous mode is designed to form a training strategy for the local model and the global model of asynchronous hierarchical federated learning, thereby highlighting the downstream task customization capabilities of the local models and data of the tactical edge nodes, and the resource-saving performance of the global model, effectively enabling multi-dimensional perception of the tactical edge environment situation.
[0060] The asynchronous hierarchical federated learning strategy is based on the cloud-edge-end hierarchical architecture of the tactical edge, with the time-series asynchronous federated learning algorithm as the core, and adopts the time-series decay contribution method to allocate parameter weights. That is, the model with a longer update time has a larger update weight ratio than the model with a shorter update time. This ensures that a parameter model close to the real data change characteristics can be fitted while taking into account both historical information and new information.
[0061] In the asynchronous hierarchical federated learning strategy, data is iterated locally for a specified number of rounds from the bottom device, and then the preliminary model parameters are uploaded to the edge for aggregation. After the edge aggregates the model and performs complex data reinforcement training, the model is uploaded to the cloud. The cloud finally aggregates the data to form a global model, and broadcasts the global model within the network to ensure that each node can receive the model parameter push and perform local model updates.
[0062] During the training process, the Epoch timing method is uniformly used to discretize the continuous time. The main process uses the K variable for counting, and the threads use τ and t to synchronize the global transfer timestamp and local timestamp respectively. Through the gradient descent method, F(w) in formula (1) gradually approaches the minimum value, and finally makes w the optimal model parameter. At the same time, in order to standardize the process of gradient descent, k is specified as the index of the number of updates, and η is used as the scale of the updated gradient. The model update formula based on the gradient descent method is as follows:
[0063]
[0064] In the asynchronous hierarchical federated learning strategy, the data set is distributed on N computing terminals, so the data set of each terminal is specified as and And the gradient update is divided into local update and aggregate update. In order to balance the learning contribution between different terminal nodes, the proportion of participating in the gradient update is different according to the different proportions of the allocated data samples. Then, the functions of the two processes are respectively set to F(w) and F i (w), the definitions of the two functions are as follows:
[0065]
[0066] In the layering of the asynchronous hierarchical federated learning strategy, it is assumed that there are L edge devices in the edge layer, and l is the device index. Each edge device has C l Terminal node, let κ 1 Update the index for the terminal device layer and let κ 2 Update the index for the edge layer. If the total index k can be 1 If it is divisible by , then the local update is terminated, edge aggregation is started, and κ is reset 1; If the total index k can be 1 κ 2 If it is divisible by , then the edge gradient update is terminated, cloud aggregation is started, and κ is reset at the same time 1 and κ 2 , and start counting again from 1.
[0067] The gradient of the device No. i belonging to the edge region No. l is specified as At this point, based on the above description, the following definitions are available:
[0068]
[0069] At this time, according to the asynchronous hierarchical federated learning strategy, the number of training rounds is used to define the time. t ∈ T , that is, the total number of training rounds is T, the current number of training rounds is t, the number of training rounds where the sample is submitted is τ, and the decay function is set to s(·), then α t ←α×s(t-τ), where α t is the decay rate at time t.
[0070] At this time, set w t is the existing aggregation model, w t-1 is the aggregation model of the previous Epoch, w new For newly uploaded models, it is easy to have w t ←(1-α t ) t-1 +α t w new . Timestamp each sample data, convey the time according to the data gradient during aggregation, and achieve decay according to the overall training effect.
[0071] In order to better balance the contribution of short-term and long-term mergers in asynchronous federation to the total aggregate information composition, and to reflect the obvious changes in contribution over time, a linear function is used as the attenuation function.
[0072] The specific implementation is divided into two modules, namely Server and Client. The cloud is Server, the terminal is Client, and the edge device is both Server and Client according to the upload and download situations. When information is pushed to the cloud for aggregation, it is Client, and when the terminal information needs to be aggregated to the edge device, it is Server. r , in short, it is subject and object.
[0073] Two threads are started in the server. One is the Scheduler responsible for pushing the model to the global and refreshing the global timestamp. The other is the Updater responsible for asynchronous aggregation. When information is aggregated from the bottom, it first undergoes the weighted equivalent conversion of the Updater, and then accumulates the information layer by layer. The algorithm framework is as follows: Figure 3 shown.
[0074] In addition, in the tactical edge environment, information is highly dynamic, and the known requirements and unknown situations of the task are complex and complex, so it is difficult to fully define the execution details of a task, resulting in an incomplete processing flow framework, which directly leads to a sharp drop in the confidence level of the correctness of the predicted answer, making it difficult to meet the needs of solving tasks and assisting decision-making. Multimodality mainly focuses on fusing the feature layer with the neural network layer and the output layer, waiting for the query request, and then returning the result, while multitasking receives the request at the beginning, performs correlation priors in the process of processing the request task, and then performs synthesis, and finally directly outputs the result.
[0075] For example, when an unknown enemy threat unit appears, the threat threshold of this unit can be quickly estimated based on surface features and related behaviors combined with the previous analysis database. At the same time, the enemy combat unit type can be quickly screened based on existing unit information and some fuzzy features to finally get the result.
[0076] Therefore, in the asynchronous hierarchical federated learning strategy, the Encoder-Decoder form and the weighted network hybrid structure are specifically used to achieve multimodal and multi-task. Figure 4 Schematic diagram of a weighted hybrid network. Figure 5 Encoder-Decoder structure.
[0077] In a weighted network, the size of the weighted weight is evaluated using the data confidence method.
[0078] Data confidence is calculated based on the quantity and quality of core data elements. The specific evaluation method is to first determine the data elements that need to be monitored and improved through the funnel method, and then monitor the core data elements with improved structure in real time, adjust the data confidence at any time according to existing indicators, and calculate the weighted weights, so as to achieve dynamic multimodal data merging.
[0079] The funnel method mainly includes two parts: identification and optimization of core data elements:
[0080] 1. Identification: The core data elements are initially screened out through business experts and scoring matrices, but in fact there are still too many elements and further screening is still needed.
[0081] 2. Optimization: Filter the core data elements to be selected again through the relevance principle and data signal-to-noise ratio to obtain the final core data elements.
[0082] However, due to the large volume of tactical edge data, it is impossible for business experts to uniformly mark and screen the data value and quality of each core element. Therefore, a supervised learning approach is adopted, with business experts providing standard and sample data sets, and then the federated computing aggregation end learns an evaluator neural network to achieve the function of data optimization.
[0083] At the same time, the evaluator neural network also needs to have a certain degree of robustness, because the accuracy of data value assessment will seriously affect the processing of subsequent links, so a certain amount of erroneous data sets should be mixed in to enhance the self-correction ability of the neural network.
[0084] At the same time, data evaluation indicators will be updated and changed within a certain period of time through changes in data consistency, accuracy, completeness, uniqueness, relevance, and timeliness, and the value of data elements will be quickly sorted. Within limited communication and processing time, according to the principle of priority, high-value elements will be processed first, and low-value data elements will be processed later.
[0085] Example 1: Example of a federated learning method for tactical edge intelligence based on the MNIST dataset
[0086] In the MNIST dataset, images are 28x28 pixel grayscale images with uniform aspect ratios and stored in bytes. They are composed of SD-3 and SD-1, which contain binary images of handwritten numbers. The SD-3 data comes from employees of the Census Bureau, whose handwriting is relatively neat and the data is relatively recognizable. The SD-1 data comes from high school students, whose handwriting is relatively sloppy and difficult to recognize.
[0087] The MNIST training set consists of 30,000 writing patterns of SD-3 and 30,000 writing patterns of SD-1, and the test set consists of 5,000 writing patterns of SD-3 and 5,000 writing patterns of SD-1. This training set consists of writing patterns of about 250 authors, which serve as examples. Also, the author sets of the training set and the test set are disjoint.
[0088] SD-1 contains 58,527 handwritten digital images from 500 different authors. In SD-3, the data blocks from each writer appear in sequence and are readable normally, while the data in SD-1 is encrypted, but the writer's identity tag in SD-1 is not encrypted, and this information can be used to decipher the writer.
[0089] The specific file structure is as follows:
[0090] train-images-idx3-ubyte: training images
[0091] train-labels-idx1-ubyte: training labels
[0092] t10k-images-idx3-ubyte: test image
[0093] t10k-labels-idx1-ubyte: test label
[0094] Within the file, the tag takes a value from 0 to 9.
[0095] Image pixel represents grayscale, and its value range is 0 to 255. 0 represents pure white, that is, the background color, and 255 represents complete black.
[0096] Magic number is a magic number, which is a number used to indicate the file format.
[0097] The structure of the tag file is as follows:
[0098]
[0099] Image file structure:
[0100]
[0101] Basic operating environment:
[0102] In the tactical edge environment, the hardware is run by a ThinkPad P52s and multiple Raspberry pi 4Bs. The P52s acts as the host and the Raspberry pi 4B acts as the edge slave. In this network, the MNIST dataset and asynchronous hierarchical federated learning strategy are used to realize the image-based handwritten digit recognition function.
[0103] According to the actual equipment, in the simulation setting of the tactical edge, the number of layers n = 2. The main framework is as follows Figure 6 :
[0104] The data layer is based on the MNIST dataset. The hardware cluster is divided into ThinkPad P52s and multiple Raspberrypi 4Bs. The neural network model uses the Relu function as the hidden layer function and the Softmax function as the normalized classification function. It has four neural layers, two convolutional layers and two fully connected linear layers.
[0105] The tactical edge nodes are divided according to partitions. The client is mainly responsible for the functional definition of the client and downloads the MNIST expert evaluation dataset for model evaluation.
[0106] custom_server mainly defines the specific functional parameters of the server, including delay time, port and address, etc.
[0107] Servers mainly implement model aggregation functions and asynchronous communication functions.
[0108] split_dataset is mainly responsible for cutting the MNIST dataset and assigning it to sub-nodes.
[0109] The usage layer is divided into two modules: train and train_batch, which are single-step training and batch training respectively. Different modules are used for different situations.
[0110] Accuracy Verification:
[0111] like Figure 7 As shown in the figure, after relevant experiments, it is determined that asynchronous hierarchical federated learning is better than the original federated average learning algorithm.
[0112] Example 2: CLIP-based multimodal image-text joint retrieval example in tactical edge scenarios
[0113] In a tactical environment, a large number of multimedia snapshots are generally accumulated at the terminal layer through the perception of images, sounds, etc. At this time, for the convenience of storage, these snapshots are generally stored in the form of text-image pairs. In order to extract valuable information from these large number of snapshots and filter out redundant information, it is necessary to retrieve them.
[0114] CLIP is a text-image joint multimodal model trained by OpenAI using a large dataset. It is in the form of an encoder-decoder and can perform semantic segmentation on the input image, thereby outputting the semantic features in the image and expressing the nested relationship in the form of text description. Alternatively, it can retrieve the image collection through the characteristic elements and mutual relationships in the text description, and output the target image in the order of similarity. It can even use text and images as input at the same time to retrieve the target image.
[0115] Compared with traditional models, retrieval is better than classification problems, because classification requires continuous training and updating of the model as the number of categories increases, so as to ensure the accuracy of classification of new types of data, while retrieval is based on image semantic features, and has a certain migration ability when processing and analyzing. For categories that have never been encountered before, feature splicing can be used to approximate their semantics. If a certain similarity ratio is reached, it can be done. At the same time, CLIP also breaks out of the limitations of the bag-of-words model and uses the concept method to split the text. Compared with the bag-of-words model, the concept model not only includes image elements, but also contains the relationship between image elements. For example, "two dogs playing in the snow" is "Two dogs playing in the snow". The specific relationship is as follows: Figure 8 To verify the ability of the federated learning method for tactical edge intelligence to target downstream tasks and global applications, CLIP is used as the main method to learn by calling the target retrieval library through the WEB API. Specifically, the Google Colab interactive cloud virtualization environment is selected to run the upper-layer application part, and the federated learning environment is uploaded to the cloud in a packaged form. Pysyft is connected to CLIP through package reference to achieve real-time reasoning.
[0116] The hierarchical hardware resource reference setup for simulating a tactical edge intelligence environment is allocated as follows: Fig. 9 As shown, the interface API of CLIP is as follows Fig.10 In order to simplify the module, we do not use crawlers to crawl and preprocess the data here, but directly use Unplash, which provides an official dataset. This dataset is the full version of Unplash Dataset. The scale of Unplash Dataset is as follows: Fig.11 shown.
[0117] The data layer of the tactical edge intelligence architecture is based on the Unplash image website dataset. The data exists in the form of text-image pairs, which is naturally suitable for simulating text-image pair data in tactical environments. Therefore, this dataset is selected for simulation.
[0118] The cosine similarity value of text-image features is calculated using the norm, and the norm calculation formula is:
[0119]
[0120] The norm defaults to p=2.
[0121] The practical verification search results of the federated learning method for tactical edge intelligence are as follows Fig.12 shown.
Claims
1. A federated learning method for tactical edge intelligence, characterized by: Build a three-layer architecture of tactical edge intelligence, including a data layer covering multimodal local data, a model layer of available neural network models deployed at the tactical edge, and a tactical edge application layer for diverse tasks; The data layer "federates" the tactical edge computing nodes, divides the cluster partitions, strengthens the retention of data locally, and clusters and scores them according to the amount of multimodal task data and the computing resources and network resources of the tactical edge computing nodes to form a customized federated cluster partitioning strategy; The model layer trains the local neural network model based on the multimodal data of tactical edge nodes in different partitions. is the local data training set, where is the number of training samples, x j is the jth input training sample, y j is the label of the corresponding sample, and the vector w is defined as the full model parameter vector, and f(x j ,y j , w) is a simplified representation of the loss function, which specifically means the loss function of the jth sample; the goal of the training process is to minimize the average loss F(w), which is specifically defined as follows: After the local model is trained, the model layer randomly samples the model parameters and customizes the local model parameters to be aggregated and uploaded to the upper-layer server, forming a distributed federated aggregation framework for differentiated data in different partitions until the model converges globally. The application layer calls the asynchronous hierarchical federated learning strategy of the model layer, providing a global model upward and a customized and personalized neural network model for edge node adaptation downward. It focuses on adapting the close application of downstream tasks and local data to form customized downstream task auxiliary decision-making capabilities, and builds a complete tactical edge intelligence framework of local data partitioning-distributed model training aggregation-customized and personalized global applications.
2. The federated learning method for tactical edge intelligence according to claim 1, characterized in that: The data layer performs clustering and scoring based on the amount of multimodal task data and the computing resources and network resources of the tactical edge computing nodes. For customized cluster i, the amount of multimodal data is vi, the computing resources are ci, and the network resources are ni. The cluster score is si=x*vi+y*ci+z*ni, where the coefficient weight satisfies x+y+z=1. It is dynamically adjusted according to the data, computing, and network resources of the tactical edge environment to meet the convergence requirements of downstream customized tasks and upstream general models.
3. The federated learning method for tactical edge intelligence according to claim 1, characterized in that: The asynchronous hierarchical federated learning strategy is based on the cloud-edge-end hierarchical architecture of the tactical edge. In the asynchronous hierarchical federated learning strategy, data is iterated locally for a specified number of rounds from the bottom device, and then the preliminary model parameters are uploaded to the edge for aggregation. After the edge aggregates the model and conducts complex data reinforcement training, the model is uploaded to the cloud. The cloud finally aggregates the data to form a global model, and broadcasts the global model within the network to ensure that each node can receive the model parameter push and perform local model updates. With the time-series asynchronous federated learning algorithm as the core, the time-series decay contribution method is used to allocate parameter weights. That is, the model with a longer last update time has a larger update weight ratio than the model with a shorter update time, ensuring that a parameter model close to the real data change characteristics is fitted while taking into account both historical information and new information.
4. The federated learning method for tactical edge intelligence according to claim 1, characterized in that: In the asynchronous hierarchical federated learning strategy, the Encoder-Decoder form and the weighted network hybrid structure are used to achieve multi-modality and multi-task; During the training process, the Epoch timing method is used to discretize the continuous time. The main process uses the K variable for counting, and the threads use τ and t to synchronize the global transfer timestamp and local timestamp respectively; Through the gradient descent method, F(w) in formula (1) gradually approaches the minimum value, and finally w is made the optimal model parameter, k is the index of the update number, and η is used as the scale of the update gradient. The model update formula based on the gradient descent method is as follows: In the asynchronous hierarchical federated learning strategy, the data set is distributed on N computing terminals, and the data set of each terminal is and And the gradient update is divided into local update and aggregate update, and the functions of the two processes are F(w) and F i (w), the definitions of the two functions are as follows: In the layering of the asynchronous hierarchical federated learning strategy, it is assumed that there are L edge devices in the edge layer, and l is the device index. Each edge device has Terminal node, let κ1 be the terminal device layer update index, let κ2 be the edge layer update index, if the total index k is divisible by κ1, then terminate the local update, start edge aggregation, and reset κ1; if the total index k is divisible by κ1κ2, then terminate the edge gradient update, start cloud aggregation, and reset κ1 and κ2 at the same time, and start counting again from 1; The gradient of the device No. i belonging to the edge region No. l is specified as At this point, based on the above description, there are the following definitions: At this time, according to the asynchronous hierarchical federated learning strategy, the number of training rounds is used to define the time. Set t∈T, that is, the total number of training rounds is T, the current number of training rounds is t, the number of training rounds where the sample is submitted is τ, and the attenuation function is set to s(·), then α t ←α×s(t-τ), where α t is the decay rate at time t; At this time, set w t is the existing aggregation model, w t-1 is the aggregation model of the previous Epoch, w new For newly uploaded models, there is w t ←(1-α t ) t-1 +α t w new ; Timestamp each sample data, and convey the time according to the data gradient during aggregation, as well as the overall training effect to achieve attenuation.
Citation Information
Cited By
Wood cutting control method based on artificial intelligence
CN120950945A