Wireless federation training scheduling method based on double-channel policy network and related device
Patent Information
- Application Number
- CN202610800467.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-01
AI Technical Summary
然而,实际网络中各终端收集到的数据一般是非独立同分布的,且彼此之间的算力和通信能力也有很大差异,这些特性会对模型的收敛性和联邦过程的系统开销带来显著影响
本发明所述基于双通道策略网络的无线联邦训练调度方法及相关装置在具体操作时,以分布式模型协同训练系统中各轮开销小于预设门限为目标,构建资源调度问题,再基于强化学习以及基于双通道架构的决策网络求解所述资源调度问题,得到二元节点激活掩码及归一化带宽资源划分结果,在保障模型精度的前提下,有效降低系统的总体开销,实用性极强。
Smart Images

Figure CN122679461A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless resource management technology, and relates to a wireless federated training scheduling method and related apparatus based on a dual-channel policy network. Background Technology
[0002] In wireless network systems, smart terminals at the network edge can collect vast amounts of data and collaboratively train machine learning models. Traditional centralized methods require uploading data collected by each device to a central server for unified processing, often facing high communication costs and difficulties in guaranteeing data privacy. In recent years, distributed training methods have become increasingly popular. Federated Learning (FL), as a typical distributed training paradigm in this field, allows multiple terminals to jointly train a global model without sharing raw data, balancing data privacy and model performance. However, in real-world networks, the data collected by each terminal is generally not independent and identically distributed, and their computing power and communication capabilities vary significantly. These characteristics significantly impact model convergence and the system overhead of the federated process. Furthermore, wireless communication networks are characterized by limited bandwidth resources and time-varying channel conditions, exacerbating communication uncertainty. Based on these combined issues, model performance and system overhead require further optimization. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a wireless federated training scheduling method and related apparatus based on a dual-channel policy network. This method and related apparatus can improve model performance and reduce system overhead.
[0004] To achieve the above objectives, this invention discloses a wireless federated training scheduling method based on a dual-channel policy network, comprising: A resource scheduling problem is constructed with the goal of minimizing the overhead of each round in a distributed model collaborative training system to a preset threshold. By introducing a reinforcement learning paradigm, the resource scheduling problem is uniformly incorporated into a hybrid decision space, and then optimized based on a shared representation learning module to obtain an enhanced representation sequence. ; The enhanced representation sequence The input is fed into the optimized decision network based on the dual-channel architecture to obtain the binary node activation mask and the normalized bandwidth resource allocation result; Resource scheduling is performed based on the binary node activation mask and the normalized bandwidth resource allocation results.
[0005] Furthermore, the resource scheduling problem is as follows: (8) in, Indicates model accuracy. and These represent the preference weights for different types of costs. For the total delay, Total energy consumption, This is the preset energy consumption threshold for the k-th node. The preset delay threshold for the k-th node is... This represents the percentage of resource blocks allocated to the k-th node. For the first A subset selected in a round-robin process This refers to the total number of rounds.
[0006] Furthermore, obtain multidimensional observation sequences. The multidimensional observation sequence Depend on It is formed by concatenating the corresponding feature vectors of each node, thus forming a multidimensional observation sequence. The input is a multi-head attention layer, and the outputs of the multi-head attention layer are concatenated, then subjected to a linear transformation and a feedforward network to obtain the enhanced representation sequence. for: (13) in, ,in, Indicates the first Data distribution characteristics, channel gain, historical participation frequency, and system overhead of each node.
[0007] Furthermore, the decision network based on the dual-channel architecture includes a selection channel for admission determination and an allocation channel for resource allocation. The selection channel for admission determination outputs a binary node activation mask, and the allocation channel for resource allocation outputs a normalized bandwidth resource allocation result.
[0008] Furthermore, the selection channel for admission determination, during operation, provides enhanced representation for each node. First, it is mapped through two layers of fully connected networks and then transformed into selection confidence. ; Introducing threshold hyperparameters Make a judgment based on the selected confidence level. Select mask : (16) in, Indicates the first One node was selected to participate in this round of training. This indicates that it was not selected.
[0009] Furthermore, the total loss function of the decision network based on the dual-channel architecture during the optimization process is as follows: (twenty one) in, To align the weights of the regularization terms, To select the weights of the guiding regularization terms, To align regular expression terms, To select the guiding regular expression, To make policy networks By maximizing expectation The main loss during the optimization process.
[0010] Furthermore, align regular expressions and select guiding regular expression They are respectively: (twenty two) (twenty three) in, Represents the maximum value in the current state. value, To select only the first Approximation at 1 node value, To select the confidence level.
[0011] This invention discloses a wireless federated training scheduling system based on a dual-channel policy network, comprising: The module is used to construct a resource scheduling problem with the goal of minimizing the overhead of each round in a distributed model collaborative training system to a preset threshold. The optimization module introduces a reinforcement learning paradigm, unifying the resource scheduling problem into a hybrid decision space, and then optimizes it based on the shared representation learning module to obtain an enhanced representation sequence. ; Decision module, used to process the enhanced representation sequence The input is fed into the optimized decision network based on the dual-channel architecture to obtain the binary node activation mask and the normalized bandwidth resource allocation result; The scheduling module is used to perform resource scheduling based on the binary node activation mask and the normalized bandwidth resource allocation result.
[0012] This invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the wireless federated training scheduling method based on a dual-channel policy network.
[0013] The present invention discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the wireless federated training scheduling method based on a dual-channel policy network.
[0014] The present invention has the following beneficial effects: The wireless federated training scheduling method and related device based on dual-channel policy network described in this invention, in specific operation, aims to reduce the overhead of each round in the distributed model collaborative training system to a preset threshold, constructs a resource scheduling problem, and then solves the resource scheduling problem based on reinforcement learning and a decision network based on dual-channel architecture, to obtain the binary node activation mask and normalized bandwidth resource allocation results. Under the premise of ensuring model accuracy, it effectively reduces the overall overhead of the system and has strong practicality. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0019] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0020] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.
[0021] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0022] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0024] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0025] This invention designs a dual-channel policy network. The dual-channel policy network first extracts the shared context relationship representation of each candidate node, then outputs binary node activation masks and normalized bandwidth resource allocation through discrete and continuous channels in parallel, and ensures the integrity of gradient flow through gradient pass-through technology. To further improve learning stability, this invention introduces two auxiliary objectives within the TD3 framework: alignment regularization and selection-guided regularization, which are used to promote semantic consistency of the dual-channel output and alleviate search degradation caused by positive feedback effects, respectively. Compared to traditional segmented or serial scheduling strategies, this invention achieves end-to-end joint optimization of the complex action space, significantly reducing system overhead while improving model accuracy. Example 1 refer to Figure 1 The bandwidth resource scheduling method for wireless federated training described in this invention includes the following steps: 1) Construct a distributed model collaborative training system; Assuming the system deployment includes Each edge service node possesses computing capabilities, and its private data pool is denoted as . Due to the large scale of edge service nodes and limited wireless bandwidth resources, it is unrealistic to require all edge service nodes to participate in each round of training. In addition, the data distribution, computing power and wireless environment of different edge service nodes vary significantly. Indiscriminate participation of all nodes will introduce noise gradients and slow down the convergence process. Therefore, it is necessary to perform adaptive selection of participating nodes and bandwidth allocation.
[0026] Considering that downlink broadcast links typically have large power budgets and bandwidth margins, this invention focuses resource adaptation on the uplink. To suppress frequency-selective fading caused by multipath reflections during wireless signal propagation and to enable parallel access for multiple users, the physical layer employs orthogonal frequency division multiple access (OFDMA), finely dividing the total available spectrum bandwidth of the system into multiple mutually orthogonal subcarrier units. These subcarrier units are dynamically assigned by the central fusion node based on real-time requirements in each round. Specifically, for the first... Wheel, given a selected subset and total bandwidth resources For each activated node ,node The transmission power is The transmission distance and channel gain with the central fusion node are respectively and And given the proportion of the resource blocks that are divided. Then the uplink transmission rate achievable by this link is... for: (1) in, For noise power density, Let the large-scale path loss attenuation index be denoted by the node. The number of parameters transmitted is Then the delay on the communication side and energy consumption for: (2) (3) In addition, the overall system overhead also includes the computational load generated by local model updates. Specifically, let's assume nodes... The CPU clock speed is The number of CPU cycles and energy coefficient required to process a unit sample are respectively and The latency of local model training Energy consumption for: (4) (5) in, Let be the number of local iterations. Considering synchronization characteristics, the th iteration... The total latency generated by the round depends on the slowest node to complete the task, while the total energy consumption is the sum of the costs of all participating nodes, i.e.: (6) (7) To flexibly balance the impact of latency and energy consumption on different scenarios, the following is set and Let represent the preference weights for different types of costs, then the resource scheduling problem to be solved is: (8) in, Representing model accuracy, the core objective of this optimization problem is to find a set of node activation strategies and bandwidth allocation strategies that minimize long-term cumulative overhead without significantly degrading the model's convergence quality, provided that the overhead in each round does not exceed a preset threshold.
[0027] 2) Construct a reinforcement learning mechanism; Because the resource scheduling problem is non-convex and involves multiple constraints, traditional search strategies based on preset rules or metaheuristics are difficult to approximate the Pareto optimal solution for all objectives within a finite computation time. Therefore, this invention introduces a reinforcement learning paradigm to uniformly incorporate the resource scheduling problem into a hybrid decision space for optimization. The specific process is as follows: 21) Construct a multi-dimensional observation vector, with the central decision-making unit at the . Multidimensional observation vector of round observation for: (9) in, It is used to measure the degree of data distribution skewness of each node, and is indirectly reflected by the loss obtained after several rounds of local training at the beginning; This is used to represent the amount of data and the cumulative activation frequency of each node, providing valuable reference for subsequent selection, thereby balancing fairness and efficiency in decision-making; This will characterize the current channel gain, communication time and energy consumption in recent rounds, providing a direct reference for subsequent bandwidth allocation.
[0028] 22) Constructing a hybrid decision vector Hybrid decision vector of the central decision-making unit for: (10) in, Choose a mask for the binary discrete element, indicating whether a node is activated in the current round. If for the ... There are nodes This indicates that it will not participate in training in the current round, and its corresponding A value of 0 indicates that no wireless bandwidth is allocated to it. Furthermore, to maximize the utilization of bandwidth resources in the wireless network, this action space needs to satisfy constraints. .
[0029] 23) Designing the reward signal: To balance the quality of the optimization model with the system cost, the reward signal is designed as a multi-objective fusion form, namely: (11) in, This represents a regularization function used to smooth out the magnitude of precision changes and avoid misleading signals caused by precision saturation in the later stages of training. This represents a hyperparameter that balances various rewards, used to adjust the weights between model accuracy and system overhead, guiding the decision-making unit to adaptively optimize resource allocation in different scenarios.
[0030] 3) Optimization based on a dual-channel strategy network; Traditional reinforcement learning methods struggle to directly optimize the aforementioned hybrid decision-making problems, while common action discretization or step-by-step decision-making methods introduce suboptimal results. To address this, this invention proposes a decision network based on a dual-channel architecture. First, it extracts globally context-aware enhancement features from the heterogeneous states of multiple nodes. Second, it designs a parallel dual-channel structure, responsible for outputting discrete node selection masks and continuous bandwidth allocation ratios, respectively. Finally, within the TD3 framework, it introduces alignment regularization and selection-guided regularization as auxiliary optimization objectives, effectively enhancing policy exploration while maintaining semantic consistency, thus achieving efficient end-to-end optimization.
[0031] 31) Construct a shared representation learning module; As shown in equation (9), the multidimensional observation vector Depend on It is formed by concatenating the corresponding feature vectors of each node, and can be represented as a sequence. ,in, For the first Information such as data distribution characteristics, channel gain, historical participation frequency, and system overhead of each node. For Each component undergoes independent Z-score normalization before being input into the shared representation learning module to eliminate the natural differences in numerical magnitude between different features. To perceive the competitive and cooperative relationships among candidate nodes in the shared spectrum pool, this invention employs a multi-head self-attention module for relationship modeling. First, the multi-dimensional observation sequence... Input a multi-head attention layer, i.e.: (12) in, To be The query matrix, key matrix, and value matrix obtained after linear transformation The dimension of the feature vector.
[0032] The outputs of the multi-head attention layer are concatenated, then subjected to linear transformation and a feedforward network to output the enhanced representation sequence. for: (13) Each of them By integrating global contextual information, subsequent decisions can be made based on an understanding of the overall resource competition relationship.
[0033] 32) Construct a dual-channel decision-making network architecture; get Subsequently, this invention designs two independent and parallel policy output channels. The purpose of this design is that, although the two different decisions in equation (10) are highly coupled, they have essential differences in output properties and constraint structures. That is, the former belongs to a discrete decision space, while the latter belongs to a constrained continuous decision space. Using two dedicated channel networks to model separately can avoid representational conflicts caused by shared parameters, while allowing the decision network to design differentiated activation functions and exploration strategies.
[0034] 321) Selection path for admission determination: Enhanced representation for each node First, it is mapped through two layers of fully connected networks and then transformed into selection confidence. : (14) (15) The confidence level is the [number]. The probability estimate of each node being selected into the training subset. To obtain a hard binary discrete decision, a threshold hyperparameter needs to be introduced. Perform discrimination to obtain the selection mask. : (16) in, Indicates the first One node was selected to participate in this round of training. This indicates that the input was not selected. The threshold-based hard decision operation introduced a non-differentiable step function, blocking the gradient backpropagation path. If this non-differentiable point is directly ignored, this channel will not be able to obtain a valid training signal, resulting in the network outputting fuzzy intermediate values instead of clear decision boundaries. To solve this problem, this invention employs a gradient pass-through estimation technique, strictly using a selection mask during forward propagation. Subsequent calculations are performed to ensure the discreteness of the decision; during backpropagation, the gradient is approximately directly propagated back to... Bypassing the non-differentiable point of the threshold function, that is, let Through this approximation, the selection network is able to receive signals from... The gradient signal of the value estimation module and the auxiliary loss regularization term enables end-to-end training, effectively solving the gradient breakage problem while maintaining the discreteness of decision-making.
[0035] 322) Resource Allocation Channel: This channel is responsible for allocating bandwidth proportionally among activated nodes and outputting the allocation weight for each node. For each First, it needs to be mapped through a network with the same architecture as formula (14) to... Then, when adding exploration noise, directly adding it to the final action vector would violate the sparsity constraint (i.e., the partition weight corresponding to the unselected node must be zero). Therefore, this invention only adds Gaussian noise to the allocation channel, that is: (17) in, Let be the noise standard deviation, whose amplitude gradually decreases with the number of iterations. Then, let... For the selected server index set of discrete channel outputs, only the values within that set are considered. Perform the Softmax operation to maximize resource utilization: (18) Finally, combining the outputs of the two channels, the action of the decision-making unit is represented as follows: ,satisfy And for nodes that are not selected, .
[0036] 33) Parameter updating and auxiliary regularization term design; To optimize the policy network This invention first maintains two independent sets Value estimation network and and through double Smaller values are used to mitigate overestimation bias. Specifically, The value network will display the current state. and the complete action As input, and from the experience replay pool Random sampling of small batches of samples Then update using the following formula: (19) in, This is the reward discount factor. Then, By maximizing expectation The value is optimized, and the corresponding main loss is: (20) However, relying solely on this main loss is insufficient to address two key issues in this invention: first, the semantic consistency between outputs from different channels is difficult to guarantee; second, the positive feedback effect during training can lead to insufficient exploration. Therefore, this invention introduces an alignment regularization term and a selection guidance regularization term in addition to the main loss, which together constitute the total loss function: (twenty one) in, and For the corresponding weights.
[0037] Alignment regularization terms This regularization term constrains the semantic consistency of the two actions, requiring nodes with higher selection probabilities to receive greater weight in bandwidth allocation. Without this constraint, the two channels may produce contradictory outputs, meaning a node might be assigned a high selection probability but receive only a very low allocation weight, thus weakening decision consistency and introducing noisy gradients. To quantitatively measure the degree of semantic alignment, this invention defines the alignment loss in the form of cross-entropy, i.e.: (twenty two) This loss term can effectively adapt to the difference in the range between the selected confidence level and the bandwidth division ratio and provide a stable gradient signal, which can effectively promote the semantic alignment of the two channels.
[0038] Select guiding regular expression In the early stages of training, if certain nodes are frequently selected, their corresponding actions will be assigned higher priority. Value estimation. In subsequent policy updates, the policy network, in order to maximize expected returns, will further increase the selection probability of these nodes, forming a "high" value. Value → High selection frequency → Higher This creates a positive feedback loop of "value". This loop causes these nodes to maintain a long-term selection advantage, while other nodes, lacking the opportunity to be evaluated, cannot have their potential value accurately estimated. As a result, the policy network loses the possibility of discovering better combinations and eventually converges to a local optimum.
[0039] Based on the above, this invention designs a guiding loss to encourage the policy network to explore underestimated candidate nodes. Specifically, when the policy network repeatedly favors a few nodes due to positive feedback, it forcibly injects a reverse gradient signal into them, compelling those nodes that have been neglected for a long time but have independent training value to have a higher probability of being selected, thereby increasing the exploratory nature of the network. The guiding loss is: (twenty three) in, Represents the maximum value in the current state. value, To select only the first Approximation at 1 node The difference between the two values measures the node's value. The larger the return gap between the current valuation system and the optimal choice, the more underestimated the marginal gain that the node might bring after being selected into the subset. This loss enhances the exploration capability of the policy network without disrupting the convergence direction of the main task, prompting the final resource scheduling scheme to approach the global optimum more closely.
[0040] It should be noted that the present invention has the following characteristics: Traditional federated learning resource scheduling methods based on heuristic search or mathematical solutions are ill-suited to dynamic wireless environments and typically involve high computational complexity. This invention, based on reinforcement learning algorithms, first designs a state space that integrates edge server data distribution characteristics, real-time channel states, and historical scheduling costs, enabling the agent to perceive multi-dimensional environmental information. Building upon this, the invention integrates discrete device selection decisions and continuous bandwidth allocation decisions into a unified end-to-end decision framework, and optimizes it using a multi-objective reward function, effectively reducing the overall system overhead while ensuring model accuracy.
[0041] Conventional methods typically employ two-stage strategies, which struggle to model the complex distributions in mixed action spaces, leading to suboptimal results. This invention proposes a dual-channel policy network structure. Based on multi-head attention extracting state features, the selection channel uses threshold-gated output discrete decisions with gradient pass-through estimation, while the allocation channel performs continuous resource allocation based on the selected subset, forming an end-to-end unified joint decision-making framework. Furthermore, this invention further designs alignment loss and selection guidance auxiliary loss based on the TD3 algorithm, effectively improving training efficiency and policy robustness.
[0042] Example 2 The wireless federated training scheduling system based on a dual-channel policy network described in this invention is characterized by comprising: The module is used to construct a resource scheduling problem with the goal of minimizing the overhead of each round in a distributed model collaborative training system to a preset threshold. The optimization module introduces a reinforcement learning paradigm, unifying the resource scheduling problem into a hybrid decision space, and then optimizes it based on the shared representation learning module to obtain an enhanced representation sequence. ; Decision module, used to process the enhanced representation sequence The input is fed into the optimized decision network based on the dual-channel architecture to obtain the binary node activation mask and the normalized bandwidth resource allocation result; The scheduling module is used to perform resource scheduling based on the binary node activation mask and the normalized bandwidth resource allocation result.
[0043] The module division in this embodiment is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in each embodiment of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0044] Example 3 A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the wireless federated training scheduling method based on a dual-channel policy network. For example, the steps include: constructing a resource scheduling problem with the objective that the overhead of each round in the distributed model collaborative training system is less than a preset threshold; introducing a reinforcement learning paradigm to unify the resource scheduling problem into a hybrid decision space; and then optimizing it based on a shared representation learning module to obtain an enhanced representation sequence. The enhanced representation sequence The input is fed into the optimized decision network based on a dual-channel architecture to obtain the binary node activation mask and normalized bandwidth resource allocation results; resource scheduling is then performed based on these results. The memory may include main memory, such as high-speed random access memory, or it may also include non-volatile memory, such as at least one disk storage device. The processor, network interface, and memory are interconnected via an internal bus, which can be an industry-standard architecture bus, a peripheral component interconnection standard bus, or an extended industry-standard architecture bus. The bus can be categorized as an address bus, data bus, or control bus. The memory stores programs; specifically, the programs may include program code, which includes computer operation instructions. The memory may include main memory and non-volatile memory, and it provides instructions and data to the processor.
[0045] Example 4 A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the wireless federated training scheduling method based on a dual-channel policy network. For example, the method includes: constructing a resource scheduling problem with the objective that the overhead of each round in the distributed model collaborative training system is less than a preset threshold; introducing a reinforcement learning paradigm to uniformly incorporate the resource scheduling problem into a hybrid decision space; and then optimizing it based on a shared representation learning module to obtain an enhanced representation sequence. The enhanced representation sequence The input is fed into the optimized decision network based on a dual-channel architecture to obtain binary node activation masks and normalized bandwidth resource allocation results; resource scheduling is then performed based on these results. Specifically, the computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0046] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0047] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0048] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0049] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0050] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and disclosure of the invention. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0051] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
[0052] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any simple modifications, alterations, or equivalent structural changes made to the above embodiments based on the technical essence of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A wireless federated training scheduling method based on a dual-channel policy network, characterized in that, include: A resource scheduling problem is constructed with the goal of minimizing the overhead of each round in a distributed model collaborative training system to a preset threshold. By introducing a reinforcement learning paradigm, the resource scheduling problem is uniformly incorporated into a hybrid decision space, and then optimized based on a shared representation learning module to obtain an enhanced representation sequence. ; The enhanced representation sequence The input is fed into the optimized decision network based on the dual-channel architecture to obtain the binary node activation mask and the normalized bandwidth resource allocation result; Resource scheduling is performed based on the binary node activation mask and the normalized bandwidth resource allocation results.
2. The wireless federated training scheduling method based on a dual-channel policy network according to claim 1, characterized in that, The resource scheduling problem is as follows: (8) in, Indicates model accuracy. and These represent the preference weights for different types of costs. For the total delay, Total energy consumption, This is the preset energy consumption threshold for the k-th node. The preset delay threshold for the k-th node is... This represents the percentage of resource blocks allocated to the k-th node. For the first A subset selected in a round-robin process This refers to the total number of rounds.
3. The wireless federated training scheduling method based on a dual-channel policy network according to claim 1, characterized in that, Obtaining multidimensional observation sequences The multidimensional observation sequence Depend on It is formed by concatenating the corresponding feature vectors of each node, thus forming a multidimensional observation sequence. The input is a multi-head attention layer, and the outputs of the multi-head attention layer are concatenated, then subjected to a linear transformation and a feedforward network to obtain the enhanced representation sequence. for: (13) in, ,in, Indicates the first Data distribution characteristics, channel gain, historical participation frequency, and system overhead of each node.
4. The wireless federated training scheduling method based on a dual-channel policy network according to claim 1, characterized in that, The decision network based on the dual-channel architecture includes a selection channel for admission determination and an allocation channel for resource allocation. The selection channel for admission determination outputs a binary node activation mask, and the allocation channel for resource allocation outputs a normalized bandwidth resource allocation result.
5. The wireless federated training scheduling method based on a dual-channel policy network according to claim 4, characterized in that, The selection channel for admission determination, during operation, provides enhanced representation for each node. First, it is mapped through two layers of fully connected networks and then transformed into selection confidence. ; Introducing threshold hyperparameters Make a judgment based on the selected confidence level. Select mask : (16) in, Indicates the first One node was selected to participate in this round of training. This indicates that it was not selected.
6. The wireless federated training scheduling method based on a dual-channel policy network according to claim 4, characterized in that, The total loss function of the decision network based on the dual-channel architecture during the optimization process is as follows: (21) in, To align the weights of the regularization terms, To select the weights of the guiding regularization terms, To align regular expression terms, To select the guiding regular expression, To make policy networks By maximizing expectation The main loss during the optimization process.
7. The wireless federated training scheduling method based on a dual-channel policy network according to claim 6, characterized in that, Alignment regularization terms and select guiding regular expression They are respectively: (22) (23) in, Represents the maximum value in the current state. value, To select only the first Approximation at 1 node value, To select the confidence level.
8. A wireless federated training scheduling system based on a dual-channel policy network, characterized in that, include: The module is used to construct a resource scheduling problem with the goal of minimizing the overhead of each round in a distributed model collaborative training system to a preset threshold. The optimization module introduces a reinforcement learning paradigm, unifying the resource scheduling problem into a hybrid decision space, and then optimizes it based on the shared representation learning module to obtain an enhanced representation sequence. ; Decision module, used to process the enhanced representation sequence The input is fed into the optimized decision network based on the dual-channel architecture to obtain the binary node activation mask and the normalized bandwidth resource allocation result; The scheduling module is used to perform resource scheduling based on the binary node activation mask and the normalized bandwidth resource allocation result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the wireless federated training scheduling method based on a dual-channel policy network as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the wireless federated training scheduling method based on a dual-channel policy network as described in any one of claims 1-7.