Method and device for constructing a combat communication system
By employing the UCB method and MA-MAB model in the combat communication system, and dynamically selecting and allocating communication nodes, the problems of limited communication resources and variable node locations on the battlefield are solved, achieving efficient and fair communication connections and maximizing QoE, thereby improving combat effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2024-04-26
- Publication Date
- 2026-04-28
AI Technical Summary
Existing combat communication systems are unable to effectively build flexible and efficient communication connections on the battlefield, especially under conditions of limited resources, variable node locations, and heterogeneous requirements. They cannot meet communication needs, and existing methods neglect the measurement and improvement of user experience quality (QoE).
A target communication node is determined from multiple communication nodes using a UCB-based target strategy. Communication connections are optimized through a resource allocation algorithm, and the MA-MAB model is combined for learning and dynamic adjustment to ensure efficient allocation of communication resources and maximize QoE.
It enables the dynamic construction of efficient and fair communication connections on the battlefield, meeting the communication needs of combat nodes and improving combat effectiveness and overall QoE value.
Smart Images

Figure CN118828908B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, and specifically relates to a method and apparatus for constructing an operational communication system. Background Technology
[0002] The development of information technology has enabled real-time data transmission between combat nodes (such as tanks and fighter jets), facilitating the evolution of warfare from platform-based to system-based confrontation. In system-based confrontation, battlefield information is typically acquired independently by a wide range of combat nodes. Effective sharing of situational awareness information among these nodes is essential for achieving situational awareness. Specifically, in time-critical combat missions such as air defense and missile defense, the combat system urgently needs rapid response to implement immediate combat actions, which places higher demands on the operational communication architecture.
[0003] The operational communication architecture is the foundation for building an operational system and a crucial guarantee for effective command and control. A typical operational system usually consists of various interdependent and interacting component systems, including operational nodes and communication nodes. To better accomplish assigned tasks, the composition and structure of the operational system should dynamically change according to mission requirements. Simultaneously, communication nodes must provide high-quality services to operational nodes, ensuring effective and uninterrupted communication. Therefore, a flexible and efficient operational communication architecture is extremely important in modern operational systems, as it can significantly enhance operational capabilities to accomplish missions.
[0004] However, designing a flexible and efficient combat communication architecture presents several challenges. First, communication resources are limited. Unlike well-developed civilian communication network infrastructure, battlefields are typically far from cities, with weak infrastructure. Furthermore, for reasons of concealment and security, communication services for combat nodes are usually provided by communication nodes within the combat system, resulting in limited communication resources. Second, resource requirements are heterogeneous. Due to the heterogeneity of the combat system's composition, combat nodes have varying time sensitivities to combat operations and different communication resource requirements. This is especially true for intelligent devices (such as drones), which require transmitting larger volumes of data (such as video streams) within specific time windows. Finally, the movement of combat nodes is random. To avoid detection and potential enemy attacks, combat nodes often employ tactics involving chaos and randomness, leading to poor performance of traditional offline methods in designing communication architectures. Therefore, centralized methods are impractical, necessitating an online approach to effectively construct combat communication architectures without prior knowledge of node states.
[0005] Furthermore, many existing methods only consider the communication coverage of operational nodes, neglecting how to measure and improve their Quality of Experience (QoE). Some methods only consider maximizing coverage area and number of users, without addressing communication connection construction, especially in large-scale scenarios. Other methods focus primarily on resource allocation at the individual base station level, without considering the collaborative allocation across multiple base stations in large-scale environments. Consequently, the designed methods generally result in poor communication performance in practical applications. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and apparatus for constructing an operational communication system that can adapt to the changing locations of combat nodes and communication nodes on the battlefield, provide a flexible and efficient way to establish communication connections for combat nodes and communication nodes, and meet communication requirements.
[0007] The present invention includes a flexible construction method for an operational communication system, applied to operational nodes, the method comprising:
[0008] Under a target time slot, candidate communication nodes that meet the preset constraints are determined from multiple communication nodes. The target time slot is obtained by dividing the combat duration of the combat node by a specified step size. Under different time slots, the positions of the combat nodes are the same or different, the positions of the communication nodes are the same or different, and the determined candidate communication nodes are at least not completely the same.
[0009] Based on the target strategy, a target communication node is determined from the candidate communication nodes, and a communication connection is established with the target communication node. The target strategy is generated based on the UCB method and is used to determine the target communication node based on the first sub-strategy in the first scenario and to determine the target communication node based on the second sub-strategy in the second scenario.
[0010] Obtain the communication resources allocated by the target communication node.
[0011] In some embodiments, the method further includes:
[0012] Obtain the QoE result of the combat node in the target time slot fed back by the target communication node, the QoE result being determined based on the communication resources required by the combat node and the actual communication resources obtained;
[0013] The parameters of the target policy are updated based on the QoE result so that, in the next time slot, if the target communication node is selected based on the first sub-policy, the generated QoE result is improved.
[0014] In some embodiments, determining the target communication node based on the first sub-policy in the first scenario includes:
[0015] If there is no communication node among the candidate communication nodes that has never been selected as a target communication node, the upper confidence bound of each candidate communication node is evaluated based on the number of times each candidate communication node has been selected as a target communication node, the historical QoE results fed back by the candidate communication node, and the standard deviation calculated from the mean of the historical QoE results, respectively.
[0016] The candidate communication node with the highest confidence upper bound is selected as the target communication node for the current time slot.
[0017] In some embodiments, updating the parameters of the target policy based on the QoE result includes:
[0018] Update the stored data of candidate communication nodes selected as target communication nodes in the current time slot, including the number of times a node has been selected as a target communication node, the QoE result of the most recent time node selected as a target communication node, and the standard deviation.
[0019] In some embodiments, determining the target communication node based on the second sub-policy in the second scenario includes:
[0020] If there is a communication node among the candidate communication nodes that has never been selected as the target communication node, then that communication node is directly determined as the target communication node.
[0021] In some embodiments, the preset limiting conditions include at least one or more of the following conditions:
[0022] The communication distance coverage condition, the maximum number of connections a communication node can make to an operating node, and the unique communication condition that restricts an operating node to only one communication node in each time slot;
[0023] When determining the target communication node based on the first sub-strategy, the following is also included:
[0024] If the number of connections established by the candidate communication node corresponding to the highest confidence upper bound already meets the maximum number of connections condition, then the candidate communication node with the second highest confidence upper bound is selected as the target communication node.
[0025] Another embodiment of the present invention also provides a flexible construction method for a combat communication system, applied to a communication node, the method comprising:
[0026] Under the target time slot, respond to the connection request of the combat node that meets the preset restrictions and establish a communication connection with the combat node;
[0027] The multiple combat nodes that simultaneously establish the communication connection under the target time slot are divided into a cell;
[0028] Based on the principle of maximizing the overall QoE value of the cell and the principle of fairness, communication resources are allocated to each combat node in the cell. The principle of fairness indicates that the difference in communication resources obtained between any two combat nodes meets the preset requirements.
[0029] In some embodiments, allocating communication resources to each combat node within the cell based on the principles of maximizing the overall QoE value of the cell and fairness includes:
[0030] Allocable communication resources are divided into multiple resource time slots;
[0031] Calculate the QoE increment for each combat node within the cell when it acquires a resource time slot;
[0032] Based on the QoE increment of each combat node and the principle of fairness, the multiple resource time slots are allocated one by one to complete the allocation of communication resources for each combat node in the cell.
[0033] The method further includes:
[0034] The QoE value of each combat node under the target time slot is calculated and fed back to the corresponding combat node so that the combat node can select a communication node to establish a communication connection in the next time slot.
[0035] Another embodiment of the present invention provides a flexible construction device for an operational communication system, applied to an operational node, the device comprising:
[0036] The first determining module is used to determine, under a target time slot, candidate communication nodes that meet the preset constraints among multiple communication nodes. The target time slot is obtained by dividing the combat duration of the combat node by a specified step size. Under different time slots, the positions of the combat nodes are the same or different, the positions of the communication nodes are the same or different, and the determined candidate communication nodes are at least not completely the same.
[0037] The second determining module is used to determine the target communication node from the candidate communication nodes according to the target strategy, and to establish a communication connection with the target communication node. The target strategy is generated based on the UCB method and is used to determine the target communication node based on the first sub-strategy in the first scenario and based on the second sub-strategy in the second scenario.
[0038] The first acquisition module is used to acquire the communication resources allocated by the target communication node.
[0039] Another embodiment of the present invention provides a flexible construction device for an operational communication system, applied to a communication node, the device comprising:
[0040] The response module is used to respond to the connection requests of combat nodes that meet preset restrictions in the target time slot and establish a communication connection with the combat nodes.
[0041] The partitioning module is used to divide multiple combat nodes that simultaneously establish the communication connection under the target time slot into a cell;
[0042] The allocation module is used to allocate communication resources to each combat node in the cell according to the principle of maximizing the overall QoE value of the cell and the fairness principle, wherein the fairness principle indicates that the difference in communication resources obtained between any two combat nodes meets the preset requirements.
[0043] The beneficial effects of this invention are that it can address the challenges of limited communication resources, constantly changing resource requirements and node locations during the construction of combat communication architecture. It enables combat nodes and communication nodes to dynamically build connections through mutual interaction and learning, allowing new and suitable communication nodes to be selected for connection as needed in each combat time slot. Communication nodes can allocate communication resources according to the needs and locations of combat nodes. This allocation of communication resources is not only efficient and fair, but also strives to maximize the overall QoE value of combat nodes, thereby improving combat effectiveness. Attached Figure Description
[0044] Figure 1 This is a flowchart of the method for constructing the combat communication system of the present invention.
[0045] Figure 2 The learning process result curve is shown for the construction method of the combat communication system of the present invention.
[0046] Figure 3 This is a flowchart of a method for constructing another combat communication system according to the present invention.
[0047] Figure 4 This is a structural block diagram of the device for constructing the combat communication system of the present invention.
[0048] Figure 5 This is a structural block diagram of a device for constructing another combat communication system according to the present invention. Detailed Implementation
[0049] The embodiments in this invention primarily consider two types of forces: communication forces and combat forces. Specifically, the combat system comprises M communication nodes and N combat nodes. Communication nodes are equipped with communication equipment (such as 4G / LTE base stations) and can provide communication services (such as uploading / downloading combat data) to combat nodes. Combat forces (combat nodes) include early warning and detection forces, command and control forces, and firepower strike forces, etc. These forces will generate different communication resource requirements (such as bandwidth) according to operational needs.
[0050] In this embodiment, the division of combat time slots is based on a discrete-time model. For example, the duration of an operation (e.g., one hour) is divided into T consecutive time slots of length Δt (a specified step size), denoted by T = {1, 2, ..., T}. For simplicity, this embodiment assumes the battlefield is a two-dimensional plane divided into identical squares (e.g., 50m × 50m). To avoid detection and attack, combat nodes will frequently change positions, moving from one square to another.
[0051] A combat system with M communication nodes can be represented by a set M = {1, 2, ..., M}. For any time slot t, the state of communication node m can be represented by the tuple m(t) = ... <x m (t),y m (t),b m ,e m ,s m ,p m > indicates. Specifically, i) as a high-value target in combat, communication nodes need to frequently change their location, which can be represented by x. m (t) and y m (t) to measure, ii)b m Let m represent the total communication resources (i.e., bandwidth) of the communication node m. For simplicity, the total bandwidth can be discretized into several equal sub-bandwidths (such as resource slots as described later). iii)e m Indicates the effective communication coverage radius, iv)s m p represents the maximum number of operational nodes that a communication node can serve. m Indicates the transmission power.
[0052] Assuming there are N combat nodes in the combat system, they can be represented by a set N = {1, 2, ..., N}. For any time slot t, the state of combat node n can be represented by the tuple n(t) = <x n (t),y n (t),r n (t),c n (t)> represents this. Specifically, 1) the combat node n will move to a specific block in time slot t, and its position can be determined by x.n (t) and y n (t) indicates that, 2) because combat nodes will perform different tasks at different times, they will have heterogeneous communication resource requirements r. n (t), 3)c n (t) refers to the actual communication resources received.
[0053] Based on the above content, such as Figure 1 As shown, this invention provides a flexible construction method for a combat communication system to address the problem of the inability to flexibly construct communication connections on the battlefield. This method is applied to each combat node, such as by intelligent devices within each combat node. The method includes:
[0054] S1: Under the target time slot, candidate communication nodes that meet the preset constraints are determined among multiple communication nodes. The target time slot is obtained by dividing the combat duration of the combat node by a specified step size. Under different time slots, the positions of the combat nodes are the same or different, the positions of the communication nodes are the same or different, and the determined candidate communication nodes are at least not completely the same.
[0055] S2: Determine the target communication node from the candidate communication nodes based on the target strategy, and establish a communication connection with the target communication node. The target strategy is generated based on the UCB method and is used to determine the target communication node based on the first sub-strategy in the first scenario and to determine the target communication node based on the second sub-strategy in the second scenario.
[0056] S3: Obtain the communication resources allocated by the target communication node.
[0057] Based on this communication resource, the combat node can complete the combat mission in the current time slot.
[0058] The preset restrictions include at least one or more of the following conditions:
[0059] The communication distance coverage condition, the maximum number of connections a communication node can make to an operator node, and the unique communication condition that restricts an operator node to only one communication node in each time slot.
[0060] Specifically, the construction of communication links between combat nodes is the foundation of combat communication architecture design, ensuring the effective transmission and exchange of information such as early warning detection, command and control, and fire strikes within the combat system. With basic communication infrastructure in place, once a combat node connects to a communication node, it will be connected to the combat communication network and can communicate with other combat nodes. The connection relationships and resource allocation matrix between operational nodes and communication nodes within time slot t are as follows:
[0061]
[0062] Among them o m,n (t) is a non-negative real number, representing the communication resources allocated by communication node m to operational node n in time slot t. The prerequisite for a communication node to allocate communication resources to an operational node is that a communication connection is established between them, and the following constraints must be met:
[0063] Communication coverage is limited. Because communication strength gradually decreases with increasing distance, communication nodes typically have an effective communication coverage range, i.e., e. m The operating node n can only receive communication services within the effective communication coverage area of the communication node m, which can be represented as:
[0064]
[0065] Where d m,n (t) is the distance between nodes, which can be represented as:
[0066]
[0067] Maximum number of connections. To ensure that service nodes within the communication coverage area can provide high QoE services, the number of combat nodes that each communication node can serve cannot exceed the maximum number of connections of the communication node, which can be specifically expressed as:
[0068]
[0069] Where I m,n (t) is an indicator that can be measured using the following formula:
[0070]
[0071] Unique communication connection. To ensure stable data transmission, each combat node can only connect to one communication node, which can be represented as:
[0072]
[0073] Furthermore, the method described above in this embodiment is based on deploying a multi-agent multi-armed bandit (MA-MAB) model within the intelligent device of each operational node to continuously learn the first strategy, thereby selecting and determining a more suitable optimal target communication node in the next time slot. The MA-MAB model allows multiple agents (intelligent devices of the combat nodes) to iteratively learn the environmental state and make decisions to minimize cumulative losses (i.e., the difference between the current decision and the optimal decision). In the context of this embodiment, each combat node can act as an agent, and the decision is to select the optimal communication node to provide communication services. Since the current state of the operational node (e.g., communication resources) is unknown to the communication node, its current state must be obtained through historical experience and continuous interaction.
[0074] Specifically, for each agent, `arm` represents a candidate agent (candidate communication node) for the combat node to choose from. The combat node needs to select the optimal `arm` from the candidate agents in each time slot. Correspondingly, from the perspective of the operating node `n`, each communication node is considered as a robotic arm. The combat node needs to select the optimal robotic arm (i.e., the optimal communication node, the target communication node) from the candidate robotic arms (i.e., the set of candidate communication nodes). For each combat node, its action vector (i.e., connection relationship) is defined as `j`. n (t)=[j n:1 (t),…,j n:M [(t)], where j n:m (t) being 0 or 1 indicates a connected or disconnected relationship, respectively, and its meaning is the same as that of O. :,n (t) are the same.
[0075] After selecting a target communication node, that node m will allocate resources to the connected combat node n based on a resource allocation function. The combat node will then receive a reward to evaluate the decision made in this round (i.e., whether the selected communication node is the optimal communication node) and update its first strategy. For each arm, the reward g(t) of the arm in time slot t is defined as the QoE value q of the combat node in that time slot. n (t), that is, g(t) = q n (t), where g(t)∈[0,1]. Since the goal of MAB is to maximize the total benefit in the long run, the operational node needs to select the optimal arm-target communication node in each round. Therefore, each operational node needs to continuously interact with the environment to improve its decision-making ability in order to select the optimal robotic arm-target communication node in the next round. To this end, this embodiment designs a learning strategy (i.e., a target strategy) for each operational node in each time slot to select the optimal arm (target communication node), thereby maximizing the total reward, that is, maximizing the QoE value, throughout the entire operation duration.
[0076] Based on the above, the method in this embodiment further includes:
[0077] S4: Obtain the QoE result of the combat node in the target time slot fed back by the target communication node. The QoE result is determined based on the communication resources required by the combat node and the actual communication resources obtained.
[0078] S5: Update the parameters of the target policy based on the QoE results so that in the next time slot, if the target communication node is selected based on the first sub-policy, the corresponding generated QoE results will be improved.
[0079] In the first scenario, determining the target communication node based on the first sub-strategy includes:
[0080] S6: In the case that there is no communication node among the candidate communication nodes that has never been selected as the target communication node, the confidence upper bound of each candidate communication node is evaluated by retrieving the stored number of times each candidate communication node has been used as the target communication node, the historical QoE results fed back by the candidate communication nodes, and the standard deviation calculated from the mean of the historical QoE results, respectively.
[0081] S7: Select the candidate communication node with the highest confidence upper bound as the target communication node for the current time slot.
[0082] Furthermore, when determining the target communication node based on the first sub-strategy, it also includes:
[0083] S8: If the number of connections established by the candidate communication node corresponding to the highest confidence upper bound already meets the maximum number of connections condition, then select the candidate communication node with the second highest confidence upper bound as the target communication node.
[0084] The parameters of the target policy are updated based on the QoE results, including:
[0085] S9: Update the stored data of the candidate communication nodes selected as the target communication node in the current time slot. The data includes the number of times the node has been selected as the target communication node, the QoE result and standard deviation of the most recent time the node was selected as the target communication node.
[0086] In the second scenario, the target communication node is determined based on the second sub-policy, including:
[0087] S10: If there is a communication node among the candidate communication nodes that has never been selected as the target communication node, directly determine the communication node as the target communication node.
[0088] Specifically, given instantaneous knowledge of the environment, combat nodes can apply the UBC-based algorithm to solve the connection relationship construction problem. This algorithm quantifies the uncertainty of the environment, comprehensively considers exploration and utilization, weighs the trade-off between exploring unknown weapons and potential high returns, and weighs the trade-off between utilizing known weapons and prior knowledge, in order to obtain long-term high returns.
[0089] The UCB-based strategy selects the arm with the highest weight, which includes the arm's average cumulative reward over the first t-1 time steps (slots) plus an additional reward. This additional reward is determined by the standard deviation of the average cumulative reward over the first t-1 time steps (slots), representing the instability of the candidate arm. The highest weight formed based on the above rewards serves as the upper bound of the confidence interval. The target policy generated by the UCB method operates under the optimistic principle of face-to-face uncertainty, enabling the combat node to select the communication node with the highest upper confidence boundary, defined as:
[0090]
[0091] in, It is the arm j up to time slot t. n The average return, is the number of times it is selected as an arm by the combat node n, and c is a parameter that balances exploration and utilization.
[0092] The connection relationship construction algorithm based on UBC, i.e., the content of the target strategy, is described as follows: It is used to explore the distribution of communication resources and provide an asymptotically optimal compromise for exploration and utilization as an online strategy. Specifically, since the combat node initially has no prior knowledge of the communication nodes, each communication node must be immediately selected within its coverage area to initialize the weapon's weight. First, each combat node must determine, based on preset constraints, which communication nodes it falls within the coverage area of for each time slot, thus determining the set of connectable communication nodes, i.e., the candidate node set. If there is a communication node in the set that combat node n has never selected before, then that communication node is directly selected as the target communication node in this round. This process is mainly used for initial exploration. Otherwise, the combat node will select the arm with the highest upper bound confidence for exploration and utilization. At the beginning of each time slot, the combat node will evaluate the upper bound confidence of each candidate communication node and select the candidate communication node with the highest upper bound confidence. If the selected candidate communication node has reached the maximum number of connections (e.g., the constraint connection count is 4), the combat node will select the candidate communication node with the second upper bound confidence, and so on. After each operational node selects a target communication node, the connection matrix between the operational and communication nodes is immediately determined. At this point, the communication node will allocate resources to connected operational nodes within the same cell according to a resource allocation function, which will be described in detail later. Based on this, the target communication node can be determined according to the following formula: (r n (t) represents the required communication resources, and c n (t) represents the actual received communication resources, α is a parameter taking values of (0,+∞), and q n (t) represents the QoE result of the combat node n in time slot t. The QoE results of each operational node in that time slot are calculated and sent back to the corresponding combat node for use in the average cumulative reward for updating combat node experience (i.e., ...). This is used to evaluate the optimistic bound of the communication node in the next time slot. The above process is repeated in each time slot (or round) until the last T time slot ends.
[0093] The method in this embodiment runs a target policy based on the UCB algorithm, and... The regret boundary is established in the form of T. The regret boundary and T are clearly sublinear. Therefore, the proposed UCB-based strategy has adaptive learning capabilities, and can make near-optimal connection construction decisions in real time, even when some network state information is unavailable.
[0094] like Figure 2 As shown in the figure (the horizontal axis represents Episodes, and the vertical axis represents Cumulative QoE), it illustrates the learning process of an operational node in determining a target communication node in an application instance. The operational node iteratively calculates and updates the weight of each communication node based on the feedback QoE results, thereby selecting the optimal communication node in the next round. Figure 2 As shown by the curve, the cumulative QoE of the system rises rapidly in the early stage of learning. However, as the number of learning iterations increases, after many learning iterations, such as 20 as shown in the figure, the algorithm can effectively converge to a relatively stable performance, with high overall learning efficiency and excellent learning results.
[0095] like Figure 3 As shown, another embodiment of the present invention also provides a flexible construction method for a combat communication system. This method is applied to each communication node, such as by intelligent devices within each communication node. The method includes:
[0096] S1: Under the target time slot, respond to the connection request of the combat node that meets the preset restrictions and establish a communication connection with the combat node;
[0097] S2: Multiple combat nodes that simultaneously establish communication connections under the target time slot will be divided into a cell;
[0098] S3: Based on the principle of maximizing the overall QoE value of the cell and the principle of fairness, communication resources are allocated to each combat node in the cell. The principle of fairness indicates that the difference in communication resources obtained between any two combat nodes meets the preset requirements.
[0099] For example, based on the above method applied to combat nodes, when a combat node identifies a target communication node in a target time slot using the aforementioned method, it sends a connection request to the target communication node. Upon receiving the connection request, the target communication node responds and establishes a communication connection with the combat node. When the number of communication connections established by a communication node reaches the maximum number of connections that can be established, a connection relationship matrix centered on that communication node is determined in that time slot. Based on this matrix, the communication node can consider connected combat nodes as a cell, meaning each cell mainly contains one communication node and several operational nodes. The communication node then allocates communication resources to combat nodes within the same cell. Since combat nodes have different resource requirements, the communication node must allocate resources rationally to ensure that each operational node meets its minimum resource requirements (considering fairness) while maximizing the overall QoE value of the cell.
[0100] Specifically, after the combat node connects to the communication node, the communication node can obtain the combat node's status, such as resource requirements. Since the resource allocation problem for each cell is essentially a convex optimization problem, it needs to be solved using optimization algorithms to achieve reasonable resource allocation. To facilitate allocation, in this embodiment, the total communication resources (i.e., bandwidth) of the communication node are discretized into several equal sub-bandwidths (i.e., resource time slots). Therefore, the resource allocation objective is to reasonably allocate these resource time slots to connected operational nodes within the same cell, thereby maximizing the cell's total QoE value.
[0101] To achieve the above objectives, this embodiment proposes a QoE-incrementally aware resource allocation algorithm. This algorithm can efficiently allocate resources to combat nodes in a greedy manner, thereby maximizing the overall QoE value. Simultaneously, by setting a resource allocation condition—that the resource difference between the combat node with the largest resource and the combat node with the smallest resource must meet a threshold requirement—fairness in the resource allocation process is ensured.
[0102] Specifically, communication resources are allocated to each combat node within the cell based on the principles of maximizing the overall QoE value of the cell and fairness, including:
[0103] S4: Divide the allocable communication resources into multiple resource time slots;
[0104] S5: Calculate the QoE increment for each combat node in the cell when it acquires a resource slot;
[0105] S5: Based on the QoE increment of each combat node and the principle of fairness, multiple resource time slots are allocated one by one to complete the allocation of communication resources for each combat node in the cell.
[0106] The method further includes:
[0107] S6: Calculate the QoE value of each combat node in the target time slot and feed it back to the corresponding combat node so that the combat node can select a communication node to establish a communication connection in the next time slot.
[0108] For example, communication node m first identifies the set of combat nodes to be allocated resources through the connection matrix and obtains their status through established communication connections. To maximize the overall QoE value, the communication node needs to allocate the discretized resource slots one by one to the combat node that receives the highest QoE increment after acquiring that resource slot. Therefore, before allocating each resource slot, the communication node needs to calculate which combat node in the cell will receive the largest QoE increment after receiving that resource slot. Specifically, the QoE increment of a combat node n when allocating a resource slot can be calculated using the following formula:
[0109]
[0110] Where λ n If the communication resources currently allocated to combat node n are used, then the communication node will allocate that resource time slot to the combat node with the largest QoE increment, i.e. in Let n be the set of combat nodes, and n be the selected combat nodes for resource allocation. In this way, resources can be efficiently allocated to combat nodes, thereby maximizing the overall QoE value.
[0111] Combining the above methods with the resource requirements of each combat node, communication nodes will allocate resources as uniformly as possible to reduce the probability of combat nodes failing to obtain the resources needed for effective combat missions, ensuring fairness in resource allocation and promoting efficiency. Furthermore, at the beginning of allocating each resource slot, the communication node needs to determine the distribution of resources among the various combat nodes within the cell. If the combat node currently receiving the most resources (i.e., max{O... m,: (t)}) and the combat node that obtains the least resources (i.e., min{O m,: If the current difference between (t) and θ is greater than θ, then communication node m will directly allocate the resource time slot to the combat node with the fewest resources. This ensures the uniformity of resource allocation and maintains the fairness of the resource allocation process.
[0112] Based on the two methods applied to combat nodes and communication nodes respectively, it is evident that their beneficial effects include addressing the challenges of limited communication resources, constantly changing resource requirements, and node locations during the construction of combat communication architecture. This allows combat nodes and communication nodes to dynamically build connections through mutual interaction and learning, enabling them to select new and suitable communication nodes for connection as needed during each combat time slot. Furthermore, communication nodes can allocate communication resources according to the needs and locations of combat nodes. This allocation of communication resources is not only efficient and fair but also strives to maximize the overall QoE value of combat nodes, thereby improving combat effectiveness.
[0113] like Figure 4 As shown, another embodiment of the present invention also provides a flexible construction device for a combat communication system, applied to combat nodes, the device 100 comprising:
[0114] The first determining module is used to determine, under a target time slot, candidate communication nodes that meet the preset constraints among multiple communication nodes. The target time slot is obtained by dividing the combat duration of the combat node by a specified step size. Under different time slots, the positions of the combat nodes are the same or different, the positions of the communication nodes are the same or different, and the determined candidate communication nodes are at least not completely the same.
[0115] The second determining module is used to determine the target communication node from the candidate communication nodes according to the target strategy, and to establish a communication connection with the target communication node. The target strategy is generated based on the UCB method and is used to determine the target communication node based on the first sub-strategy in the first scenario and based on the second sub-strategy in the second scenario.
[0116] The first acquisition module is used to acquire the communication resources allocated by the target communication node.
[0117] In some embodiments, the apparatus further includes:
[0118] The second acquisition module is used to acquire the QoE result of the combat node fed back by the target communication node in the target time slot, wherein the QoE result is determined based on the communication resources required by the combat node and the actual communication resources acquired.
[0119] An update module is used to update the parameters of the target policy based on the QoE result, so that in the next time slot, if the target communication node is selected based on the first sub-policy, the corresponding generated QoE result is improved.
[0120] In some embodiments, determining the target communication node based on the first sub-policy in the first scenario includes:
[0121] If there is no communication node among the candidate communication nodes that has never been selected as a target communication node, the upper confidence bound of each candidate communication node is evaluated based on the number of times each candidate communication node has been selected as a target communication node, the historical QoE results fed back by the candidate communication node, and the standard deviation calculated from the mean of the historical QoE results, respectively.
[0122] The candidate communication node with the highest confidence upper bound is selected as the target communication node for the current time slot.
[0123] In some embodiments, updating the parameters of the target policy based on the QoE result includes:
[0124] Update the stored data of candidate communication nodes selected as target communication nodes in the current time slot, including the number of times a node has been selected as a target communication node, the QoE result of the most recent time node selected as a target communication node, and the standard deviation.
[0125] In some embodiments, determining the target communication node based on the second sub-policy in the second scenario includes:
[0126] If there is a communication node among the candidate communication nodes that has never been selected as the target communication node, then that communication node is directly determined as the target communication node.
[0127] In some embodiments, the preset limiting conditions include at least one or more of the following conditions:
[0128] The communication distance coverage condition, the maximum number of connections a communication node can make to an operating node, and the unique communication condition that restricts an operating node to only one communication node in each time slot;
[0129] When determining the target communication node based on the first sub-strategy, the following is also included:
[0130] If the number of connections established by the candidate communication node corresponding to the highest confidence upper bound already meets the maximum number of connections condition, then the candidate communication node with the second highest confidence upper bound is selected as the target communication node.
[0131] like Figure 5 As shown, another embodiment of the present invention also provides a flexible construction device 200 for a combat communication system, applied to a communication node, the device comprising:
[0132] The response module is used to respond to the connection requests of combat nodes that meet preset restrictions in the target time slot and establish a communication connection with the combat nodes.
[0133] The partitioning module is used to divide multiple combat nodes that simultaneously establish the communication connection under the target time slot into a cell;
[0134] The allocation module is used to allocate communication resources to each combat node in the cell according to the principle of maximizing the overall QoE value of the cell and the fairness principle, wherein the fairness principle indicates that the difference in communication resources obtained between any two combat nodes meets the preset requirements.
[0135] In some embodiments, allocating communication resources to each combat node within the cell based on the principles of maximizing the overall QoE value of the cell and fairness includes:
[0136] Allocable communication resources are divided into multiple resource time slots;
[0137] Calculate the QoE increment for each combat node within the cell when it acquires a resource time slot;
[0138] Based on the QoE increment of each combat node and the principle of fairness, the multiple resource time slots are allocated one by one to complete the allocation of communication resources for each combat node in the cell.
[0139] The device further includes:
[0140] The statistics module is used to count the QoE value of each combat node under the target time slot and feed it back to the corresponding combat node so that the combat node can select a communication node to establish a communication connection in the next time slot.
[0141] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.
[0142] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.
Claims
1. A method for constructing a combat communication system, applied to combat nodes, characterized in that, The method includes: Under a target time slot, candidate communication nodes that meet the preset constraints are determined from multiple communication nodes. The target time slot is obtained by dividing the combat duration of the combat node by a specified step size. Under different time slots, the positions of the combat nodes are the same or different, the positions of the communication nodes are the same or different, and the determined candidate communication nodes are at least not completely the same. Based on the target strategy, a target communication node is determined from the candidate communication nodes, and a communication connection is established with the target communication node. The target strategy is generated based on the UCB method and is used to determine the target communication node based on the first sub-strategy in the first scenario and to determine the target communication node based on the second sub-strategy in the second scenario. Obtain the communication resources allocated by the target communication node; The method further includes: Obtain the QoE result of the combat node in the target time slot fed back by the target communication node, the QoE result being determined based on the communication resources required by the combat node and the actual communication resources obtained; The parameters of the target policy are updated based on the QoE result so that, in the next time slot, if the target communication node is selected based on the first sub-policy, the generated QoE result is improved.
2. The method for constructing a combat communication system according to claim 1, characterized in that, The step of determining the target communication node based on the first sub-strategy in the first scenario includes: If there is no communication node among the candidate communication nodes that has never been selected as a target communication node, the upper confidence bound of each candidate communication node is evaluated based on the number of times each candidate communication node has been selected as a target communication node, the historical QoE results fed back by the candidate communication node, and the standard deviation calculated from the mean of the historical QoE results, respectively. The candidate communication node with the highest confidence upper bound is selected as the target communication node for the current time slot.
3. The method for constructing a combat communication system according to claim 2, characterized in that, The parameters for updating the target policy based on the QoE result include: Update the stored data of candidate communication nodes selected as target communication nodes in the current time slot, including the number of times a node has been selected as a target communication node, the QoE result of the most recent time node selected as a target communication node, and the standard deviation.
4. The method for constructing a combat communication system according to claim 1, characterized in that, The determination of the target communication node based on the second sub-strategy in the second scenario includes: If there is a communication node among the candidate communication nodes that has never been selected as the target communication node, then that communication node is directly determined as the target communication node.
5. The method for constructing a combat communication system according to claim 2, characterized in that, The preset restrictions include at least one or more of the following conditions: The communication distance coverage condition, the maximum number of connections a communication node can make to an operating node, and the unique communication condition that restricts an operating node to only one communication node in each time slot; When determining the target communication node based on the first sub-strategy, the following is also included: If the number of connections established by the candidate communication node corresponding to the highest confidence upper bound already meets the maximum number of connections condition, then the candidate communication node with the second highest confidence upper bound is selected as the target communication node.
6. A method for constructing a combat communication system, applied to communication nodes, characterized in that, The method includes: Under the target time slot, respond to the connection request of the combat node that meets the preset restrictions and establish a communication connection with the combat node; The multiple combat nodes that simultaneously establish the communication connection under the target time slot are divided into a cell; Based on the principle of maximizing the overall QoE value of the cell and the principle of fairness, communication resources are allocated to each combat node in the cell. The principle of fairness indicates that the difference in communication resources obtained between any two combat nodes meets the preset requirements. The allocation of communication resources to each combat node within the cell based on the principles of maximizing the overall QoE value of the cell and fairness includes: Allocable communication resources are divided into multiple resource time slots; Calculate the QoE increment for each combat node within the cell when it acquires a resource time slot; Based on the QoE increment of each combat node and the principle of fairness, the multiple resource time slots are allocated one by one to complete the allocation of communication resources for each combat node in the cell. The method further includes: The QoE value of each combat node under the target time slot is calculated and fed back to the corresponding combat node so that the combat node can select a communication node to establish a communication connection in the next time slot.
7. A device for constructing an operational communication system, applied to operational nodes, characterized in that, The device includes: The first determining module is used to determine, under a target time slot, candidate communication nodes that meet the preset constraints among multiple communication nodes. The target time slot is obtained by dividing the combat duration of the combat node by a specified step size. Under different time slots, the positions of the combat nodes are the same or different, the positions of the communication nodes are the same or different, and the determined candidate communication nodes are at least not completely the same. The second determining module is used to determine the target communication node from the candidate communication nodes according to the target strategy, and to establish a communication connection with the target communication node. The target strategy is generated based on the UCB method and is used to determine the target communication node based on the first sub-strategy in the first scenario and based on the second sub-strategy in the second scenario. The first obtaining module is used to obtain the communication resources allocated by the target communication node; The device further includes: The second acquisition module is used to acquire the QoE result of the combat node fed back by the target communication node in the target time slot, wherein the QoE result is determined based on the communication resources required by the combat node and the actual communication resources acquired. An update module is used to update the parameters of the target policy based on the QoE result, so that in the next time slot, if the target communication node is selected based on the first sub-policy, the corresponding generated QoE result is improved.
8. A device for constructing an operational communication system, applied to a communication node, characterized in that, The device includes: The response module is used to respond to the connection requests of combat nodes that meet preset restrictions in the target time slot and establish a communication connection with the combat nodes. The partitioning module is used to divide multiple combat nodes that simultaneously establish the communication connection under the target time slot into a cell; The allocation module allocates communication resources to each combat node in the cell based on the principle of maximizing the overall QoE value of the cell and the principle of fairness. The principle of fairness indicates that the difference in communication resources obtained between any two combat nodes meets the preset requirements. The allocation of communication resources to each combat node within the cell based on the principles of maximizing the overall QoE value of the cell and fairness includes: Allocable communication resources are divided into multiple resource time slots; Calculate the QoE increment for each combat node within the cell when it acquires a resource time slot; Based on the QoE increment of each combat node and the principle of fairness, the multiple resource time slots are allocated one by one to complete the allocation of communication resources for each combat node in the cell. The device further includes: The statistics module is used to count the QoE value of each combat node under the target time slot and feed it back to the corresponding combat node so that the combat node can select a communication node to establish a communication connection in the next time slot.
Citation Information
Patent Citations
Online dispatching method and system for edge computing task
CN112799823A