A method for optimizing radio frequency front-end for strong interference-resistant ground-penetrating communication
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-11
AI Technical Summary
[0002]随着地下物联网技术的快速发展,其在安全生产监测与环境感知领域的应用日益深入,例如在煤矿安全监测传感器网络、金属矿开采环境监控系统以及智慧农业中的土壤墒情与养分长期监测网络等场景中,这些网络通常由大规模、低成本、电池供电的传感节点构成,并被长期部署于地下或贴近地面的复杂介质环境中,节点面临的电磁环境异常复杂:一方面,地层介质因季节更替、含水量变化而呈现显著且缓慢时变的电特性,导致信号传播路径损耗与多径效应不断变化;另一方面,应用现场常见的大型机电设备,如矿用采掘机、运输机或农业灌溉电机,其启停与运行会周期性地产生强烈的宽频谱电磁干扰,这些因素共同构成了一个动态、非平稳且资源苛刻的无线传输环境,对节点的可靠通信与持久续航构成了双重挑战
[0014]本发明的有益效果是:通过分布式节点间的协同学习与自主优化,显著提升了复杂电磁环境下透地通信链路的整体性能与可靠性,其核心在于使每个节点不仅能依据本地环境实时优化自身射频前端策略,还能通过策略协同、选择性知识迁移与演化预警,与邻居节点形成高效协作,这有效抑制了共信道干扰,降低了网络总体能耗,延长了节点续航,同时,方案赋予网络持续适应环境慢变与抗未知干扰的强鲁棒性,实现了从个体智能到群体协同的跃升,保障了大规模地下物联网长期稳定运行。
Smart Images

Figure CN122340517B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of radio frequency front-end optimization for through-ground communication, and more specifically, to a method for optimizing radio frequency front-ends for through-ground communication that is resistant to strong interference. Background Technology
[0002] With the rapid development of underground IoT technology, its application in safety production monitoring and environmental perception is becoming increasingly in-depth. For example, in scenarios such as coal mine safety monitoring sensor networks, metal mining environmental monitoring systems, and long-term soil moisture and nutrient monitoring networks in smart agriculture, these networks are usually composed of large-scale, low-cost, battery-powered sensor nodes, which are deployed for a long time in complex media environments underground or close to the ground. The electromagnetic environment faced by the nodes is extremely complex: on the one hand, the geological medium exhibits significant and slowly time-varying electrical characteristics due to seasonal changes and changes in water content, resulting in constantly changing signal propagation path loss and multipath effects; on the other hand, large electromechanical equipment commonly used in the application field, such as mining excavators, transport machines, or agricultural irrigation motors, will periodically generate strong broadband electromagnetic interference during their start-up, shutdown, and operation. These factors together constitute a dynamic, non-stationary, and resource-intensive wireless transmission environment, posing a dual challenge to the reliable communication and long-term battery life of the nodes.
[0003] Under severely limited resources, the anti-interference design of existing through-ground communication nodes faces a serious performance bottleneck. Currently, a common technical solution is to use an energy detection mechanism based on a fixed threshold. When the received signal strength exceeds a certain preset value, it is determined that strong interference has occurred, and the node then switches to one of several preset "strong interference operating modes," such as narrowing the channel bandwidth, enhancing analog filtering, or reducing transmission power. Another commonly used strategy is to design fixed sleep and listening duty cycles to reduce average power consumption. However, these methods have inherent drawbacks: First, the types of preset interference modes are limited and cannot cover the infinite combinations of interference intensity, spectral characteristics, and time-domain patterns in the actual environment. Especially when the interference spectrum partially overlaps with the useful signal, simple mode switching often leads to a sharp deterioration in communication performance or an increase in ineffective power consumption. Second, a fixed sleep and listening cycle cannot keep up with dynamically changing interference. Matching the scenario with actual business traffic, ineffective listening during interference intervals may waste energy, or critical data may be lost due to dormancy during business surges. More fundamentally, existing solutions lack the ability to learn and utilize long-term, predictable interference patterns in the environment. Their decision-making is instantaneous, passive, and short-sighted, unable to extract the temporal correlation, statistical characteristics, and mapping relationship with node energy efficiency from historical operational data, thus failing to make forward-looking strategy adjustments. This often results in nodes maintaining basic connectivity at the expense of unnecessary energy consumption during long-term operation, leading to battery life far below theoretical expectations, high network maintenance costs, and threats to the continuity and reliability of monitoring tasks. Therefore, there is an urgent need for an RF front-end anti-interference strategy optimization method that can adapt to long-term environmental changes, has autonomous learning and decision-making capabilities, and meets low power consumption constraints. Summary of the Invention
[0004] This invention addresses the technical problems existing in the prior art by providing an optimization method for radio frequency front-end of strong interference-resistant ground-penetrating communication, thereby solving the problems mentioned in the background art.
[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: specifically, it includes the following steps: Step S1: Each node continuously collects local environmental observation data, which includes at least the instantaneous signal-to-interference-plus-noise ratio and the operating parameters of the local radio frequency front-end. At the same time, it receives policy identifiers from neighboring nodes that have direct communication connections with it. Each node maintains a policy model for deciding on the operating parameters of the radio frequency front-end and, based on a preset utility evaluation model, performs a quantitative evaluation of the radio frequency front-end anti-interference strategy currently adopted by the policy model to generate a local policy utility value. Step S2: Each node calculates the policy gradient term based on its local policy utility value, and infers the local policy distribution characteristics of its neighboring nodes' policy models based on the neighboring nodes' policy identifiers. Each node constructs a composite loss function, which includes a policy gradient term and a policy consistency regularization term. The policy consistency regularization term is used to constrain the difference between the local policy distribution output by the node's policy model and the local policy distribution characteristics of its neighboring nodes' policy models. Each node updates the parameters of its policy model by minimizing the composite loss function. Step S3: Each node periodically compares the local policy distribution of its policy model on a preset state set with the local policy distribution of the policy models of neighboring nodes corresponding to the policy identifiers of neighboring nodes on the same preset state set, and calculates the policy similarity. When the policy similarity is higher than the preset first similarity threshold, and the historical utility record corresponding to the policy identifier of the neighboring node, recorded locally or obtained from the neighboring node, is better than the current local policy utility value of the current node, the current node obtains the policy soft label output by the policy model of the corresponding neighboring node, and uses the policy soft label to perform knowledge distillation training on the policy model of the current node. Step S4: After applying the updated current policy model in step S2 for one time period, each node compares the actual measured communication performance indicators with the local policy utility value predicted by the utility evaluation model to generate a policy confidence score. If the policy confidence score is lower than the preset confidence threshold for multiple consecutive time periods, the current node triggers the evolution update of its policy model and sends a model evolution notification signal containing an environmental feature summary to neighboring nodes whose policy similarity exceeds the preset second similarity threshold. In a preferred embodiment, in step S1, local environmental observation data is continuously collected, including the instantaneous signal-to-interference-plus-noise ratio and the operating parameters of the local radio frequency front end. The operating parameters include at least the transmit power, filter bandwidth and filter center frequency. At the same time, the percentage of remaining energy of the node and the rate of change of recent packet reception rate are collected. The instantaneous signal-to-interference-plus-noise ratio (SIR) is compressed logarithmically to the base of the natural logarithm to obtain the SIR characteristic value. The values of transmit power, filter bandwidth, and filter center frequency are mapped to a closed interval between zero and one to obtain the normalized transmit power, normalized filter bandwidth, and normalized center frequency. The signal-to-interference-plus-noise ratio (SIR) eigenvalues, normalized transmit power, normalized filter bandwidth, normalized center frequency, and percentage of remaining node energy are combined in sequence to form an environmental state feature vector.
[0006] In a preferred embodiment, the process of receiving the neighbor node policy identifier from other nodes with which it has a direct communication connection specifically includes: Nodes synchronously receive neighbor node policy identifiers periodically broadcast by neighbor nodes within their one-hop communication range. The neighbor node policy identifier is a hash digest value representing key information of the policy model maintained by the neighbor node. Nodes associate the received neighbor node policy identifier with the identity information of the neighbor node that sent the neighbor node policy identifier and store it in their local neighbor policy information table.
[0007] In a preferred embodiment, the specific process of generating a local policy utility value through quantitative evaluation based on a preset utility evaluation model is as follows: First, calculate the communication quality components: subtract a preset demodulation threshold from the signal-to-interference-plus-noise ratio characteristic value, multiply the difference by a signal quality adjustment coefficient, calculate its arctangent value, and then multiply it by a first constant for scaling to obtain the communication quality components. Secondly, calculate the link instability penalty component: obtain the rate of change of the recent packet reception rate, calculate the square of the rate of change and take the negative, multiply the obtained negative square value by a fluctuation sensitivity coefficient, calculate the exponential function value with the natural constant as the base, and subtract this exponential function value from one to obtain the link instability penalty component. Next, calculate the energy cost component: calculate the exponential function value after multiplying the node's remaining energy percentage by a negative energy influence factor, and multiply the normalized transmit power by this exponential function value to obtain the energy cost component; Next, a weighted merging process is performed: the communication quality component is multiplied by the communication quality weight, the difference after subtracting the link instability penalty component is multiplied by the link stability weight, the two products are added together, and then multiplied by the difference after subtracting a global performance trade-off factor to obtain the weighted performance result; the energy cost component is multiplied by the global performance trade-off factor to obtain the weighted energy cost result. Finally, calculate the local policy utility value: subtract the weighted energy consumption cost value from the weighted performance value to obtain the local policy utility value.
[0008] In a preferred embodiment, step S2, which involves calculating the policy gradient term based on the local policy utility value and inferring the local policy distribution characteristics of the neighboring nodes' policy models based on the neighboring node policy identifiers, specifically comprises the following steps: The node randomly selects a batch of data consisting of multiple environmental state feature vectors from the local stored history. For each environmental state feature vector in the batch of data, the node inputs it into its own maintained policy model to obtain the local policy distribution output by the policy model for different combinations of RF front-end operating parameters. At the same time, the node retrieves the local policy utility value recorded when the policy was previously executed from the historical records, corresponding to each environmental state feature vector; Based on the environmental state feature vector, the corresponding local policy distribution, and the local policy utility value in the batch data, the node calculates the policy gradient term. The calculation process for the policy gradient term is as follows: For each environmental state feature vector in the batch data, an action is sampled according to the current local policy distribution. The gradient of the log probability of the action with respect to the policy model parameters is calculated. This gradient is multiplied by the local policy utility value corresponding to the environmental state feature vector to obtain the contribution gradient under the environmental state feature vector. The average of the contribution gradients under all environmental state feature vectors is taken, and the resulting average gradient vector is the policy gradient term.
[0009] In a preferred embodiment, the specific process of constructing a composite loss function for each node is as follows: The node retrieves the neighbor policy identifiers of all neighboring nodes from its maintained neighbor policy information table; the node then uses preset decoding rules to restore each neighbor policy identifier to an approximate local policy distribution feature. Each node selects a preset state set as a comparison benchmark. For each environmental state feature vector in the preset state set, the node calculates the distribution difference between the local policy distribution output by its own policy model and the approximate local policy distribution features of each neighboring node. The calculation process of the distribution difference metric is as follows: For each environmental state feature vector in the preset state set, two first actions are sampled from the local policy distribution, the first kernel function value between the two first actions is calculated and the average value is obtained to obtain the first expected value; two second actions are sampled from the approximate local policy distribution features of the neighboring nodes, the second kernel function value between the two second actions is calculated and the average value is obtained to obtain the second expected value; a third action is sampled from the local policy distribution, and a fourth action is sampled from the approximate local policy distribution features of the neighboring nodes, the third kernel function value between the third action and the fourth action is calculated and the average value is obtained to obtain the third expected value; the first expected value and the second expected value are added together, and then twice the third expected value is subtracted. The result is the generalized energy distance between the local policy distribution and the approximate local policy distribution features of the neighboring node under the current environmental state feature vector. The average of the distribution difference measures corresponding to all neighboring nodes calculated on the preset state set is averaged again and used as the policy consistency regularization term. Next, a composite loss function is constructed: First, the expected local policy utility value is calculated, which is the average of the local policy utility values corresponding to all environmental state feature vectors in the batch data; the negative expected local policy utility value is added to the product of an adjustable regularization strength coefficient and a policy consistency regularization term to form the composite loss function. Finally, the parameters of the policy model are updated by subtracting the product of the learning rate and the gradient of the composite loss function with respect to the policy model parameters from the current values of the policy model parameters.
[0010] In a preferred embodiment, in step S3, the process of periodically comparing the local policy distribution of the current node's policy model on a preset state set with the local policy distribution of the neighbor nodes' policy models on the same preset state set based on the neighbor node's policy identifier, and calculating the policy similarity, specifically involves: In each calculation cycle, the node first uses its own maintained strategy model to perform forward calculation on each environmental state feature vector in the preset state set to obtain the local strategy distribution of the node for all optional RF front-end working parameter combinations under the current environmental state feature vector, and records it to form the local strategy distribution set of the node. At the same time, the node sends a query request to each neighbor node recorded in the neighbor policy information table, requesting to obtain the local policy distribution of that neighbor node on the same preset state set; the node receives the response from the neighbor node and obtains its corresponding neighbor local policy distribution set. For each neighboring node, the node calculates policy similarity based on its own local policy distribution set and the neighboring node's neighboring local policy distribution set. The calculation process for policy similarity is as follows: For each environmental state feature vector in the preset state set, the information entropy of the node's local policy distribution under that environmental state feature vector is calculated. The information entropy is the sum of each probability value in the distribution multiplied by the base-2 logarithm of that probability value, and the negative value is taken after summing all the products. Then, an exponential function with the natural constant as the base is calculated, and its exponent is the negative value of the information entropy, to obtain the decision confidence weight of the node on this environmental state feature vector. Next, the generalized Jensen-Shannon divergence between the two local policy distributions of the node and its neighbors under this environmental state feature vector is calculated. The calculation process is as follows: the decision confidence weight is used as the first coefficient, and the first coefficient is subtracted from the number one to obtain the second coefficient. The first intermediate distribution is calculated, which is the sum of the first coefficient multiplied by the local policy distribution of the node and the second coefficient multiplied by the logarithm of the logarithm. The sum of the local policy distributions of neighboring nodes; multiplying the local policy distribution of this node by the KL divergence between the local policy distribution and the first intermediate distribution using a first coefficient, and adding the second coefficient multiplied by the KL divergence between the local policy distribution of neighboring nodes and the first intermediate distribution. The KL divergence is calculated by multiplying the value of the element in the first probability distribution for each element in the probability distribution by the base-2 logarithm of that value, then dividing by the value of the corresponding element in the second probability distribution, and summing all the results; then, the calculated generalized Jensen-Shannon divergence is weighted by the decision confidence weights to obtain the weighted divergence value under the environmental state feature vector; the average of the weighted divergence values under all environmental state feature vectors in the preset state set is calculated, specifically by summing all weighted divergence values and summing all decision confidence weights, and dividing the sum of weighted divergence values by the sum of decision confidence weights to obtain the second average; the square root of the second average is taken, and finally, the square root result is subtracted from the number one to obtain the policy similarity with the neighboring node.
[0011] In a preferred embodiment, the process of obtaining policy soft labels from the corresponding neighboring nodes and performing knowledge distillation training when the policy similarity is higher than a preset first similarity threshold and the historical utility record corresponding to the policy identifier of the neighboring node is better than the current local policy utility value of the current node specifically includes: A node calculates its policy similarity with a neighboring node and compares it with a preset first similarity threshold. Simultaneously, it retrieves the historical utility record associated with that neighboring node from the neighbor policy information table. This historical utility record is the moving average of the neighboring node's recent local policy utility value. The node then compares the retrieved historical utility record of the neighboring node with its own recent local policy utility value moving average, based on the following criteria: The calculation determines whether the calculated policy similarity is greater than the first similarity threshold and whether the moving average of the historical utility records of the neighboring nodes is greater than the moving average of the historical utility records of this node plus a preset performance margin value. The condition is satisfied if and only if both of the above conditions are satisfied at the same time. The node then sends a knowledge distillation request to the neighboring node. After receiving the request, the neighboring node sends its local policy distribution set on the preset state set as a policy soft label back to the current node. After the current node receives the policy soft label, it starts knowledge distillation training: Each node inputs a preset state set into its own policy model to obtain the current local policy distribution set, and simultaneously extracts the corresponding neighbor local policy distributions from the received policy soft tags. A temperature parameter is introduced. For each environmental state feature vector in the preset state set, the node's current local policy distribution and the corresponding neighbor local policy distribution extracted from the policy soft tags are processed as follows: each probability value in the probability distribution is divided by the temperature parameter, and then an exponential function with the natural constant as its base is calculated, where the exponent is each probability value divided by the temperature parameter, resulting in intermediate exponent values. All intermediate exponent values are summed to obtain a normalized denominator. Each intermediate exponent value is divided by the normalized denominator to obtain a new probability distribution after temperature scaling and normalization. Then, for each environmental state feature vector in the preset state set, the KL divergence between the new probability distribution of neighbor nodes after temperature scaling and normalization and the new probability distribution of the node after temperature scaling and normalization is calculated, and this is used as the knowledge distillation loss component under that environmental state feature vector. The average value of the knowledge distillation loss components under all environmental state feature vectors is calculated to obtain the final knowledge distillation loss. Nodes update the parameters of their own policy models by performing one or more steps of gradient descent algorithm to minimize knowledge distillation loss.
[0012] In a preferred embodiment, step S4, which compares the actually measured communication performance indicators with the local policy utility value predicted by the utility evaluation model to generate a policy confidence score, specifically involves: Within a preset sliding time window, based on the environmental observation data continuously collected in step S1, the node calculates the actual measured communication performance index sequence for each time period. The actual measured communication performance index is a scalar value obtained by weighted summation of the average throughput, average packet error rate and average energy consumption within the period. At the same time, the local policy utility value predicted by the utility evaluation model based on the current environmental state is collected at the beginning of each time period. For each time period, the difference between the actual measured value and the predicted local policy utility value is calculated to obtain the predicted residual sequence; the predicted residual sequence is then weighted and averaged over a sliding time window to obtain the weighted average residual. Simultaneously, the normalized median absolute deviation of the predicted residual sequence is calculated. The calculation process of the normalized median absolute deviation is as follows: first, the median of the entire residual sequence is calculated; then, the absolute value of the difference between each residual and the median is calculated to obtain the absolute deviation sequence; then, the median of the absolute deviation sequence is calculated; finally, the median is divided by a constant 0.6745. The absolute value of the weighted average residual is subtracted from the product of a significance level coefficient and the normalized median absolute deviation to obtain the deviation detection value. The deviation detection value is input into the S-shaped growth function for mapping, and its output value is between zero and one, and has a scaling factor that controls the steepness of the function curve. The mapping result is then subtracted from the number one to obtain the policy confidence score.
[0013] In a preferred embodiment, the process of triggering the evolution and update of the policy model of the current node and sending a model evolution notification signal containing an environmental feature summary to neighboring nodes whose policy similarity exceeds a preset second similarity threshold specifically includes: When the policy confidence score is lower than the confidence threshold for three consecutive time periods, the node starts with the parameters of the current policy model and performs model evolution based on adaptive Cauchy mutation. Nodes use recently collected, unused mini-batch environment state feature vectors as a validation set to quickly evaluate each candidate variant. The evaluation method involves inputting the environment state feature vectors from the validation set into the candidate variant to obtain the recommended action, then calculating the predicted utility of the action through a utility evaluation model, and taking the average predicted utility of all validation samples. The candidate variant with the highest average predicted utility is selected as the parameter for the next-generation policy model. After completing its own model evolution, the node selects all neighboring nodes whose policy similarity exceeds a preset second similarity threshold from the stored policy similarity records, and unicasts a model evolution announcement signal to these neighboring nodes. The model evolution announcement signal contains an environment feature summary, which is a statistical summary of the environment state feature vectors within the most recent time window before triggering the evolution, including the mean and principal components of each dimension.
[0014] The beneficial effects of this invention are as follows: through collaborative learning and autonomous optimization among distributed nodes, the overall performance and reliability of ground-penetrating communication links in complex electromagnetic environments are significantly improved. The core of this invention is that each node can not only optimize its own radio frequency front-end strategy in real time according to the local environment, but also form efficient cooperation with neighboring nodes through strategy collaboration, selective knowledge transfer and evolution warning. This effectively suppresses co-channel interference, reduces the overall energy consumption of the network, and extends the node's endurance. At the same time, the solution endows the network with strong robustness to continuously adapt to slow environmental changes and resist unknown interference, realizing a leap from individual intelligence to group collaboration, and ensuring the long-term stable operation of large-scale underground Internet of Things. Attached Figure Description
[0015] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0018] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0019] Example 1 This embodiment provides, for example Figure 1 The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication includes the following steps: Step S1: Each node continuously collects local environmental observation data, which includes at least the instantaneous signal-to-interference-plus-noise ratio and the operating parameters of the local radio frequency front-end. At the same time, it receives policy identifiers from neighboring nodes that have direct communication connections with it. Each node maintains a policy model for deciding on the operating parameters of the radio frequency front-end and, based on a preset utility evaluation model, performs a quantitative evaluation of the radio frequency front-end anti-interference strategy currently adopted by the policy model to generate a local policy utility value. Step S2: Each node calculates the policy gradient term based on its local policy utility value, and infers the local policy distribution characteristics of its neighboring nodes' policy models based on the neighboring nodes' policy identifiers. Each node constructs a composite loss function, which includes a policy gradient term and a policy consistency regularization term. The policy consistency regularization term is used to constrain the difference between the local policy distribution output by the node's policy model and the local policy distribution characteristics of its neighboring nodes' policy models. Each node updates the parameters of its policy model by minimizing the composite loss function. Step S3: Each node periodically compares the local policy distribution of its policy model on a preset state set with the local policy distribution of the policy models of neighboring nodes corresponding to the policy identifiers of neighboring nodes on the same preset state set, and calculates the policy similarity. When the policy similarity is higher than the preset first similarity threshold, and the historical utility record corresponding to the policy identifier of the neighboring node, recorded locally or obtained from the neighboring node, is better than the current local policy utility value of the current node, the current node obtains the policy soft label output by the policy model of the corresponding neighboring node, and uses the policy soft label to perform knowledge distillation training on the policy model of the current node. Step S4: After applying the updated current policy model from step S2 for one time period, each node compares the actual measured communication performance indicators with the local policy utility value predicted by the utility evaluation model to generate a policy confidence score. If the policy confidence score is lower than a preset confidence threshold for multiple consecutive time periods, the current node triggers the evolution update of its policy model and sends a model evolution notification signal containing an environmental feature summary to neighboring nodes whose policy similarity exceeds a preset second similarity threshold.
[0020] In this embodiment, it is specifically necessary to explain step S1, which involves continuously collecting local environmental observation data, including collecting instantaneous signal-to-interference-plus-noise ratio (SIR) and local RF front-end operating parameters. The operating parameters include at least transmit power, filter bandwidth, and filter center frequency. Simultaneously, the percentage of remaining node energy and the rate of change of recent packet reception rate are collected. The instantaneous SIR can be calculated from the received signal strength indicator and the noise floor power. The transmit power, filter bandwidth, and filter center frequency are real-time configuration values of the RF front-end programmable device. The percentage of remaining node energy is provided by the power management unit. The rate of change of recent packet reception rate is obtained by statistically analyzing the number of successfully received data packets within a past time window and calculating its percentage change relative to the previous equal-length window. The time window length can be configured to be 10 to 100 communication cycles. The instantaneous signal-to-interference-plus-noise ratio (SNR) is subjected to logarithmic compression with the natural logarithm as the base to obtain SNR eigenvalues. The logarithmic compression is achieved using the formula log(1+SINR), where SINR is the instantaneous SNR. Adding a constant of 1 avoids taking the logarithm of zero and maps the typical linear SNR range (e.g., 0-30dB) to an approximately linear feature space. This enhances the distinguishability of the low SNR region while preventing the explosive growth of eigenvalues in the high SNR region. The values of transmit power, filter bandwidth, and filter center frequency are respectively... Mapping to a closed interval between zero and one yields normalized transmit power, normalized filter bandwidth, and normalized center frequency. The mapping is achieved by dividing by the maximum value allowed by each hardware component. For example, normalized transmit power = current transmit power / maximum allowed transmit power; normalized filter bandwidth = current bandwidth / maximum adjustable bandwidth; normalized center frequency = (current center frequency - minimum adjustable frequency) / (maximum adjustable frequency - minimum adjustable frequency). This normalization process eliminates the influence of different dimensions, providing a unified scale input for subsequent strategy models. The signal-to-interference-plus-noise ratio (SINR) eigenvalue, normalized transmit power, normalized filter bandwidth, normalized center frequency, and the percentage of remaining energy of the node are combined in sequence to form an environmental state feature vector, which is used by the policy model and the utility evaluation model. The combined environmental state feature vector is a five-dimensional vector that comprehensively represents the current channel quality, self-configuration status, and energy level of the node, and serves as the standardized input basis for policy decision-making and utility evaluation. The process of receiving policy identifiers from neighboring nodes with which it has a direct communication connection is as follows: Nodes synchronously receive neighbor node policy identifiers periodically broadcast by neighbor nodes within their one-hop communication range. The interval between periodic broadcasts can be set to 1 to 10 policy update cycles to achieve a balance between communication overhead and information freshness. The neighbor node policy identifier is a hash digest value representing the key information of the policy model maintained by the neighbor node. The hash digest value can be calculated using cryptographic hash functions such as SHA-256. Its input is all trainable parameter values of the neighbor node's policy model at the current time, or the output probability distribution vector of the model on a set of standard environment state vectors. This digest value has fixed length and collision resistance characteristics, and can uniquely represent the current state of the policy model with minimal communication overhead. Nodes associate the received neighbor node policy identifier with the identity information of the neighbor node that sent the policy identifier and store it in their local neighbor policy information table. The neighbor policy information table is a dynamically updated data structure that contains at least the fields "neighbor node ID", "policy identifier", "last update timestamp", and "associated historical utility record". This neighbor policy information table is the basic database for nodes to perform policy coordination, similarity calculation, and knowledge transfer. The specific process of generating a local policy utility value through quantitative evaluation based on a pre-defined utility evaluation model is as follows: First, calculate the communication quality component: subtract a preset demodulation threshold from the signal-to-interference-plus-noise ratio (SINR) characteristic value, multiply the difference by a positive signal quality adjustment coefficient, calculate its arctangent value, and then multiply it by a first constant for scaling to obtain the communication quality component. The preset demodulation threshold is determined according to the modulation and coding scheme used. For example, it can be set to 8dB under QPSK modulation. The signal quality adjustment coefficient is used to control the sensitivity of the function near the demodulation threshold and can be set to 0.5. The first constant is 2 / π, which is used to scale the output range of the arctangent function from (-π / 2, π / 2) to (-1, 1). The final value range of the communication quality component is approximately (-1, 1). This arctangent function form makes the evaluation value change significantly when the SINR is close to the threshold, while the evaluation value tends to saturate when the SINR is much higher or lower than the threshold, which is consistent with the nonlinear characteristics of the communication system. Secondly, calculate the link instability penalty component: obtain the recent rate of change of packet reception rate, calculate the square of the rate of change and take the negative, multiply the resulting negative square value by a positive fluctuation sensitivity coefficient, calculate the exponential function value with the natural constant as the base, and subtract this exponential function value from one to obtain the link instability penalty component. The fluctuation sensitivity coefficient controls the penalty intensity for link fluctuations and can be set to 10. When the rate of change is zero, the penalty component is close to 0; when the rate of change increases, the penalty component quickly approaches 1. This design enables the utility evaluation to effectively penalize large jitters in the communication link and encourages the selection of strategies that can maintain a stable connection. Next, the energy cost component is calculated: the exponential function value obtained by multiplying the percentage of remaining node energy by a negative energy influence factor is calculated, and the normalized transmit power is multiplied by this exponential function value to obtain the energy cost component. The negative energy influence factor can be set to -3. This design makes the energy cost component proportional to the normalized transmit power, but it is adjusted by the negative exponential decay of the remaining energy. When the node energy is sufficient, the exponential term is close to 1, and the cost mainly depends on the power. When the node energy is scarce, the exponential term decreases sharply, resulting in a very high cost assessment for the same transmit power, thereby guiding the strategy to prioritize energy saving when the power is low. Next, a weighted merging process is performed: the communication quality component is multiplied by the communication quality weight, and the difference between one and the link instability penalty component is multiplied by the link stability weight. These two products are added together, and then multiplied by the difference between one and a global performance trade-off factor to obtain the weighted performance result. The energy cost component is multiplied by the global performance trade-off factor to obtain the weighted energy cost result. The sum of the communication quality weight and the link stability weight is one. The communication quality weight and the link stability weight can be set according to the application scenario. For example, in security monitoring requiring high reliability, they can be set to 0.7 and 0.3 respectively. The global performance trade-off factor is used to dynamically adjust the priority of performance and energy consumption throughout the system's lifecycle. Its value range is [0,1]. In the early stages of network deployment, it can be set to a smaller value (such as 0.2) to prioritize performance; in the later stages of network operation or when energy is scarce, it can be increased to 0.5 or higher to prioritize energy saving. Finally, the local policy utility value is calculated: the weighted energy consumption cost result is subtracted from the weighted performance result to obtain the local policy utility value. The final local policy utility value is a dimensionless scalar. The higher the value, the better the overall performance of the RF front-end configuration recommended by the current policy model in terms of communication performance, link stability and energy efficiency under the current environment. This value is directly used as the basis for calculating the policy gradient term in step S2.
[0021] In this embodiment, it is specifically necessary to explain that in step S2, the process of calculating the policy gradient term based on the local policy utility value and inferring the local policy distribution characteristics of the neighboring nodes' policy models based on the neighboring node policy identifiers is as follows: The node randomly selects a batch of data consisting of multiple environmental state feature vectors from the locally stored historical records. The size of the batch data can be set to 32, 64, or 128 to achieve a balance between computational efficiency and gradient estimation stability. For each environmental state feature vector in the batch data, the node inputs it into its own maintained policy model to obtain the local policy distribution output by the policy model for different combinations of RF front-end operating parameters. This local policy distribution is a probability density distribution over all discrete or continuous action spaces. At the same time, the node retrieves the local policy utility value recorded when the policy was previously executed from the historical records, corresponding to each environmental state feature vector; Based on the environmental state feature vector, the corresponding local policy distribution, and the local policy utility value in the batch data, the node calculates the policy gradient term. The calculation process for the policy gradient term is as follows: For each environment state feature vector in the batch data, an action is sampled according to the current local policy distribution. The sampling method is probability sampling, where high-probability actions are more likely to be selected. The gradient of the log probability of this action with respect to the policy model parameters is calculated. This multiplication operation enhances the gradient direction corresponding to high utility values and suppresses the gradient direction corresponding to low utility values. This gradient is multiplied by the local policy utility value corresponding to the environment state feature vector to obtain the contribution gradient under the environment state feature vector. The average of the contribution gradients under all environment state feature vectors is taken, and the resulting average gradient vector is the policy gradient term. This policy gradient term indicates the direction of updating the policy model parameters, thereby increasing the probability of selecting actions that can generate higher local policy utility values. This is the basic application of the policy gradient theorem in reinforcement learning, and its technical effect is to drive the policy model to improve itself in a direction that performs better in historical experience. The specific process of constructing a composite loss function for each node is as follows: The node obtains the neighbor policy identifiers of all neighbor nodes from the neighbor policy information table it maintains; the node restores each neighbor policy identifier to an approximate local policy distribution feature through a preset decoding rule. The decoding rule can be a shared, pre-trained lightweight neural network, or it can query a policy prototype library stored locally on the node and issued by the aggregation node by using the policy identifier as an index. Each node selects a preset state set as a comparison benchmark. The preset state set contains 50 to 200 typical environmental state feature vectors. These vectors are distributed by the sink node during node initialization or collected by the node during long-term operation. Their purpose is to provide a common policy evaluation benchmark that covers typical scenarios. For each environmental state feature vector in the preset state set, the node calculates the local policy distribution output by its own policy model and measures the distribution difference between it and the approximate local policy distribution features of each neighboring node. The calculation process of the distribution difference measure is as follows: For each environmental state feature vector in the preset state set, two first actions are sampled from the local policy distribution, the first kernel function value between the two first actions is calculated and the average value is obtained to obtain the first expected value; two second actions are sampled from the approximate local policy distribution features of the neighboring node, the second kernel function value between the two second actions is calculated and the average value is obtained to obtain the second expected value; a third action is sampled from the local policy distribution, and a fourth action is sampled from the approximate local policy distribution features of the neighboring node, the third kernel function value between the third action and the fourth action is calculated and the average value is obtained to obtain the third expected value; the first expected value and the second expected value are added together, and then twice the third expected value is subtracted. The result is the generalized energy distance between the local policy distribution and the approximate local policy distribution features of the neighboring node under the current environmental state feature vector. Compared with the traditional KL divergence, this measurement method is more robust to the case of non-overlapping support sets of two distributions and can better measure the overall morphological difference of policy behavior. The kernel function value is calculated as follows: Calculate the square of the Euclidean distance between two actions, take a negative value, and divide it by twice the square of the Gaussian kernel bandwidth coefficient (which can be 0.1). Then calculate the exponent with the natural constant as the base to obtain the Gaussian kernel component, which measures the absolute similarity of the action values. Calculate the dot product of the two action vectors, divide it by the product of their respective magnitudes, subtract this quotient from the number one, and multiply it by the cosine weighting coefficient (which can be 0.5) to obtain the cosine similarity component. This component measures the similarity of the action vector directions, i.e., whether the parameter adjustment trends are consistent. Add the Gaussian kernel component and the cosine similarity component to obtain the final kernel function value. Combining the kernel function can comprehensively measure the similarity of actions from both numerical and directional dimensions. The average value of the distribution difference measure corresponding to all neighboring nodes is calculated again on the preset state set and used as the policy consistency regularization term. The smaller the value of this regularization term, the more similar the overall behavior of the node policy and all neighbor policies is in the typical state. Its technical effect is to implicitly coordinate the policies of nodes in the local area and reduce mutual interference. Next, a composite loss function is constructed: First, the expected local policy utility value is calculated, which is the average of the local policy utility values corresponding to all environmental state feature vectors in the batch data. The negative expected local policy utility value is then multiplied by an adjustable regularization strength coefficient and a policy consistency regularization term, and the composite loss function is formed. This composite loss function realizes the joint optimization of two objectives: individual node performance optimization and neighbor group collaboration. The adjustable regularization strength coefficient is an adaptive coefficient, with an initial value of a preset fixed positive number (e.g., 0.1). It can be dynamically adjusted according to the ratio of the average historical local policy utility value of the node itself to the average historical local policy utility value of neighboring nodes. When the node's own average value is consistently lower than the neighbor's average value, the coefficient is increased (to be more inclined to learn collaboration from the neighbors), and vice versa (to focus more on individual performance optimization). (This adaptive mechanism ensures that the weight of collaborative optimization matches environmental adaptability and the node's own competitiveness.) Finally, the policy model parameters are updated by subtracting the product of the learning rate and the gradient of the composite loss function with respect to the policy model parameters from the current values of the policy model parameters. The learning rate is a preset fixed positive number (e.g., 0.001) or a positive number dynamically adjusted according to the optimization algorithm (e.g., the adaptive learning rate when using the Adam optimizer). This parameter update step finally completes one round of iterative optimization of the policy model, enabling it to simultaneously consider higher individual utility and better neighbor cooperation in the next decision.
[0022] In this embodiment, it is specifically necessary to explain that in step S3, the process of periodically comparing the local policy distribution of the current node's policy model on the preset state set with the local policy distribution of the neighbor nodes' policy models on the same preset state set based on the neighbor node's policy identifier, and calculating the policy similarity, is as follows: In each computation cycle, the node first uses its own maintained strategy model to perform forward computation on each environmental state feature vector in the preset state set to obtain the local strategy distribution of the node for all possible RF front-end operating parameter combinations under the current environmental state feature vector, and records it to form the local strategy distribution set of the node. The local strategy distribution set is stored in the node memory in the form of a matrix. The rows of the matrix correspond to different environmental state feature vectors in the preset state set, the columns correspond to different RF front-end operating parameter combinations, and the matrix element values are the probabilities of selecting the corresponding parameter combinations output by the strategy model. Meanwhile, the node sends a query request to each neighbor node recorded in the neighbor policy information table, requesting to obtain the local policy distribution of that neighbor node on the same preset state set; the node receives the response from the neighbor node and obtains its corresponding set of neighbor local policy distributions. This step is completed through point-to-point communication, ensuring that the distribution being compared is the real-time output of the neighbor node's policy model at the current moment, rather than a historical cache or approximation, thereby improving the timeliness and accuracy of similarity assessment. For each neighboring node, the node calculates policy similarity based on its own local policy distribution set and the neighboring node's neighbor's local policy distribution set. The calculation process for policy similarity is as follows: For each environmental state feature vector in the preset state set, the information entropy of the node's local policy distribution under that environmental state feature vector is calculated. The information entropy is the sum of each probability value in the distribution multiplied by the base-2 logarithm of that probability value, and then the summation of all products is taken as the negative value. Then, an exponential function with the base of the natural constant is calculated, and its exponent is the negative value of the information entropy, to obtain the node's decision confidence weight on this environmental state feature vector. This weight reflects the node's confidence in its own decision under this state; the lower the entropy, the higher the confidence. The more explicit the decision, the higher the weight, making the similarity calculation focus more on comparing the strategy behavior of nodes in a "confident" state. Next, the generalized Jensen-Shannon divergence between the local policy distributions of the current node and its neighbors under the environmental state feature vector is calculated. The calculation process is as follows: the decision confidence weight is used as the first coefficient, and the first coefficient is subtracted by the number one to obtain the second coefficient. This makes the divergence calculation biased towards the node's own distribution. For example, the first coefficient can be 0.6 and the second coefficient can be 0.4. The first intermediate distribution is calculated, which is the sum of the first coefficient multiplied by the local policy distribution of the current node and the second coefficient multiplied by the local policy distribution of the neighbors. The K value between the local policy distribution of the current node and the first intermediate distribution is used. The L-divergence is added to the second coefficient multiplied by the KL divergence between the local policy distribution of neighboring nodes and the first intermediate distribution. The KL divergence is calculated by multiplying the value of that element in the first probability distribution for each element in the probability distribution by the base-2 logarithm of that value, then dividing by the value of the corresponding element in the second probability distribution, and summing all the results. Then, the calculated generalized Jensen-Shannon divergence is weighted by the decision confidence weights to obtain the weighted divergence value under the environmental state feature vector. This weighting operation further strengthens the influence of high-confidence states in the overall similarity assessment. The average value of the weighted divergence values under all environmental state feature vectors in the preset state set is calculated, specifically by calculating the sum of all weighted divergence values and calculating all decision confidence weights. The sum of the decision confidence weights is divided by the sum of the weighted divergence values to obtain the second average. The square root of the second average is then taken, and the result is subtracted from the square root by the number one to obtain the policy similarity with the neighboring node. The square root operation makes the final similarity value less sensitive to distribution differences, enhancing the stability of the measurement. This policy similarity is a value between zero and one. The larger the value, the more similar the core decision behaviors of the two node policy models are on the preset state set. This calculation method can effectively capture the similarity of policy models in "decision style" or "behavioral preference", rather than just the simple overlap of output distribution, providing an accurate basis for subsequent selective knowledge transfer. When the policy similarity is higher than a preset first similarity threshold, and the historical utility record corresponding to the policy identifier of the neighboring node is better than the current local policy utility value of this node, the process of obtaining policy soft labels from the corresponding neighboring nodes and performing knowledge distillation training is as follows: A node calculates its policy similarity with a neighboring node and compares it to a preset first similarity threshold, which can be set to 0.75. This threshold defines a "highly similar" policy behavior. Simultaneously, the node retrieves the historical utility record associated with that neighboring node from the neighbor policy information table. This historical utility record is the moving average of the neighboring node's recent local policy utility value; for example, it can be the exponential moving average of the local policy utility value over the past 20 policy update cycles. The node then compares the retrieved historical utility record of the neighboring node with its own recent local policy utility value moving average, based on the following criteria: The calculation determines whether the calculated policy similarity is greater than the first similarity threshold, and whether the moving average of the historical utility records of the neighboring nodes is greater than the moving average of the historical utility records of this node plus a preset performance margin value. The performance margin value can be set to 0.05 to avoid triggering unnecessary distillation when the performance is extremely close, ensuring that the migration has a clear performance improvement expectation. The condition is satisfied if and only if the above two conditions are satisfied at the same time. The node then sends a knowledge distillation request to the neighboring node. After receiving the request, the neighboring node sends its local policy distribution set on the preset state set as a policy soft label back to the current node. After the current node receives the policy soft label, it starts knowledge distillation training: Nodes input a preset state set into their own policy model to obtain the current local policy distribution set, and simultaneously extract the corresponding neighbor local policy distributions from the received policy soft tags. Nodes introduce a temperature parameter greater than one (usually set to 3.0). This temperature parameter softens the probability distribution; increasing the temperature reduces the probability differences between different actions, thus highlighting the relative merits of "suboptimal actions," i.e., "hidden knowledge." For each environmental state feature vector in the preset state set, the node's current local policy distribution and the corresponding neighbor local policy distribution extracted from the policy soft tags are processed as follows: each probability value in the probability distribution is divided by the temperature parameter, and then an exponential function with the natural constant as its base is calculated, where the exponent is the probability value after dividing by the temperature parameter, yielding an intermediate exponent value. For all... The intermediate exponent values are summed to obtain the normalized denominator. Each intermediate exponent value is divided by the normalized denominator to obtain a new probability distribution after temperature scaling and normalization. This process is the core step of knowledge distillation, which generates a set of smoother and more information-rich "softened" target distributions. Then, for each environmental state feature vector in the preset state set, the KL divergence between the new probability distributions of neighboring nodes after temperature scaling and normalization and the new probability distribution of the current node after temperature scaling and normalization is calculated. This divergence is used as the knowledge distillation loss component under that environmental state feature vector. The average value of the knowledge distillation loss components under all environmental state feature vectors is taken to obtain the final knowledge distillation loss. This loss function drives the output probability distribution of the current node's policy model to approximate the excellent, softened distribution of neighboring nodes in terms of morphology, thereby learning their decision preferences. Nodes minimize knowledge distillation loss by executing one or more gradient descent steps (usually 1 to 3 steps), thereby updating the parameters of their own policy model. This makes the output distribution of the node's policy model approximate the decision preferences reflected by the policy soft labels of neighboring nodes in terms of morphology. This update is local and small-scale, aiming to absorb the "essence" of neighboring policies rather than copying them entirely. This enables the effective and low-risk migration of high-performance policies between nodes with similar behaviors, accelerating the overall performance improvement and equalization of the network.
[0023] In this embodiment, it is specifically necessary to explain that in step S4, the process of comparing the actually measured communication performance indicators with the local policy utility value predicted by the utility evaluation model to generate a policy confidence score is as follows: Within a preset sliding time window, the node calculates a sequence of communication performance indicators for each time period. The sliding time window length can be set to 10 to 20 time periods to balance the immediacy and stability of the evaluation. Based on the environmental observation data continuously collected in step S1, the node calculates a sequence of communication performance indicators for each time period. The actual measured communication performance indicators are scalar values obtained by weighted summation of average throughput, average packet error rate and average energy consumption within the period. For example, the weights can be set to 0.5, 0.3 and 0.2 respectively, so that the evaluation focuses on throughput while taking into account reliability and energy efficiency. At the same time, the local policy utility value predicted by the utility evaluation model based on the current environmental state is collected at the beginning of each time period. For each time period, the difference between the actual measured value and the predicted local policy utility value is calculated to obtain the predicted residual sequence. The predicted residual sequence is then weighted and averaged over a sliding time window to obtain the weighted average residual. The weight of each residual is determined by an exponential function with a base of the natural constant. The exponent of this exponential function is a decay coefficient with a negative exponent multiplied by the difference between the current time and the time corresponding to the residual, so that recent residuals receive higher weights. The decay coefficient can be set to 0.2 to ensure that residuals in recent periods occupy the main weights, thereby quickly reflecting the latest performance trends. Simultaneously, the normalized median absolute deviation of the predicted residual sequence is calculated. The calculation process is as follows: first, the median of the entire residual sequence is calculated; then, the absolute value of the difference between each residual and the median is calculated to obtain the absolute deviation sequence; next, the median of this absolute deviation sequence is calculated; finally, the median is divided by a constant 0.6745. Dividing by 0.6745 is to ensure that the estimator is approximately consistent with the standard deviation when the residuals follow a normal distribution. This constant is derived from the interquartile range of the standard normal distribution. The absolute value of the weighted average residual is subtracted from the product of a significance level coefficient and the normalized median absolute deviation to obtain the deviation detection value. The significance level coefficient is usually taken as 2, which is equivalent to examining whether the residuals exceed two quartiles under the assumption of normal distribution. The standard deviation range is calculated. The deviation detection value is input into a sigmoid growth function for mapping. The sigmoid growth function is a logistic function whose output value is between zero and one. It has a scaling factor that controls the steepness of the function curve. The scaling factor can be set to 0.5 to adjust the sensitivity of the score to the deviation detection value. The smaller the factor, the faster the score decreases. The mapping result is then subtracted from the number one to obtain the strategy confidence score. The strategy confidence score is a value between zero and one. The higher the score, the higher the consistency between the model prediction and the actual performance. The confidence threshold for judging whether the score is too low is set to 0.7. This threshold provides a clear performance degradation judgment line. A score that is consistently below 0.7 indicates a significant and systematic deviation between the prediction and the actual measurement. The process of triggering the evolution and update of the policy model of this node, and sending a model evolution notification signal containing a summary of environmental features to neighboring nodes whose policy similarity exceeds a preset second similarity threshold, is as follows: When the policy confidence score falls below the confidence threshold for three consecutive time periods, the node starts with the parameters of the current policy model and performs model evolution based on adaptive Cauchy mutation. Specifically, the node generates multiple candidate model parameter variants in parallel (e.g., 8 or 16 candidate variants). Each candidate variant is obtained by applying a Cauchy distribution perturbation to the current policy model parameters. The Cauchy distribution perturbation is achieved by adding a perturbation step size coefficient to the current parameters and multiplying it by a random vector following a Cauchy distribution. The Cauchy distribution has a heavy-tailed property, and the resulting perturbation is highly likely to be... Small fluctuations, but with a certain probability of producing large jumps, help to retain the potential to escape local optima while conducting local searches. The Cauchy distribution has a location parameter of zero, and the scale parameter is an adaptive parameter negatively correlated with the current policy confidence score. Specifically, the scale parameter is equal to a preset base scale value multiplied by a number one minus the difference between the current policy confidence score and the base scale value. For example, if the base scale value is set to 0.01, and the current score is 0.6, then the scale parameter is 0.01(1-0.6)=0.004. The lower the score, the larger the perturbation range and the stronger the exploratory nature. The node uses recently collected, unused mini-batch environment state feature vectors as a validation set. The validation set size is typically half the batch data size, such as 16 or 32 samples. Each candidate variant is quickly evaluated by inputting the environment state feature vectors from the validation set into the candidate variant to obtain the recommended action. The predicted utility of this action is then calculated using a utility evaluation model, and the average predicted utility of all validation samples is taken. The candidate variant with the highest average predicted utility is selected as the parameter for the next-generation policy model. This selection process achieves "survival of the fittest," directing the search towards the parameter region that performs better in the current validation environment. After its own model evolves, the node selects all neighboring nodes whose policy similarity exceeds a preset second similarity threshold from the stored policy similarity records, and sends a model evolution notification signal to these neighboring nodes via unicast. The model evolution notification signal contains an environmental feature summary, which is a statistical summary of the environmental state feature vector within the most recent window before the evolution is triggered, including the mean of each dimension and the principal components. For example, the arithmetic mean of each dimension of the environmental state feature vector is calculated, and the eigenvectors corresponding to the first 1-2 largest eigenvalues of its covariance matrix are calculated as principal components, thereby compressing the main patterns that characterize environmental changes. After receiving the model evolution notification signal, neighboring nodes take it as an early warning that their environment may undergo similar changes. Based on this, they selectively increase the verification frequency of their own policy confidence scores. For example, they may shorten the verification cycle to half of the original time or slightly increase the scale parameter of the Cauchy distribution perturbation when they execute model evolution in the future. For example, they may temporarily increase the base scale value by 20%. This enables them to passively and collaboratively adapt to potential changes in the environment. The second similarity threshold is set to 0.6, which is lower than the first similarity threshold used in knowledge distillation. This expands the coverage of the early warning network, allowing nodes with certain similar behaviors to be aware of signals of possible environmental changes in advance.
[0024] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0025] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0026] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0027] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0028] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0029] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0030] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication, characterized in that, Specifically, the steps include the following: Step S1: Each node continuously collects local environmental observation data, which includes at least the instantaneous signal-to-interference-plus-noise ratio (SIR) and the operating parameters of the local RF front-end. Simultaneously, it receives policy identifiers from neighboring nodes with which it has a direct communication connection. Each node maintains a policy model for deciding on the RF front-end operating parameters and, based on a preset utility evaluation model, quantitatively evaluates the RF front-end anti-interference strategy currently adopted by the policy model, generating a local policy utility value, specifically: First, calculate the communication quality components: subtract a preset demodulation threshold from the signal-to-interference-plus-noise ratio characteristic value, multiply the difference by a signal quality adjustment coefficient, calculate its arctangent value, and then multiply it by a first constant for scaling to obtain the communication quality components. Secondly, calculate the link instability penalty component: obtain the rate of change of the recent packet reception rate, calculate the square of the rate of change and take the negative, multiply the obtained negative square value by a fluctuation sensitivity coefficient, calculate the exponential function value with the natural constant as the base, and subtract this exponential function value from one to obtain the link instability penalty component. Next, calculate the energy cost component: calculate the exponential function value after multiplying the node's remaining energy percentage by a negative energy influence factor, and multiply the normalized transmit power by this exponential function value to obtain the energy cost component; Next, a weighted merging process is performed: the communication quality component is multiplied by the communication quality weight, the difference after subtracting the link instability penalty component is multiplied by the link stability weight, the two products are added together, and then multiplied by the difference after subtracting a global performance trade-off factor to obtain the weighted performance result. Multiply the energy cost component by the global efficiency trade-off factor to obtain the weighted energy cost result; Finally, calculate the local policy utility value: subtract the weighted energy cost value from the weighted performance value to obtain the local policy utility value; Step S2: Each node calculates the policy gradient term based on its local policy utility value, and infers the local policy distribution characteristics of its neighboring nodes' policy models based on the neighboring nodes' policy identifiers. Each node constructs a composite loss function, which includes a policy gradient term and a policy consistency regularization term. The policy consistency regularization term is used to constrain the difference between the local policy distribution output by the node's policy model and the local policy distribution characteristics of its neighboring nodes' policy models. Each node updates the parameters of its policy model by minimizing the composite loss function. Step S3: Each node periodically compares the local policy distribution of its policy model on a preset state set with the local policy distribution of the policy models of neighboring nodes corresponding to the policy identifiers of neighboring nodes on the same preset state set, and calculates the policy similarity. When the policy similarity is higher than the preset first similarity threshold, and the historical utility record corresponding to the policy identifier of the neighboring node, recorded locally or obtained from the neighboring node, is better than the current local policy utility value of the current node, the current node obtains the policy soft label output by the policy model of the corresponding neighboring node, and uses the policy soft label to perform knowledge distillation training on the policy model of the current node. Step S4: After applying the updated current policy model in step S2 for one time period, each node compares the actual measured communication performance indicators with the local policy utility value predicted by the utility evaluation model to generate a policy confidence score. If the policy confidence score is lower than the preset confidence threshold for multiple consecutive time periods, the current node triggers the evolution update of its policy model and sends a model evolution notification signal containing an environmental feature summary to neighboring nodes whose policy similarity exceeds the preset second similarity threshold.
2. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 1, characterized in that: In step S1, local environmental observation data is continuously collected, including the instantaneous signal-to-interference-plus-noise ratio and the operating parameters of the local radio frequency front end. The operating parameters include at least the transmit power, filter bandwidth and filter center frequency. At the same time, the percentage of remaining energy of the node and the rate of change of recent packet reception rate are collected. The instantaneous signal-to-interference-plus-noise ratio (SINR) is subjected to logarithmic compression with the natural logarithm as the base to obtain the SINR characteristic value; The values of transmit power, filter bandwidth, and filter center frequency are mapped to a closed interval between zero and one to obtain normalized transmit power, normalized filter bandwidth, and normalized center frequency. The signal-to-interference-plus-noise ratio (SIR) eigenvalues, normalized transmit power, normalized filter bandwidth, normalized center frequency, and percentage of remaining node energy are combined in sequence to form an environmental state feature vector.
3. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 2, characterized in that: The process of receiving policy identifiers from neighboring nodes with which it has a direct communication connection specifically involves: A node synchronously receives neighbor node policy identifiers periodically broadcast by neighbor nodes within its one-hop communication range. The neighbor node policy identifier is a hash digest value that represents key information of the policy model maintained by the neighbor node. The node associates the received neighbor node policy identifier with the identity information of the neighbor node that sent the neighbor node policy identifier, and stores it in the local neighbor policy information table.
4. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 3, characterized in that: In step S2, the process of calculating the policy gradient term based on the local policy utility value and inferring the local policy distribution characteristics of the neighboring nodes' policy models based on the neighboring node policy identifiers is as follows: The node randomly selects a batch of data consisting of multiple environmental state feature vectors from the local stored history. For each environmental state feature vector in the batch of data, the node inputs it into its own maintained policy model to obtain the local policy distribution output by the policy model for different combinations of RF front-end operating parameters. At the same time, the node retrieves the local policy utility value recorded when the policy was previously executed from the historical records, corresponding to each environmental state feature vector; Based on the environmental state feature vector, the corresponding local policy distribution, and the local policy utility value in the batch data, the node calculates the policy gradient term. The calculation process for this policy gradient term is as follows: For each environmental state feature vector in the batch data, an action is sampled according to the current local policy distribution. The gradient of the log probability of the action with respect to the policy model parameters is calculated. This gradient is multiplied by the local policy utility value corresponding to the environmental state feature vector to obtain the contribution gradient under the environmental state feature vector. The average of the contribution gradients under all environmental state feature vectors is taken, and the resulting average gradient vector is the policy gradient term.
5. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 4, characterized in that: The specific process of constructing a composite loss function for each node is as follows: The node retrieves the neighbor policy identifiers of all neighboring nodes from its maintained neighbor policy information table; the node then uses preset decoding rules to restore each neighbor policy identifier to an approximate local policy distribution feature. Each node selects a preset state set as a comparison benchmark. For each environmental state feature vector in the preset state set, the node calculates the distribution difference between the local policy distribution output by its own policy model and the approximate local policy distribution features of each neighboring node. The calculation process of the distribution difference metric is as follows: For each environmental state feature vector in the preset state set, two first actions are sampled from the local policy distribution, the first kernel function value between the two first actions is calculated and the average value is obtained to obtain the first expected value; two second actions are sampled from the approximate local policy distribution features of the neighboring nodes, the second kernel function value between the two second actions is calculated and the average value is obtained to obtain the second expected value; a third action is sampled from the local policy distribution, and a fourth action is sampled from the approximate local policy distribution features of the neighboring nodes, the third kernel function value between the third action and the fourth action is calculated and the average value is obtained to obtain the third expected value; the first expected value and the second expected value are added together, and then twice the third expected value is subtracted. The result is the generalized energy distance between the local policy distribution and the approximate local policy distribution features of the neighboring node under the current environmental state feature vector. The average of the distribution difference measures corresponding to all neighboring nodes calculated on the preset state set is averaged again and used as the policy consistency regularization term. Next, a composite loss function is constructed: First, the expected local policy utility value is calculated, which is the average of the local policy utility values corresponding to all environmental state feature vectors in the batch data; the negative expected local policy utility value is added to the product of an adjustable regularization strength coefficient and a policy consistency regularization term to form the composite loss function. Finally, the parameters of the policy model are updated by subtracting the product of the learning rate and the gradient of the composite loss function with respect to the policy model parameters from the current values of the policy model parameters.
6. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 5, characterized in that: In step S3, the process of periodically comparing the local policy distribution of the current node's policy model on a preset state set with the local policy distribution of the neighbor nodes' policy models on the same preset state set based on the neighbor node's policy identifier, and calculating the policy similarity, is as follows: In each calculation cycle, the node first uses its own maintained strategy model to perform forward calculation on each environmental state feature vector in the preset state set to obtain the local strategy distribution of the node for all optional RF front-end working parameter combinations under the current environmental state feature vector, and records it to form the local strategy distribution set of the node. At the same time, the node sends a query request to each neighbor node recorded in the neighbor policy information table, requesting to obtain the local policy distribution of that neighbor node on the same preset state set; A node receives responses from neighboring nodes and obtains the corresponding set of local policy distributions for those neighbors. For each neighboring node, the node calculates policy similarity based on its own local policy distribution set and the neighboring node's neighboring local policy distribution set. The calculation process for policy similarity is as follows: For each environmental state feature vector in the preset state set, the information entropy of the node's local policy distribution under that environmental state feature vector is calculated. The information entropy is the sum of each probability value in the distribution multiplied by the base-2 logarithm of that probability value, and the negative value is taken after summing all the products. Then, an exponential function with the natural constant as the base is calculated, and its exponent is the negative value of the information entropy, to obtain the decision confidence weight of the node on this environmental state feature vector. Next, the generalized Jensen-Shannon divergence between the two local policy distributions of the node and its neighbors under this environmental state feature vector is calculated. The calculation process is as follows: the decision confidence weight is used as the first coefficient, and the first coefficient is subtracted from the number one to obtain the second coefficient. The first intermediate distribution is calculated, which is the sum of the first coefficient multiplied by the local policy distribution of the node and the second coefficient multiplied by the logarithm of the logarithm. The sum of the local policy distributions of neighboring nodes; multiplying the local policy distribution of this node by the KL divergence between the local policy distribution and the first intermediate distribution using a first coefficient, and adding the second coefficient multiplied by the KL divergence between the local policy distribution of neighboring nodes and the first intermediate distribution. The KL divergence is calculated by multiplying the value of the element in the first probability distribution for each element in the probability distribution by the base-2 logarithm of that value, then dividing by the value of the corresponding element in the second probability distribution, and summing all the results; then, the calculated generalized Jensen-Shannon divergence is weighted by the decision confidence weights to obtain the weighted divergence value under the environmental state feature vector; the average of the weighted divergence values under all environmental state feature vectors in the preset state set is calculated, specifically by summing all weighted divergence values and summing all decision confidence weights, and dividing the sum of weighted divergence values by the sum of decision confidence weights to obtain the second average; the square root of the second average is taken, and finally, the square root result is subtracted from the number one to obtain the policy similarity with the neighboring node.
7. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 6, characterized in that: The process of obtaining policy soft labels from the corresponding neighboring nodes and performing knowledge distillation training when the policy similarity is higher than a preset first similarity threshold, and the historical utility record corresponding to the policy identifier of the neighboring node is better than the current local policy utility value of the current node, specifically involves: A node calculates its policy similarity with a neighboring node and compares it with a preset first similarity threshold. Simultaneously, it retrieves the historical utility record associated with that neighboring node from the neighbor policy information table. This historical utility record is the moving average of the neighboring node's recent local policy utility value. The node then compares the retrieved historical utility record of the neighboring node with its own recent local policy utility value moving average, based on the following criteria: Whether the calculated strategy similarity is greater than the first similarity threshold, and whether the moving average of the historical utility records of the neighboring nodes is greater than the moving average of the historical utility records of this node plus a preset performance margin value; The condition is satisfied if and only if both of the above conditions are met simultaneously; the node then sends a knowledge distillation request to the neighboring node. After receiving the request, the neighboring node sends its local policy distribution set on the preset state set as a policy soft label back to the current node. After the current node receives the policy soft label, it starts knowledge distillation training: The node inputs the preset state set into its own policy model to obtain the current local policy distribution set, and at the same time extracts the corresponding neighbor local policy distribution from the received policy soft tags. A temperature parameter is introduced for each node. For each environmental state feature vector in the preset state set, the current local policy distribution of this node and the corresponding neighbor local policy distribution extracted from the policy soft tag are processed as follows: each probability value in the probability distribution is divided by the temperature parameter, and then an exponential function with the natural constant as the base is calculated, where the exponent is each probability value divided by the temperature parameter, to obtain intermediate exponent values; all intermediate exponent values are summed to obtain a normalized denominator; each intermediate exponent value is divided by the normalized denominator to obtain a new probability distribution after temperature scaling and normalization; then, for each environmental state feature vector in the preset state set, the KL divergence between the new probability distribution of neighbor nodes after temperature scaling and normalization and the new probability distribution of this node after temperature scaling and normalization is calculated, and this is used as the knowledge distillation loss component under this environmental state feature vector; the average value of the knowledge distillation loss components under all environmental state feature vectors is calculated to obtain the final knowledge distillation loss. Nodes update the parameters of their own policy models by performing one or more steps of gradient descent algorithm to minimize knowledge distillation loss.
8. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 7, characterized in that: In step S4, the process of comparing the actually measured communication performance indicators with the local policy utility value predicted by the utility evaluation model to generate a policy confidence score is as follows: Within a preset sliding time window, based on the environmental observation data continuously collected in step S1, the node calculates the actual measured communication performance index sequence for each time period. The actual measured communication performance index is a scalar value obtained by weighted summation of the average throughput, average packet error rate and average energy consumption within the period. At the same time, the local policy utility value predicted by the utility evaluation model based on the current environmental state is collected at the beginning of each time period. For each time period, the difference between the actual measured value and the predicted local policy utility value is calculated to obtain the predicted residual sequence; the predicted residual sequence is then weighted and averaged over a sliding time window to obtain the weighted average residual. Simultaneously, the normalized median absolute deviation of the predicted residual sequence is calculated. The calculation process of the normalized median absolute deviation is as follows: first, the median of the entire residual sequence is calculated; then, the absolute value of the difference between each residual and the median is calculated to obtain the absolute deviation sequence; then, the median of the absolute deviation sequence is calculated; finally, the median is divided by a constant 0.6745. The absolute value of the weighted average residual is subtracted from the product of a significance level coefficient and the normalized median absolute deviation to obtain the deviation detection value. The deviation detection value is input into the S-shaped growth function for mapping, and its output value is between zero and one, and has a scaling factor that controls the steepness of the function curve. The mapping result is then subtracted from the number one to obtain the policy confidence score.
9. The method for optimizing a radio frequency front-end for strong interference-resistant, through-ground communication according to claim 8, characterized in that: The process of triggering the evolution and update of the policy model of this node, and sending a model evolution notification signal containing an environmental feature summary to neighboring nodes whose policy similarity exceeds a preset second similarity threshold, is specifically as follows: When the policy confidence score is lower than the confidence threshold for three consecutive time periods, the node starts with the parameters of the current policy model and performs model evolution based on adaptive Cauchy mutation. Nodes use recently collected, unused mini-batch environment state feature vectors as a validation set to quickly evaluate each candidate variant. The evaluation method involves inputting the environment state feature vectors from the validation set into the candidate variant to obtain the recommended action, then calculating the predicted utility of the action through a utility evaluation model, and taking the average predicted utility of all validation samples. The candidate variant with the highest average predicted utility is selected as the parameter for the next-generation policy model. After completing its own model evolution, the node selects all neighboring nodes whose policy similarity exceeds a preset second similarity threshold from the stored policy similarity records, and unicasts a model evolution announcement signal to these neighboring nodes. The model evolution announcement signal contains an environment feature summary, which is a statistical summary of the environment state feature vectors within the most recent time window before triggering the evolution, including the mean and principal components of each dimension.
Citation Information
Patent Citations
Interference confrontation method, system and device for wireless communication network, and medium
CN121194223A
Wireless communication anti-interference strategy optimization method and system based on reinforcement learning
CN121841544A