Network configuration methods, systems, devices, and media based on radio frequency environment fingerprinting
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]相关技术CN113302980B(一种信道分配方法、系统、电子设备及存储介质,发明人雷永成、吴方)中,公开了一种利用强化学习模型根据WLAN网络状态信息为接入点推荐最优信道的方案,其方案定义了基于信道切换QoE的奖励函数,通过捕获信道切换前后邻居RSSI变化、流量拥塞变化及最大发射功率变化来动态评估信道质量,以解决传统自动信道选择(ACS)无法适应变化网络环境的问题
[0016]本申请公开了一种基于射频环境指纹的网络配置方法,本方法根据当前无线网络接入点的射频遥测数据生成当前射频环境指纹向量,结合上述当前射频环境指纹向量和强化学习核心实体的当前策略生成决策信息。上述决策信息中包含目标配置动作、置信度评分和候选作用域,置信度评分为所述目标配置动作的置信度,候选作用域为已应用过所述目标配置动作的无线网络接入点的集合;在置信度评分大于第一置信度阈值时,本申请根据射频环境指纹向量的相似度从候选作用域中选取无线网络接入点作为有效作用域,进而在上述有效作用域内应用所述目标配置动作。上述过程以射频环境指纹作为配置决策的依据,通过置信度和候选作用域的限制,实现了配置动作与实际射频环境的精准匹配,并将配置变更范围严格限定在环境一致且经验证的范围内。因此,本申请能够基于射频环境对无线网络接入点进行准确配置,提高无线网络配置的安全性和可靠性。本申请同时还提供了一种基于射频环境指纹的网络配置系统、一种存储介质和一种电子设备,具有上述有益效果,在此不再赘述。
Smart Images

Figure CN122579178A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless local area network technology, and in particular to a network configuration method, system, device and medium based on radio frequency environment fingerprinting. Background Technology
[0002] In wireless local area networks (WLANs), radio spectrum is an important shared network resource that needs to be carefully managed to optimize the end-user experience.
[0003] The related technology CN113302980B (A channel allocation method, system, electronic device and storage medium, inventors Lei Yongcheng and Wu Fang) discloses a scheme that uses a reinforcement learning model to recommend the optimal channel for access points based on WLAN network state information. This scheme defines a reward function based on channel handover QoE and dynamically evaluates channel quality by capturing changes in neighbor RSSI, traffic congestion, and maximum transmit power before and after channel handover, thus addressing the problem that traditional Automatic Channel Selection (ACS) cannot adapt to changing network environments. However, the aforementioned related technologies only address the single configuration dimension of channel selection and lack a multi-dimensional perception and confidence quantification constraint mechanism for the overall RF (Radio Frequency) physical environment characteristics of the site. This results in association failures and coverage gaps at some sites due to differences in local RF environments.
[0004] Therefore, how to accurately configure wireless network access points based on the radio frequency environment and improve the security and reliability of wireless network configuration is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this application is to provide a network configuration method, system, device, and medium based on radio frequency environment fingerprinting, which can accurately configure wireless network access points based on the radio frequency environment, thereby improving the security and reliability of wireless network configuration.
[0006] To address the aforementioned technical problems, this application provides a network configuration method based on radio frequency environment fingerprinting, the network configuration method comprising: Acquire the radio frequency telemetry data of the current wireless network access point, and generate the current radio frequency environment fingerprint vector based on the radio frequency telemetry data; Decision information is generated based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity; wherein, the decision information includes multiple configuration description fields, and the configuration description fields include at least the target configuration action, confidence score and candidate scope; the confidence score is the confidence of the target configuration action, and the candidate scope is the set of wireless network access points that have applied the target configuration action; If the confidence score is greater than the first confidence threshold, then the reference radio frequency environment fingerprint vector of each wireless network access point in the candidate scope when performing the target configuration action is determined, and wireless network access points are selected from the candidate scope as effective scopes based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector. Apply the target configuration action within the effective scope.
[0007] Optionally, selecting a wireless network access point as an effective scope from the candidate scopes based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector includes: Determine the fingerprint similarity between each of the reference radio frequency environment fingerprint vectors and the current radio frequency environment fingerprint vector; Based on the fingerprint similarity, all wireless network access points in the candidate scope are divided into a first type of wireless network access point, a second type of wireless network access point, and a third type of wireless network access point; wherein, the fingerprint similarity corresponding to the first type of wireless network access point is greater than a first fingerprint similarity threshold, the fingerprint similarity corresponding to the second type of wireless network access point is less than or equal to the first fingerprint similarity threshold and greater than a second fingerprint similarity threshold, and the fingerprint similarity corresponding to the third type of wireless network access point is less than or equal to the second fingerprint similarity threshold; Add all first-type wireless network access points to the candidate set; If a manual confirmation message is received, the second type of wireless network access point corresponding to the manual confirmation message is added to the candidate set; Determine whether the number of wireless network access points in the candidate set is greater than the number threshold Y; If so, then the set of the Y wireless network access points with the highest fingerprint similarity in the candidate set is set as the effective scope; If not, then the candidate set is set as the effective scope.
[0008] Optionally, decision information is generated based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity, including: The current radio frequency environment fingerprint vector is input into the reinforcement learning core entity so that the reinforcement learning core entity determines the target configuration action based on the current policy. Obtain multiple historical status information entries; wherein each historical status information entry includes the configuration action a executed at time step t. t Application configuration action a t The obtained performance index values and the RF environment fingerprint vector at time step t; Calculate the similarity between the reference radio frequency environment fingerprint vector in the historical state information and the current radio frequency environment fingerprint vector, and add the K historical state information entries with the highest similarity to the historical matching set; Configure action a in the historical matching set t The historical state information of the actions configured for the target is added to the reference set; Using the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector as weights, the confidence score is obtained by weighted statistical analysis of the cases in the reference set where the performance index value meets the preset performance conditions after the target configuration action is performed. Set the set of wireless network access points corresponding to each historical state information in the reference set as the candidate scope; Construct the decision information including the target configuration action, the confidence score, and the candidate scope.
[0009] Optional, also includes: Generate summary information for the historical state information in the historical matching set, and add the summary information to the decision information; If the confidence score is less than or equal to the first confidence threshold and greater than the second confidence threshold, then the decision information is sent to the management device. If an authorization instruction for the decision information is received from the management device, the effective scope is determined, and the target configuration action is applied within the effective scope.
[0010] Optional, also includes: If the confidence score is less than the second confidence threshold, the current radio frequency environment fingerprint vector is added to the sample pool, and radio frequency telemetry data and configuration execution results of other wireless network access points with a similarity higher than the preset collection threshold to the current radio frequency environment fingerprint vector are obtained in order to increase historical status information; When the newly added historical state information exceeds the preset accumulated amount, the confidence score of the target configuration action is recalculated.
[0011] Optionally, after applying the target configuration action within the effective scope, the method further includes: Obtain performance metrics after applying the target configuration action; wherein, the performance metrics include signal-to-noise ratio change, roaming duration change, channel utilization offset, number of coverage holes, association failure rate, and number of abnormal devices; the channel utilization offset is used to describe the degree of deviation between the actual channel utilization and the target channel utilization, and the number of abnormal devices is used to describe the number of wireless network access points that are abnormal after applying the target configuration action; The reward function value is calculated based on the performance index, and the reward function value is normalized using the L2 norm of the current radio frequency environment fingerprint vector; The current policy of the core entity in the reinforcement learning is updated based on the normalized reward function value.
[0012] Optionally, after applying the target configuration action within the effective scope, the method further includes: Within a preset time window, collect performance metrics before and after applying the target configuration action; Determine whether the performance indicators before and after applying the target configuration action meet the anomaly judgment conditions; If so, a rollback operation is performed to restore the configuration to its state before the target configuration action was applied, and the reward function value corresponding to the target configuration action is modified to a preset penalty value; The current radio frequency environment fingerprint vector, the target configuration action, and the modified reward function value are set as penalty samples so that the current policy of the reinforcement learning core entity is updated based on the penalty samples.
[0013] This application also provides a network configuration system based on radio frequency environmental fingerprinting, the network configuration system comprising: The fingerprint generation module is used to acquire the radio frequency telemetry data of the current wireless network access point and generate the current radio frequency environment fingerprint vector based on the radio frequency telemetry data. A decision module is used to generate decision information based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity; wherein, the decision information includes multiple configuration description fields, and the configuration description fields include at least a target configuration action, a confidence score, and a candidate scope; the confidence score is the confidence of the target configuration action, and the candidate scope is a set of wireless network access points that have applied the target configuration action; The scope determination module is used to determine the reference radio frequency environment fingerprint vector when each wireless network access point in the candidate scope performs the target configuration action if the confidence score is greater than the first confidence threshold, and select the wireless network access point as the effective scope from the candidate scope based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector. The configuration application module is used to apply the target configuration action within the effective scope.
[0014] This application also provides a storage medium storing a computer program thereon, which, when executed, implements the steps of the network configuration method based on radio frequency environment fingerprinting described above.
[0015] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor, when calling the computer program in the memory, implements the steps of the network configuration method based on radio frequency environment fingerprinting described above.
[0016] This application discloses a network configuration method based on radio frequency (RF) environment fingerprints. This method generates a current RF environment fingerprint vector based on the RF telemetry data of the current wireless network access point, and combines this current RF environment fingerprint vector with the current policy of the reinforcement learning core entity to generate decision information. The decision information includes a target configuration action, a confidence score, and a candidate scope. The confidence score represents the confidence level of the target configuration action, and the candidate scope is the set of wireless network access points that have already applied the target configuration action. When the confidence score is greater than a first confidence threshold, this application selects wireless network access points from the candidate scope as effective scopes based on the similarity of the RF environment fingerprint vectors, and then applies the target configuration action within the effective scope. This process uses the RF environment fingerprint as the basis for configuration decisions. Through the constraints of confidence and candidate scopes, it achieves accurate matching between the configuration action and the actual RF environment, and strictly limits the scope of configuration changes to a consistent and verified range. Therefore, this application can accurately configure wireless network access points based on the RF environment, improving the security and reliability of wireless network configuration. This application also provides a network configuration system based on radio frequency environment fingerprinting, a storage medium, and an electronic device, which have the above-mentioned beneficial effects, and will not be elaborated here. Attached Figure Description
[0017] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a network configuration method based on radio frequency environmental fingerprinting provided in this application embodiment; Figure 2 An adaptive configuration system architecture diagram based on wireless radio frequency environment fingerprinting is provided for an embodiment of this application; Figure 3 This is a diagram illustrating another adaptive configuration system architecture based on wireless radio frequency environment fingerprinting, provided in an embodiment of this application. Figure 4 This is a schematic diagram of the architecture of a radio frequency environment fingerprint generation engine provided in an embodiment of this application; Figure 5 This is a schematic diagram of a reinforcement learning principle provided in an embodiment of this application; Figure 6 This is a schematic diagram of a two-dimensional decision-making framework for bounded confidence and explosion radius control provided in an embodiment of this application; Figure 7 This is a comparison chart of the measured effects of this application and related technologies; Figure 8 This is a schematic diagram illustrating the design principle of a confidence-bounded configuration reward function provided in an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] Please see below. Figure 1 , Figure 1 This is a flowchart illustrating a network configuration method based on radio frequency environment fingerprinting, provided in an embodiment of this application.
[0021] Specific steps may include: S101: Obtain the radio frequency telemetry data of the current wireless network access point, and generate the current radio frequency environment fingerprint vector based on the radio frequency telemetry data.
[0022] This embodiment can be applied to cloud servers or cloud AI (Artificial Intelligence) platforms connected to wireless network access points (APs).
[0023] This step acquires the radio frequency telemetry (RFT) data uploaded by the current wireless network access point. This RFT data may include channel interference rate (CNR), noise floor, channel utilization, client signal-to-noise ratio (SNR), and time-series heatmaps. As a feasible implementation, this step acquires the raw data uploaded by the current wireless network access point and extracts the five quantitative indicators—channel interference rate, noise floor, channel utilization, client SNR, and time-series heatmaps—from the raw data. The raw data includes driver logs, event data, and network performance statistics.
[0024] After obtaining the aforementioned radio frequency telemetry data, the multi-dimensional radio frequency telemetry data can be input into the corresponding fingerprint extraction model to obtain the current radio frequency environment fingerprint vector. Specifically, the fingerprint extraction model may include a convolutional neural network (CNN), a fully connected layer, and a long short-term memory (LSTM) network. The process of generating the current radio frequency environment fingerprint vector using the fingerprint extraction model may include: inputting the current radio frequency environment fingerprint vector into the convolutional neural network to extract spatial features; after dimensionality reduction and fusion of the spatial features through a fully connected layer, inputting them into the LSTM network for temporal modeling and outputting temporal-aware features; and concatenating the spatial features with the temporal-aware features to obtain the current radio frequency environment fingerprint vector.
[0025] S102: Generate decision information based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity.
[0026] In this step, the current radio frequency environment fingerprint vector is input into the core entity of the reinforcement learning, which then generates decision information based on the current policy reasoning.
[0027] The aforementioned decision information includes multiple configuration description fields, which at least include a target configuration action, a confidence score, and a candidate scope. The confidence score is the confidence level of the target configuration action, and the candidate scope is a set of wireless network access points that have applied the target configuration action.
[0028] The aforementioned target configuration action can be directly output by the core entity of reinforcement learning. This step can be based on historical similarity retrieval, and the confidence level can be obtained by weighted statistical action success rate. The set of sites that have successfully applied the action can be taken as the candidate scope.
[0029] Specifically, this step inputs the current radio frequency environment fingerprint vector into the reinforcement learning core entity, enabling the core entity to output the target configuration action based on the current policy inference. This step retrieves historical records similar to the current radio frequency environment fingerprint vector, using similarity as weight, and performs a weighted statistical analysis on cases where the target configuration action achieves the expected improvement in experience metrics (i.e., performance metrics such as channel utilization, roaming latency, association failure rate, etc.) to obtain a confidence score. The confidence score describes the probability that the target configuration action will achieve the expected experience improvement after execution in a radio frequency environment similar to the current radio frequency environment fingerprint vector; that is, the degree of confidence that the configuration action works in the current site's radio frequency physical environment. This step can also use the set of wireless network access points that have applied the target configuration action as a candidate scope. As a feasible implementation, the similarity between the sites in the candidate scope and the current radio frequency environment fingerprint vector can be further limited to a value greater than a set value; that is, the candidate scope is the set of wireless network access points that have applied the target configuration action and whose similarity to the current radio frequency environment fingerprint vector is greater than a set value.
[0030] The aforementioned decision information, also known as a decision contract or configuration recommendation descriptor, refers to a structured output object containing four fields generated by the system for each configuration recommendation. This object serves as a complete set of decision-making criteria to support human-machine collaborative review and configuration execution. The fields included in the decision information can be the target configuration action (i.e., the proposed configuration action), the inference path (containing historical matching records and effect data of similar wireless network access points), the confidence score (ranging from 0 to 1), and the candidate scope (i.e., the set of wireless network access points).
[0031] S103: If the confidence score is greater than the first confidence threshold, then determine the reference radio frequency environment fingerprint vector when each wireless network access point in the candidate scope performs the target configuration action, and select the wireless network access point as the effective scope from the candidate scope according to the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector.
[0032] This step is based on a confidence score greater than a first confidence threshold. In this case, some or all wireless network access points from the candidate scope are selected as valid scopes. For each wireless network access point in the candidate scope, this step can extract the radio frequency environment fingerprint vector corresponding to its historical execution of the target configuration action, as a reference radio frequency environment fingerprint vector. The reference radio frequency environment fingerprint vector records the radio frequency environment state of the wireless network access point when it executed the target configuration action. This step can also calculate the similarity (e.g., cosine similarity) between each reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector, and then add wireless network access points with similarity reaching the threshold to the valid scope, while the remaining wireless network access points are excluded.
[0033] S104: Apply the target configuration action within the effective scope.
[0034] This step involves sending the instruction for the target configuration action to each wireless network access point within the effective scope, so that the wireless network access points can adjust their parameters and execute the target configuration action according to the instruction.
[0035] This embodiment generates a current radio frequency environment fingerprint vector based on the radio frequency telemetry data of the current wireless network access point. It then combines this fingerprint vector with the current policy of the reinforcement learning core entity to generate decision information. This decision information includes a target configuration action, a confidence score, and candidate scopes. The confidence score represents the confidence level of the target configuration action, and the candidate scopes are the set of wireless network access points where the target configuration action has been applied. When the confidence score is greater than a first confidence threshold, this embodiment selects a wireless network access point from the candidate scopes as an effective scope based on the similarity of the radio frequency environment fingerprint vectors, and then applies the target configuration action within the effective scope. This process uses the radio frequency environment fingerprint as the basis for configuration decisions. By limiting the confidence level and candidate scopes, it achieves precise matching between the configuration action and the actual radio frequency environment, and strictly limits the scope of configuration changes to a consistent and verified range. Therefore, this embodiment can accurately configure wireless network access points based on the radio frequency environment, improving the security and reliability of wireless network configuration.
[0036] As for Figure 1 A further description of the corresponding embodiment: the process of selecting a wireless network access point as an effective scope from the candidate scope based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector includes the following steps A1-A5: Step A1: Determine the fingerprint similarity between each of the reference RF environment fingerprint vectors and the current RF environment fingerprint vector.
[0037] This step can calculate the cosine similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector when each wireless network access point in the candidate scope performs the target configuration action, so as to determine the degree of matching between the two radio frequency environments.
[0038] Step A2: Based on the fingerprint similarity, classify all wireless network access points in the candidate scope into three categories: first type wireless network access points, second type wireless network access points, and third type wireless network access points.
[0039] Wherein, the fingerprint similarity corresponding to the first type of wireless network access point is greater than a first fingerprint similarity threshold; the fingerprint similarity corresponding to the second type of wireless network access point is less than or equal to the first fingerprint similarity threshold and greater than a second fingerprint similarity threshold; and the fingerprint similarity corresponding to the third type of wireless network access point is less than or equal to the second fingerprint similarity threshold. The first fingerprint similarity threshold is greater than the second fingerprint similarity threshold.
[0040] The above process classifies wireless network access points in the candidate scope into three levels based on fingerprint similarity: sites with a similarity greater than the first fingerprint similarity threshold are classified as Class I wireless network access points and can be directly included in the execution scope; sites with a similarity between the second and first fingerprint similarity thresholds are classified as Class II wireless network access points and require manual confirmation before inclusion; sites with a similarity less than or equal to the second fingerprint similarity threshold are classified as Class III wireless network access points and can be directly excluded. This operation enables fine-grained control over the scope of configuration changes.
[0041] Step A3: Add all first-type wireless network access points to the candidate set.
[0042] Step A4: If a manual confirmation message is received, add the second type of wireless network access point corresponding to the manual confirmation message to the candidate set.
[0043] In this step, for the second type of wireless network access point, if the administrator confirms after review that it can perform the target configuration action, a manual confirmation message can be issued; based on the received manual confirmation message, this step adds the corresponding second type of wireless network access point to the candidate set.
[0044] Step A5: Determine whether the number of wireless network access points in the candidate set is greater than the number threshold Y; if yes, set the set of Y wireless network access points with the highest fingerprint similarity in the candidate set as the effective scope; if no, set the candidate set as the effective scope.
[0045] In this step, the maximum number of sites that can be affected by a single configuration change can be predetermined, i.e., a threshold Y. If the number of wireless network access points in the candidate set is greater than Y, it indicates that the potential impact range is too large. In this case, the Y wireless network access points with the highest fingerprint similarity are selected from the candidate set to form the effective scope, prioritizing the configuration change for the sites that best match the environment. If the number of wireless network access points in the candidate set is less than or equal to Y, it indicates that the impact range is within safe limits, and the entire candidate set is directly used as the effective scope.
[0046] The effective scope, also known as the blast radius (the range of network nodes affected when a configuration change fails), or the scope of influence of a configuration change, is used to describe the range of systems, services, users, or business processes affected when a configuration change, deployment, or patch encounters a problem. By performing the above operations, the blast radius of configuration changes can be limited while ensuring configuration accuracy, effectively preventing network stability risks caused by an excessively large blast radius from a single change.
[0047] As for Figure 1 A further description of the corresponding embodiment: the process of generating decision information based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity includes the following steps B1-B7: Step B1: Input the current radio frequency environment fingerprint vector into the reinforcement learning core entity so that the reinforcement learning core entity can determine the target configuration action based on the current policy.
[0048] This step can use the current radio frequency environment fingerprint vector as the state input to the reinforcement learning core entity. The reinforcement learning core entity performs forward inference based on the current policy network, calculates the expected reward of each candidate configuration action, and selects the action with the highest reward as the target configuration action output, thus realizing the mapping from environment perception to configuration decision.
[0049] Step B2: Obtain multiple historical status information.
[0050] Each of the historical state information includes the configuration action a executed at the t-th time step. t Application configuration action a t The obtained performance index values and the RF environment fingerprint vector at time step t.
[0051] This step can retrieve multiple historical status records from the historical database. Each historical status record contains the following information for a specific time step: the configuration action a performed at that time step. t The performance metrics collected after applying the action, and the radio frequency environment fingerprint vector corresponding to the time step.
[0052] Step B3: Calculate the similarity between the reference radio frequency environment fingerprint vector in the historical state information and the current radio frequency environment fingerprint vector, and add the K historical state information entries with the highest similarity to the historical matching set.
[0053] In this step, the cosine similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector in each historical state information can be calculated. Then, the similarity is sorted from high to low, and the K historical state information with the highest similarity is selected to form a historical matching set.
[0054] Step B4: Configure action a in the historical matching set. t The historical state information of the actions configured for the target is added to the reference set.
[0055] This step involves filtering configuration action 'a' from the historical matching set. t Historical state information that matches the target configuration action will be added to the reference set after filtering.
[0056] Step B5: Using the similarity between the reference RF environment fingerprint vector and the current RF environment fingerprint vector as weights, perform weighted statistics on the cases in the reference set where the performance index value meets the preset performance conditions after executing the target configuration action, and obtain the confidence score.
[0057] In this step, for each historical state information entry in the reference set, the similarity between the reference RF environment fingerprint vector and the current RF environment fingerprint vector is used as a weight. This weight quantifies the degree of matching between the environment in which the historical record was located and the current environment; the higher the similarity, the greater the weight. This step judges whether the performance index value after executing the target configuration action meets the preset performance conditions in each historical state information entry: if it meets the conditions (i.e., configuration successful), it is marked as 1; if it does not meet the conditions (i.e., configuration failed), it is marked as 0. The weighted success marks of all historical records are summed to obtain a weighted success score; simultaneously, all weights are summed to obtain the total weight. The weighted success score is divided by the total weight to obtain the confidence score. The confidence score, between 0 and 1, reflects the proportion of expected performance improvement that can be achieved after executing the target configuration action in a similar environment. The higher the confidence score, the higher the credibility of the action in the current environment.
[0058] Step B6: Set the set of wireless network access points corresponding to each historical state information in the reference set as the candidate scope.
[0059] In this step, the wireless network access point corresponding to each historical status information in the reference set can be extracted, and the set of these sites can be taken as the candidate scope, that is, the range of sites that have successfully applied the target configuration action.
[0060] Step B7: Construct the decision information including the target configuration action, the confidence score, and the candidate scope.
[0061] This step encapsulates the determined target configuration action, confidence score, and candidate scope as three configuration description fields into structured decision information output.
[0062] As a feasible implementation, during the process of constructing decision information, summary information can also be generated for the historical state information in the historical matching set, and the summary information can be added to the decision information. Through this process, decision information including the target configuration action, the confidence score, the candidate scope, and the summary information can be constructed. The summary information, also known as the inference path, can specifically include the wireless network access point identifier corresponding to the historical state information, the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector, the performance index value after executing the target configuration action, and the determination result of whether preset performance conditions are met.
[0063] Furthermore, if the confidence score is less than or equal to the first confidence threshold and greater than the second confidence threshold, the decision information is sent to the management device. This decision information includes summary information, which allows the administrator to assess the applicability of the target configuration action in the current environment and decide whether to execute the action. If the administrator decides to execute the target configuration action, an authorization command is issued. Upon receiving the authorization command from the management device, the effective scope can be determined based on the candidate scopes in the decision information, and the target configuration action can be applied within the effective scope.
[0064] Furthermore, if the confidence score is less than the second confidence threshold, the current radio frequency environment fingerprint vector is added to the sample pool, and radio frequency telemetry data and configuration execution results of other wireless network access points with a similarity higher than the preset collection threshold are obtained to increase historical status information; when the newly added historical status information is greater than the preset accumulation quantity, the confidence score of the target configuration action is recalculated.
[0065] When the confidence score falls below the second confidence threshold, the above operation searches the entire network for other wireless network access points with a fingerprint vector similarity to the current radio frequency environment that is higher than a preset collection threshold. For these similar sites, radio frequency telemetry data and performance index results after various configuration actions were performed in the past are collected, and this newly added data is added to the sample pool as historical state information. When the number of newly added historical state information in the sample pool reaches a preset accumulation threshold (e.g., 30 entries), it indicates that there is sufficient experience with similar environments for reference. At this time, the confidence score calculation process can be re-executed based on the expanded historical state information to update the confidence score of the target configuration action in the current environment.
[0066] As for Figure 1 In a further description of the corresponding embodiment, after applying the target configuration action within the effective scope, performance metrics after applying the target configuration action can also be obtained; a reward function value can be calculated based on the performance metrics, and the reward function value can be normalized using the L2 norm of the current radio frequency environment fingerprint vector; the current policy of the reinforcement learning core entity can be updated based on the normalized reward function value.
[0067] In the above process, to eliminate the impact of differences in radio frequency environment complexity between different sites on the magnitude of the reward, the reward function value is normalized using the L2 norm of the current radio frequency environment fingerprint vector, so that the reward signal is comparable under different sites and different radio frequency complexities, which is beneficial to the convergence of reinforcement learning algorithms.
[0068] The aforementioned performance metrics include signal-to-noise ratio change, roaming duration change, channel utilization offset, number of coverage holes, association failure rate, and number of abnormal devices. The channel utilization offset describes the degree of deviation between the actual channel utilization and the target channel utilization, and the number of abnormal devices describes the number of wireless network access points that become abnormal after applying the target configuration action.
[0069] The signal-to-noise ratio change is the difference between the signal-to-noise ratio received by the client after the application target configuration action and before the configuration. The roaming time change is the change in the time spent by the terminal switching between wireless network access points before and after the configuration. The channel utilization offset is the difference between the current channel utilization and the preset target utilization. The number of coverage holes is the number of locations or areas within the effective scope where the client signal strength is continuously lower than the association threshold and cannot associate with other wireless network access points. The association failure rate is the proportion of clients that are rejected or timed out when trying to associate with wireless network access points per unit time. The number of abnormal devices is the number of wireless network access points that experience association failure, rate degradation, or disconnection due to the application target configuration action.
[0070] After applying the target configuration action within the effective scope, performance metrics before and after applying the target configuration action can be collected within a preset time window. It can then be determined whether the performance metrics before and after applying the target configuration action meet the anomaly criteria. If so, a rollback operation is performed to restore the configuration to its state before applying the target configuration action, and the reward function value corresponding to the target configuration action is modified to a preset penalty value. The current radio frequency environment fingerprint vector, the target configuration action, and the modified reward function value are set as penalty samples to update the current policy of the reinforcement learning core entity based on the penalty samples.
[0071] Specifically, this step can initiate a verification time window (e.g., 30–60 minutes) after the configuration takes effect, continuously collecting performance metrics (such as client signal-to-noise ratio, roaming latency, channel utilization, association failure rate, etc.). If performance metrics show degradation, it is determined to meet the anomaly judgment conditions. At this time, the AP radio frequency parameters can be automatically restored to their original values before executing the target configuration action, preventing the fault from continuously affecting users. This step can also set the reward function value corresponding to this decision as a preset penalty value, and set the current radio frequency environment fingerprint vector, the target configuration action, and the modified reward function value as penalty samples. Based on the penalty samples, the current policy of the reinforcement learning core entity is updated, thereby preventing the policy from repeatedly triggering the same erroneous decision on wireless network access points with the same fingerprint type.
[0072] This step can compare the mean or distribution of each performance index sequence before and after applying the target configuration action. If the test results show that the performance index has deteriorated significantly, it is determined to meet the anomaly judgment conditions.
[0073] The process described in the above embodiments is illustrated below through examples in practical applications.
[0074] With the large-scale deployment of Wi-Fi 6 (802.11ax) and Wi-Fi 6E (802.11ax 6GHz), the complexity of RF (Radio Frequency) configuration management has increased dramatically, involving coordinated decisions across multiple dimensions, including channel selection, Received Signal Threshold (RxSOP), Fast BSS (Basic Service Set) switching, Flexible Radio Architecture (FRA), and security protocols (WPA3). Wi-Fi 6 (802.11ax) and Wi-Fi 6E (802.11ax 6GHz) represent wireless local area network (WLAN) communication technologies.
[0075] The related technical solution CN113302980B (A channel allocation method, system, electronic device and storage medium, inventors Lei Yongcheng and Wu Fang) has the following limitations: First, it only addresses the single configuration dimension of channel selection and cannot cover the full set of RF configuration parameters such as 802.11r (Fast BSS switching), RxSOP (Receive Start Packet Detection Threshold), FRA (Flexible Radio Architecture), and WPA3 (Third Generation Wireless Security Protocol); Second, it uses the channel switching history at the single AP level as a state representation, lacking multi-dimensional perception of the overall RF physical environment characteristics of the site; Third, the configuration recommendation lacks a confidence quantification constraint mechanism, and the scope of the configuration change is uncontrolled, posing a risk of an excessively large blast radius; Fourth, when the same configuration strategy is applied indiscriminately to all sites (such as simultaneously pushing the same WPA3+802.11r+RxSOP configuration to administrative buildings, teaching buildings, laboratories, and stadiums in a university campus), serious problems such as association failure and coverage gaps may occur in some sites due to differences in the local RF environment.
[0076] In related technologies, existing wireless network radio frequency configuration management commonly adopts the Golden Template approach, which involves pre-defining a globally unified set of radio frequency configuration parameters and pushing it in batches to all access points and networks under its jurisdiction. This static configuration approach cannot adapt to heterogeneous radio frequency environments, leading to performance degradation. Real-world test data shows that in a university (19 buildings, 930 APs, approximately 14,600 clients), using the Golden Template deployment, the interference rate in the 2.4GHz band exceeded 36%, the roaming latency in the 5GHz band was approximately 470ms, and the utilization rate of the 6GHz band was only 1% due to client access failures. All of these problems stem from a systematic mismatch between the configuration intent and the physical radio frequency environment.
[0077] To achieve adaptive configuration recommendation with quantitative confidence constraints and precise blast radius control across the entire set of radio frequency configuration parameters, driven by the site-specific radio frequency physical environment fingerprint, this embodiment provides a confidence-bounded adaptive configuration system and method based on wireless radio frequency environment fingerprints. This system, driven by the site-specific radio frequency environment fingerprint across the entire set of radio frequency configuration parameters, achieves secure, explainable, and online self-evolving wireless network configuration recommendation through confidence gating and precise blast radius constraints. This addresses core issues in existing technologies such as mismatch between static configuration and heterogeneous radio frequency environments, uncontrollable risks in automated decision-making, and lack of online closed-loop learning.
[0078] The core problem with using the Golden Template for configuration is the systematic discrepancy between the static configuration intent and the heterogeneous RF environment, and the uncontrolled blast radius of configuration changes. Using the Golden Template in a university campus network resulted in a roaming latency of 470ms, an interference rate >36%, and a 6GHz utilization rate of 1%, leading to situations where the administration building functioned normally, classroom association failed, laboratory roaming failed, and the stadium had coverage vulnerabilities. In the RF fingerprint-driven adaptive configuration scheme provided in this application embodiment, the following operations can be executed sequentially: telemetry data (channel interference rate, noise floor, channel utilization, client signal-to-noise ratio, and time-series heatmap) acquisition, environmental fingerprint vector generation (using CNN-LSTM encoding to generate unique site identifiers), decision contract generation, and blast radius control. After applying the configuration actions, the fast roaming and WPA3 functions of the administration building functioned normally, the on-demand activation and security enhancement functions of the classrooms functioned normally, the precise configuration and IoT (Internet of Things) gate control functions of the laboratory functioned normally, and the coverage optimization and FRA (Flexible Radio Architecture) activation of the stadium functioned normally. By employing a strategy of bounded blast radius, interpretable decision-making, and continuous online evolution, this solution reduces interference rate by 41%, improves 6GHz SNR by 100%, and reduces roaming latency to 46ms (a 93% improvement). The aforementioned decision contract, or decision information, includes target configuration actions, inference paths, confidence scores, and candidate scopes. After confidence gating, this solution can determine the blast radius based on RF fingerprint similarity; the blast radius is the effective scope.
[0079] This solution provides an adaptive configuration method based on wireless radio frequency environment fingerprinting, as follows: Multidimensional radio frequency (RF) telemetry data, including channel interference rate, noise floor, channel utilization, client signal-to-noise ratio, and time-series heatmaps, is collected from multiple wireless network access points. This multidimensional RF telemetry data is then input into an RF environment fingerprint generation engine, which outputs a site-specific RF environment fingerprint vector F(s) using a CNN-LSTM combined neural network. The RF environment fingerprint vector F(s) is then input into a confidence-bounded decision engine, which generates a decision contract including a proposed configuration action A, an inference path R, a confidence score C, and a candidate scope S. A three-level gating judgment is performed based on the confidence score C, and the effective scope S is calculated based on the RF fingerprint similarity. eff Only for S eff The network subset is configured to perform changes; experience metrics are collected after the configuration is executed, reward signals are calculated, and reinforcement learning strategies are updated.
[0080] The adaptive configuration system architecture based on wireless radio frequency environment fingerprinting provided in this embodiment is compatible with the overall framework of related technology CN113302980B, and introduces two new components: a radio frequency environment fingerprint generation engine and a confidence-bounded decision engine. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 The present application provides an adaptive configuration system architecture diagram based on wireless radio frequency environment fingerprinting, which includes a wireless network access point (AP) and a cloud server / cloud AI platform. A wireless network access point includes the following sub-modules: Data forwarding agent: Used to receive various types of data from below (currently available channel scan results, network performance statistics, WLAN event data), and forward network status information upwards to the cloud server / cloud AI platform.
[0081] Wireless radio frequency: Used to support the 802.11a / b / g / n / ac / ax / be protocol and provide radio frequency side data to the data forwarding agent.
[0082] The WLAN network is used to receive WLAN driver logs and provide network-side data.
[0083] The cloud server / cloud AI platform includes the following sub-modules: Data collector: Used to receive network status information uploaded by wireless network access points.
[0084] The radio frequency environment fingerprint generation engine contains a state mapping entity, a 4-layer CNN convolutional layer, a fully connected FC layer, and an LSTM temporal layer. The radio frequency environment fingerprint generation engine outputs the radio frequency environment fingerprint F(s) to the decision engine.
[0085] The confidence-bounded decision engine contains a reward-generating entity, a confidence-evaluating entity, a decision contract generator, and an explosion radius controller. The input of the confidence-bounded decision engine is the radio frequency environment fingerprint F(s), and the output is the decision contract.
[0086] Configure the recommender publisher and database to receive contract information, combine it with the database, and output it to the reinforcement learning core.
[0087] The core entity of reinforcement learning includes the Q(s,a) policy network, which modifies the policy according to the reward R and outputs the updated policy to the configuration distribution module, which finally distributes it to the wireless network access point.
[0088] This embodiment can be applied to a cloud server connected to a wireless network access point. The cloud server can return configuration action recommendations to the wireless network access point based on the multi-dimensional network status information uploaded by the access point. Specifically, the wireless radio frequency is configured to operate the local WLAN network to provide network access services to multiple WLAN clients. The WLAN network can generate WLAN driver logs, structured WLAN event data, and network range-specific and client-specific network performance statistics, including data rate, packet statistics (retries, loss, decoding errors), and radio signal strength (SNR). WLAN stands for Wireless Local Area Network.
[0089] The data forwarding agent collects all the aforementioned telemetry data and periodically forwards the data to the data collector in the cloud server. The sum of all the data collected by the data collector is called network state information. The machine learning application manager receives notifications of network state information availability from the data collector and coordinates the training and inference operations of the RF environment fingerprint generation engine and the confidence-bounded decision engine.
[0090] Compared with the related technical solution CN113302980B, the difference of this embodiment is that: this embodiment introduces a CNN-LSTM joint neural network in the state mapping entity to map the network state information into a site-specific radio frequency environment fingerprint vector F(s) instead of a simple state vector S; a confidence evaluation mechanism and blast radius control logic are introduced in the reward generation entity; and the decision-making object is expanded from a single AP channel to the entire set of RF configuration parameters (802.11r, RxSOP, FRA, WPA3, data rate benchmark, etc.).
[0091] Please see Figure 3 , Figure 3 Another adaptive configuration system architecture based on wireless radio frequency environment fingerprinting provided in this application embodiment includes: Access Point Layer: Contains multiple access points (AP1, AP2, ..., APn). This layer is responsible for RF PHY / MAC telemetry data acquisition; RF PHY stands for Radio Frequency Physical Layer, and MAC stands for Media Access Control Layer.
[0092] RF environment fingerprint generation module: This module receives low-level telemetry data and processes it using a CNN (4-layer convolutional), FC, and LSTM network structure. The output of this module is a site-specific RF fingerprint vector F(s).
[0093] The decision contract generation module receives the input: radio frequency environment fingerprint F(s) plus RL policy (reinforcement learning policy), and outputs the decision contract containing {action, inference path, confidence score C, scope S}.
[0094] The explosion radius control module receives the output from the decision contract generation module and implements confidence gating. When the confidence score C > the first confidence threshold θ, the gating is applied. high When the second confidence threshold θ is reached, automatic execution is triggered; low <Confidence score C ≤ first confidence threshold θ high When manual review is triggered, this module can use RF fingerprint similarity to accurately identify the subset of the network in action.
[0095] The verification and continuous learning module receives input from the blast radius control module and performs short-term statistical baseline verification and automatic rollback mechanism. This module defines reward conditions: SNR improvement, roaming latency reduction, and utilization optimization; and penalty conditions: coverage gaps, association failures, and starvation of legacy equipment.
[0096] Configure the distribution layer by confidence level: Enable 802.11r (network 13), FRA (network 16), WPA3Transition (network 13), and disable 3 / 1 / 0 networks. Transition indicates transition mode.
[0097] Real-world test results: Data from a university campus (19 buildings, 930 APs, 14,600 clients); 2.4GHz roaming latency: reduced from 470ms to 46ms (a 93% decrease); 5GHz interference rate: decreased by 41%; 6GHz SNR: increased by 100%; blast radius: precise and bounded.
[0098] Figure 3 The paper demonstrates the collaborative relationship between the RF environment fingerprint generation engine, the confidence-bounded decision engine, the reinforcement learning core entity, and the verification and rollback entity.
[0099] This embodiment provides an adaptive configuration method based on wireless radio frequency environment fingerprinting, which includes the following steps C1-C5: Step C1: Multidimensional radio frequency telemetry data acquisition.
[0100] This step can collect the following multi-dimensional RF telemetry data periodically (every 5 minutes by default) from each wireless network access point as feature input: Channel Interference Ratio (CIR): This is the percentage of interference from non-local networks in the current channel and adjacent channels, ranging from 0 to 100%. In this embodiment, the initial interference rate for the 2.4 GHz band is approximately 36%, which is considered a high-interference state.
[0101] Noise Floor: The background electromagnetic noise power of the current radio frequency environment, measured in decibels and milliwatts (dBm), reflecting the electromagnetic cleanliness of the deployment environment.
[0102] Channel Utilization (CU): The air interface time occupancy rate of all devices (including neighboring APs and WLAN clients) on the current channel, with a value ranging from 0 to 100%.
[0103] Client SNR / SiNR: A downlink quality metric measured in decibels (dB). In this embodiment, the client SNR before optimization is approximately 29dB on the 2.4GHz band, while the 6GHz band has no valid data due to the lack of client access.
[0104] Temporal Heatmaps: The above indicators are distributed in a 30-minute time series over a 24-hour period to capture the network's busy / idle period (weekday peak 8:00-22:00, weekend patterns, etc.) and the dynamic evolution of interference.
[0105] In this embodiment, network status information includes, but is not limited to, the aforementioned channel interference rate, noise floor, channel utilization, client SNR, and time series heatmap, and can be further extended to include unstructured data sources such as WLAN driver logs and WLAN event data. Multidimensional RF telemetry data is extracted and calculated from raw data such as network performance statistics and WLAN event data.
[0106] Step C2: Generate the radio frequency environment fingerprint vector.
[0107] The state mapping entity receives the multidimensional radio frequency telemetry data, and the radio frequency environment fingerprint generation engine outputs the RF environment fingerprint vector F(s) through the following three-level processing: Level 1: CNN Spatial Feature Extraction. The CNN subnetwork contains at least four convolutional layers, with the number of kernels in each layer set to 64, 128, 128, or 256, and a kernel size of 1×3. ReLU (Rectified Linear Function) is used as the activation function. Using the interference rate sequence, noise floor sequence, channel utilization sequence, and client SNR sequence as four independent input channels (analogous to the RGB channels of an image), local pattern extraction is performed on the temporal data. The output dimension of the convolutional neural network (CNN) is K. CNN (Default 256) Spatial feature map. Features captured by the CNN subnetwork include: interference distribution patterns within a specific time period, spatial structure characteristics of channel occupancy, and interference coupling relationships between neighboring APs.
[0108] Level 2: Fully Connected Dimensionality Reduction and Fusion. After flattening the feature map output by the CNN, the fully connected layer (FC Layer) reduces the feature dimension to an intermediate representation vector of dimension 256 through two fully connected layers (hidden layer dimensions of 512 and 256 respectively, using random dropout to prevent overfitting).
[0109] Level 3: LSTM Temporal Modeling. The Long Short-Term Memory (LSTM) network takes a temporal sequence (window length defaults to 48 time steps, corresponding to a 24-hour / 30-minute granularity) as input, and models the temporal dynamics of the RF environment through the forget gate, input gate, and output gate mechanisms of the LSTM unit. It captures busy-idle periodic patterns, interference dynamic evolution patterns, and channel contention timing characteristics, outputting a temporal-aware feature vector. The LSTM network has a dimension of K. LSTM (Default 256).
[0110] RF environment fingerprint vector F(s) = [f1, f2, ..., f K ], where K=K CNN +K LSTM =512, representing the concatenation of CNN spatial features and LSTM temporal features. F(s) is stored in a fingerprint database indexed by site identifiers (Site IDs) and updated periodically. When the L2 norm change of F(s) exceeds a preset drift threshold (||F(s)... new )-F(s old )||2>drift threshold This triggers the policy reassessment process. F(s) new F(s) represents the currently acquired radio frequency environment fingerprint vector. old F(s) represents the previously acquired radio frequency environment fingerprint vector, ||F(s) new )- F(s old )||2 represents the L2 norm change of the radio frequency environment fingerprint vector, drift thresholdThis indicates the preset drift threshold.
[0111] Please see Figure 4 , Figure 4 This is a schematic diagram of the architecture of a radio frequency environment fingerprint generation engine provided in an embodiment of this application. The radio frequency environment fingerprint generation engine is implemented based on a CNN-LSTM joint neural network architecture, including an input layer, a CNN network layer, a fully connected layer, an LSTM layer, and a radio frequency fingerprint output layer. The input layer can respectively input the interference rate sequence, the noise floor sequence, the channel utilization sequence, and the client signal-to-noise ratio sequence.
[0112] A CNN network consists of convolutional layer 1, convolutional layer 2, convolutional layer 3, and convolutional layer 4. Fully connected layers can perform fully connected dimensionality reduction and fusion; The LSTM layer serves as a time memory unit, enabling dynamic time capture. The radio frequency environmental fingerprint output by the radio frequency fingerprint output layer is the result of spatiotemporal feature fusion and serves as a unique identifier for the site.
[0113] The spatial features extracted by CNN include interference distribution patterns and channel occupancy spatial structure; the temporal patterns fused by LSTM include busy / idle periodicity and dynamic evolution of interference. F(s) serves as a unique environmental identifier for a site, distinct from the device RF fingerprint, which characterizes hardware manufacturing differences.
[0114] Step C3: Generate a decision contract.
[0115] The reward-generating entity receives the RF environment fingerprint vector F(s) and the current policy pi(a|s) of the reinforcement learning core entity, and generates a decision contract through the following sub-entities: State history entities: Store the most recent N in a first-in, first-out (FIFO) manner. hist (Default 1000) RF fingerprint status records {s t ,a t ,r t ,F(s t This is analogous to the function of the state history entity in CN113302980B, but the storage granularity is expanded to multi-dimensional RF configuration decision records. The above s t Let a represent the environment state at time step t. t This represents the configuration action executed by the agent at time step t, r. t F(s) represents the reward signal received at time step t. t ) represents the radio frequency environment fingerprint vector at time step t.
[0116] Fingerprint matching entities: Retrieve the top K entities from the state history that have the highest cosine similarity to the current radio frequency environment fingerprint F(s) from the current radio frequency environment entities. match(Default 20) historical records are used to extract the configuration actions and corresponding experience metric results for each record, forming a historical matching set {(a j ,outcome j |j=1,...,K match}. The above a j Indicates historical configuration actions, outcome j This represents the historical experience result, and j represents the sequence number.
[0117] Confidence score calculation entity: The success rate of target action A in the historical matching set is weighted and statistically analyzed. The confidence score C = sum(w j ×I(outcome j >threshold j )) / sum(w j ), where the weight w j =sim(F(s j ),F(s)) β (β is the attenuation coefficient, default 2.0), I(·) is the indicator function, threshold j The threshold for improving the corresponding experience metrics is defined by sum, where sum represents the summation function and sim represents the cosine similarity function.
[0118] Contract output entity: Proposed configuration action A, inference path R (including K) match The summary of historical matching records, confidence score C, and candidate scope S (the union of the sets of sites in the historical matching set that successfully applied the action) are encapsulated into a standardized decision contract format for output.
[0119] The contract format can be described as follows: Contract={Action A :[(param k ,val k Rationale R :[{site id ,sim score ,outcome}],Confidence C :float,Scope S :[site id ]}.
[0120] Contract represents a decision-making contract, Action A Indicates the proposed configuration action A, (param k ,val k ) represents a list of parameter key-value pairs, Rationale R Represents the reasoning path R, siteid The site identifier, sim score The similarity score represents the similarity score, and the outcome represents the execution result (a record of the performance metrics or effect description after applying this configuration action to the historical site). Confidence C Indicates the confidence score C, float represents a floating-point number, and scope. S Denotes the candidate scope S, site id Represents a list of site identifiers.
[0121] Step C4: Three-level confidence gating.
[0122] This embodiment can set differentiated confidence thresholds to achieve the following three levels of gating: High confidence level (C>θ) high , default θ high =0.85): The system automatically sends configuration action A to the blast radius assessment process without manual intervention. In this embodiment, the confidence score recommended by the Fast BSS Transition in a certain university campus is 0.91, which meets the automatic execution conditions.
[0123] Medium confidence level (θ) low ≤C≤θ high , default θ low =0.6, θ high =0.85): The system sends the decision contract to the administrator via a message push system. The administrator can review the complete inference path and decide whether to execute it. This scenario is analogous to the judgment logic of training completion and incompleteness in CN113302980B, reflecting the principle of human-machine collaboration.
[0124] Low confidence level (C<θ) low , default θ low =0.6): The system stores the current RF environment fingerprint status into the low confidence sample pool, triggers data collection from more similar RF environment fingerprint sites, and recalculates the confidence score after accumulating enough historical matching records (at least 30 by default).
[0125] Step C5: Explosion radius assessment and precise clearance.
[0126] The explosion radius control sub-entity executes the following algorithm: RF fingerprint similarity calculation: For each station (i.e., wireless network access point) in the candidate scope S, s i Calculate the cosine similarity sim(F(s)) between its RF fingerprint and the current site's RF fingerprint. i ),F(s)): sim(F(s i ),F(s))=F(s i)·F(s) / (||F(s i )||2×||F(s)||2); F(s i F(s) represents the RF environment fingerprint vector of the i-th wireless network access point, and F(s) represents the RF environment fingerprint vector of the current wireless network access point. i )||2 represents F(s) i Let ||F(s)||2 represent the L2 norm of F(s).
[0127] Effective scope filtering: Constructing effective scope S eff ={s i ∈S:sim(F(s i ),F(s))≥δ}, where δ is the similarity threshold (default 0.85). For similarities within [δ... low Sites with similarity in the range of δ (default 0.6 to 0.85) are marked as a cautionary zone and require manual confirmation. Sites with similarity below δ... low Sites with a default version of 0.6 are forcibly excluded (application configuration is prohibited).
[0128] Explosion radius upper limit constraint: The system also sets an upper limit B for the explosion radius. max (By default, it does not exceed 80% of the entire network), when |S eff | / N total >B max At that time, from S eff Select the results in descending order of similarity, prioritizing those that satisfy B. max A subset of constraints ensures the controllability of the impact ratio of a single configuration change. total This indicates the total number of wireless network access points.
[0129] In this embodiment, when configuring Fast BSS Transition for university campuses, the system identified three networks with a radio frequency environment fingerprint similarity of less than 0.85 (these three networks contained a large number of older 802.11a / b / g devices, whose RF fingerprint characteristics showed low client SNR distribution and high device type entropy). These networks were excluded from the effective scope, and Fast BSS Transition was ultimately enabled on 13 networks. Roaming latency decreased from approximately 470ms to approximately 46ms (an improvement of approximately 93%). Simultaneously, the three excluded networks maintained their original configurations, avoiding the problem of old device association failures. 802.11a / b / g represents a wireless LAN standard.
[0130] The reward function for the core entity in reinforcement learning is explained below: This embodiment designs a reward function for the entire set of RF configuration parameters. The loop description of the entire mechanism is as follows: 1) Obtain a snapshot M of the experience metrics after configuration execution from the data collector. t ;2) The entity that generates the reward calculates the reward signal R. t 3) Reinforcement learning core entities based on R t Update the policy pi(a|s); 4) Under the guidance of the new policy, the wireless radio frequency continues to collect WLAN environment feedback to provide new input network status information.
[0131] Reward function R t The definition is as follows: R t =f(M t M t-1 )[+]g(E t ), f≥0, f nondecreasing; Where, f(M) t M t-1 ) represents the positive reward function, [+] is the summation or product operator, g(E t ) indicates a negative penalty function; "f nondecreasing" indicates that f is a non-decreasing function. M t This represents a snapshot of current experience metrics (including SNR, roaming latency, and channel utilization), M t-1 Represents a snapshot of the benchmark experience metric, E t This represents a set of negative events (such as coverage holes, association failures, old device starvation, etc.). f(M t M t-1 )=h1(dSNR t )[+]h2(dLat t )[+]h3(dUtil t ); dSNR t =SNR t -SNR t-1 dSNR t This indicates the improvement in client-side SNR, in dB; dLat t =Lat t-1 -Lat t dLat t dUtil represents the improvement in roaming latency, in milliseconds (ms). t =|CU target -CU t |-|CU target -CU t-1 |,dUtil t CU represents the amount by which channel utilization converges towards the target interval. targetThe target utilization rate is set at 30% to 45% by default. SNR t and SNR t-1 Represents the client SNR at different times, Lat t-1 -Lat t CU represents the roaming latency at different times. target CU represents the target range for channel utilization. t-1 and CU t This represents the channel utilization rate at different times.
[0132] The above h1, h2, and h3 are positive non-decreasing functions of their corresponding inputs. Their invariance means that the probability of the reinforcement learning engine continuously choosing configuration actions that improve SNR, reduce roaming latency, and rationalize channel utilization continues to increase. From the perspective of minimizing service interruption, the user experience quality under this configuration is higher, which is often overlooked by traditional static configuration schemes.
[0133] The above h1 is the SNR improvement function, which is positive and non-decreasing. When dSNR>=0, h1(dSNR) increases positively. h2 is the roaming delay improvement function, which is positive and non-decreasing. When dLat>=0, h2(dLat) increases positively. h3 is the channel utilization rationalization function, which increases positively when the channel utilization Util approaches the target range.
[0134] g(E t )=-[p1(CovHole t )[+]p2(AssocFail t )[+]p3(LegacyStarv t )]; Here, p1, p2, and p3 are positive non-decreasing functions of their respective inputs. p1 captures the negative impact of coverage holes (the more coverage holes, the greater the penalty); p2 captures the negative impact of increased 802.11 association failure rate (the higher the failure rate, the greater the penalty); p3 captures the negative impact of traditional 802.11 (a / b / g / n) device starvation (the higher the proportion of affected older devices, the greater the penalty). Specifically, p1 is the coverage hole penalty function, positive non-decreasing, with the penalty increasing when a hole appears; p2 is the association failure rate penalty function, positive non-decreasing; and p3 is the traditional 802.11 device starvation penalty function.
[0135] CovHole t This indicates the number of coverage holes, specifically the number of areas where wireless signal coverage is interrupted due to configuration changes detected at time step t. (AssocFail) t This represents the association failure rate, specifically the 802.11 association failure rate at time step t, which is the proportion of clients unable to associate with the AP due to configuration changes. (LegacyStarv) tThis indicates the hunger level of traditional devices, specifically the proportion of older 802.11 (a / b / g / n) devices that are unable to connect normally due to new configurations (such as forced WPA3) at time step t.
[0136] Reward normalization: R t_norm =R t / ||F(s t )||2, normalize the reward value using the L2 norm of the current site's RF fingerprint vector to eliminate the dimensional inconsistency caused by differences in RF environment complexity between heterogeneous sites, ensuring the convergence of cross-site reinforcement learning training. The above operation normalizes the reward value according to the RF fingerprint magnitude to eliminate dimensional differences between sites. R t_norm R represents the normalized reward value. t / ||F(s t )||2 represents the original reward value R t Divide by the current site's radio frequency environment fingerprint vector F(s) t The result obtained from the L2 norm of ).
[0137] If the core entity in reinforcement learning is considered to have achieved sufficient performance within a certain period, the online update part of the reward-generating entity can be omitted to save processing resources, and the current policy pi(a|s) can be used directly for inference output.
[0138] Please see Figure 5 , Figure 5 This is a schematic diagram of a reinforcement learning principle provided in an embodiment of this application. The process defines the state s as an RF fingerprint vector. This embodiment can use the Bellman equation for reinforcement learning.
[0139] Figure 5 The diagram illustrates a reinforcement learning agent, a configuration action delivery module, a wireless network environment, and an experience feedback module. The reinforcement learning agent has feature extraction capabilities (such as feature extraction using a CNN-LSTM network structure), state awareness capabilities, a policy output module, and an action space module.
[0140] The action space module contains several specific wireless network configuration actions, such as: enabling / disabling 802.11r, adjusting the data rate baseline, switching WPA3 Transition on and off, adjusting the RxSOP threshold, and switching FRA on and off.
[0141] The configuration action delivery module is used to push information in a hierarchical manner based on confidence level and fingerprint similarity.
[0142] The wireless network environment includes wireless network access point clusters and clients. Experience metrics include: client SNR, roaming latency, channel utilization, association success rate, and coverage hole detection results.
[0143] The reward signals for the above process include: a significant increase in client SNR (exceeding the threshold), a reduction in roaming latency (e.g., from 470ms to 46ms), a rationalization of channel utilization (approaching medium load), and an enhancement of downlink signal strength.
[0144] The penalty signals for the above process include: the appearance of coverage holes, the increase in roaming failure rate, the failure of association / starvation of traditional 802.11 equipment, and the deterioration of co-channel interference rate.
[0145] This embodiment also includes the following verification and automatic rollback mechanisms: After configuration execution, the verification and rollback entities are displayed in the statistics window T. valid (Default 30 to 60 minutes) Continuously collect key experience indicator sequences {SNR} t ,Lat t ,CU t AssocFail t SNR t Indicates the client signal-to-noise ratio, Lat t Indicates roaming latency, CU t AssocFail represents channel utilization. t This represents the association failure rate. If any core indicator shows a statistically significant deterioration (using a t-test, significance level alpha=0.05, comparing the T values before and after configuration execution),... valid If the difference between the mean and the mean of the current time window is less than or equal to 2, a rollback to the previous configuration state will be automatically triggered, and the reward for this execution will be set to R. min (Default -1.0), penalized samples are included in the RL (Reinforcement Learning) training set, and the Q-value table (action value table) is updated through the Bellman equation to prevent the policy from repeatedly triggering the same erroneous decision on sites with the same RF fingerprint type: Q(s,a)<-Q(s,a)+alpha×[R min +gamma×max a'Q(s',a') -Q(s,a)]; Q(s,a) represents the expected cumulative discounted reward obtained by performing action a in state s (RF environment fingerprint vector); Q(s',a') represents the expected cumulative discounted reward obtained by performing action a' in state s'; <- indicates assignment; alpha represents the learning rate; gamma represents the discount factor; max a'Q(s',a') R represents the optimal future value; min This indicates the minimum penalty reward.
[0146] In this embodiment, the short-term statistical baseline validation can be implemented using statistical methods including, but not limited to, t-test (Student's t-test), CUSUM (cumulative sum control chart) test, and Bayesian online change point detection. Preset operations include any one or a combination of several of the following: linear summation, logarithmic summation, linear filtering, nonlinear filtering, time series analysis, and trend analysis.
[0147] The adaptive configuration method based on wireless radio frequency environment fingerprinting in this embodiment includes the following steps: Step D1: Periodically collect multi-dimensional RF telemetry data from multiple wireless network access points and WLAN networks through a data forwarding agent. The multi-dimensional RF telemetry data includes channel interference rate, noise floor, channel utilization, client signal-to-noise ratio (SNR), and time series heatmap. Step D2: Input the multidimensional RF telemetry data into the state mapping entity in the RF environment fingerprint generation engine. The state mapping entity uses a convolutional neural network (CNN) to extract spatial features from the multidimensional RF telemetry time series. After dimensionality reduction by a fully connected layer, the temporal dynamic features are captured by a long short-term memory network (LSTM) to output an RF environment fingerprint vector F(s) that characterizes the uniqueness of the site's physical RF environment. Step D3: Input the RF environment fingerprint vector F(s) into the reward generating entity in the confidence bounded decision engine. The reward generating entity generates a decision contract {A,R,C,S}, which includes: proposed configuration action A, inference path R, confidence score C, and candidate scope S. Step D4: Confidence gating judgment: If C is greater than the high threshold θ high Then proceed to the explosion radius assessment. If θ low Less than or equal to C Less than or equal to θ high Then it will be sent to a human reviewer. If C is less than θ low This will temporarily suspend and trigger further data collection; Step D5: Explosion radius assessment: Calculate the RF fingerprint similarity sim(F(s) of each network site in the candidate scope S). i ), F(s)), to filter the effective scope S with similarity greater than or equal to the similarity threshold δ. eff Only for the effective scope S eff The network subset is issued configuration action A; Step D6: After configuration execution, collect experience metrics within a short-term statistical window and calculate the reward signal R. t And update the policy pi(a|s) in the core entity of the reinforcement learning.
[0148] The RF environmental fingerprint vector F(s) consists of K-dimensional features, which are obtained by concatenating spatial features (K1-dimensional) extracted by CNN and temporal features (K2-dimensional) extracted by LSTM, i.e., K = K1 + K2. To calculate the similarity between different states, the cosine similarity formula sim = F(s) is used. i )·F(s j ) / (||F(s i )||2×||F(s j The value of 1 is [-1, 1], with values closer to 1 indicating greater environmental similarity. The system employs different processing strategies based on the similarity score. The first fingerprint similarity threshold is set to delta. high Set the second fingerprint similarity threshold to delta. low When sim <delta low When the environment is deemed significantly different, the application configuration is prohibited; when delta low ≤sim <delta high When sim ≥ delta, it is judged to be of moderate similarity in the environment and should be applied with caution after manual confirmation; high When the environment is deemed highly similar, it is included in the effective scope S. eff The threshold value is typically set as delta. low =0.6, delta high =0.85 (configurable). Furthermore, by constructing a similarity matrix SIM[i][j]=sim(F(s) i ),F(s j The set of effective scopes S is defined as follows: i, j = 1...N. eff ={i:sim(F(s i ),F(s))≥delta high}, and set an upper bound constraint on the explosion radius |S eff | / N total ≤B max |S eff | indicates the size of the effective scope, N total B represents the total number of websites. max This indicates the configurable upper limit.
[0149] In the university case study, RF fingerprint clustering and configuration applications divided the network into four clusters: the administration building cluster and the teaching building cluster both had a similarity ≥ 0.85, and 802.11r and WPA3, and FRA and 802.11r were enabled respectively; the laboratory cluster had a similarity between 0.6 ≤ sim < 0.85, and was marked as requiring manual confirmation and subject to IoT gating; the stadium cluster had a similarity < 0.6, and was prohibited from application, with 802.11r disabled. The final measured allocation results showed that Fast BSS Transition enabled 13 networks / disabled 3 networks, FRA enabled 16 networks / disabled 1 network, and WPA3 enabled 13 networks / disabled 0 networks.
[0150] Please see Figure 6 , Figure 6 This is a schematic diagram of a two-dimensional decision-making framework for bounded confidence and explosion radius control provided in an embodiment of this application. The input of the framework includes an RF fingerprint vector F(s) and a reinforcement learning policy π. After processing, the input generates a decision contract, which contains specific elements: {action A, inference path R, confidence C, candidate scope S0}.
[0151] The confidence level internal control judgment module calculates the confidence level score; If the confidence score is less than θ low If the recommendation is not approved, more data collection will be triggered, and the evaluation will be re-evaluated after the confidence level is increased. If the confidence score is at θ low and θ high In between, a manual review mode is used to push the decision contract to the administrator, who will then confirm and execute it.
[0152] If the confidence level is greater than θ high If so, the automatic execution mode will be activated, and the explosion radius assessment process will begin.
[0153] In the blast radius assessment process, the RF fingerprint similarity of each network can be calculated, and a subset of networks with a similarity greater than a threshold δ can be selected. Configuration changes are then performed on this subset of networks, while the remaining networks remain unaffected.
[0154] Below is a data reference from a university case study: Fast BSS Transition: Enable 13 / Disable 3; FRA: Enable 16 / Disable 1; WPA3 Transition: Enable 13 / Disable 0.
[0155] In the verification and rollback mechanism, a statistical window test of 30-60 minutes is performed; if degradation is detected, automatic rollback is triggered and penalty samples are recorded to ensure system stability.
[0156] The RF environment fingerprint vector F(s) is generated as follows: the convolutional neural network part in the state mapping entity contains at least four convolutional layers, each with several convolutional kernels, and uses the interference rate sequence, noise floor sequence, channel utilization sequence, and client SNR sequence as multi-channel inputs for local feature extraction; the fully connected layer flattens the feature map output by the convolutional layer and maps it to a fixed-dimensional intermediate vector; the long short-term memory network models the temporal sequence of the intermediate vector, captures the network busy-idle periodicity, interference dynamic evolution mode, and channel competition temporal characteristics, and outputs the final RF environment fingerprint vector F(s) = [f1, f2, ..., f K ], where K=K CNN +K LSTM K CNN K represents the convolutional feature dimension. LSTM This represents the temporal feature dimension.
[0157] The RF environment fingerprint vector F(s) characterizes the uniqueness of the RF physical environment of the network deployment location. Unlike the device RF fingerprint, F(s) describes the site environment characteristics rather than the differences in terminal device hardware manufacturing. The RF fingerprints of the same site are highly similar when the physical environment is stable, while different sites show statistically significant fingerprint differences due to differences in building structure, interference patterns, and device density. The RF environment fingerprint is archived in a fingerprint database indexed by site identifiers, which supports cross-time period environment drift detection. When a statistically significant shift in the RF fingerprint is detected, a policy re-evaluation is triggered.
[0158] The fields of the decision contract {A,R,C,S} are defined as follows: Proposed configuration action A={(param k ,val k )|k=1,...,M}, where param k For configuration parameter names (including but not limited to 802.11r Fast BSS Transfer Switch, RxSOP Threshold, FRA Switch, WPA3 Transition Mode Switch, Data Rate Base), val k The target parameter value; the inference path R = {evidence} j |j=1,...,J}, where each summary information is evidence. j It includes RF fingerprint matching records that support action security, the improvement in effectiveness of historical sites with the same fingerprint type, and the estimated range of expected effects for the current action; the confidence score C belongs to [0,1] and is calculated by weighting the success rate of the same action on historical sites with the same RF fingerprint type; the candidate scope S is the set of network identifiers to be evaluated.
[0159] The RF fingerprint similarity calculation method is: sim(F(s) i),F(s))=F(s i )·F(s) / (||F(s i The effective scope S is calculated as ||2×||F(s)||2), which means calculating the cosine similarity between the RF fingerprint vectors of the two sites; eff ={s i in S:sim(F(s i The system also sets an upper bound constraint B on the explosion radius, where delta is a configurable similarity threshold. max When |S eff | / N total Greater than B max At that time, from S eff In this process, the top few networks with the highest similarity are selected for configuration execution to ensure that the impact of a single change does not exceed B. max The above s i in S represents the wireless network access point s in the candidate scope S. i .
[0160] Reward signal R t The calculation formula is: R t =f(M t M t-1 )[+]g(E t ), where [+] is the summation or product operator; f(M t M t-1 )=h1(dSNR t )[+]h2(dLat t )[+]h3(dUtil t ), where dSNR t dLat t dUtil t These represent the improvement in client SNR, the improvement in roaming latency, and the rationalization of channel utilization, respectively, with h1, h2, and h3 being positive non-decreasing functions of their corresponding inputs; g(E t )=-[p1(CovHole t )[+]p2(AssocFail t )[+]p3(LegacyStarv t ]], where p1, p2, and p3 are penalty functions, each being a positive non-decreasing function of its corresponding input; Reward Normalization: R t_norm =R t / ||F(s t )||2, normalize the RF fingerprint vector magnitude to eliminate dimensional differences between heterogeneous sites.
[0161] Please see Figure 8 , Figure 8 This is a schematic diagram of the design principle of a confidence-bounded configuration reward function provided in an embodiment of this application. The state change is calculated based on the current state and the previous state to obtain the reward component and the penalty component, and then the total reward is obtained. After the radio frequency fingerprint normalization operation, the RL policy is updated.
[0162] The current state and the previous state include RF fingerprint and experience metric snapshot; The calculation of state changes includes the following four changes: dSNR=SNR t -SNR t-1 ; dLat=Lat t-1 -Lat t ; dInt=Int t-1 -Int t ; dAssoc=Fail t-1 -Fail t ; Int t and Int t-1 dInt represents the channel utilization rate, and Fail represents the change in channel utilization rate. t and Fail t-1 dAssoc represents the association failure rate, and dAssoc represents the change in the association failure rate.
[0163] The reward components include: R1=h1(dSNR) and R2=h2(dLat); The penalty components include: P1=p1(CovHole), P2=p2(AssocFail), P3=p3(LegacyStarv); CovHole represents the number of coverage holes, AssocFail represents the association failure rate, and LegacyStarv represents the starvation level of the legacy device.
[0164] The total reward is the sum of the reward amount and the penalty amount.
[0165] Confidence gating is set to a three-level threshold: when the confidence score C is greater than θ high When θ is 0.85 (default), the system automatically performs configuration changes; when θ low (Default 0.6) Confidence score C less than or equal to θ high When the system generates a recommendation notification containing a decision contract and sends it to the administrator, the administrator confirms and executes it; when the confidence score C is less than θ lowIf this happens, the system will temporarily suspend the recommendation and store the current status in a low-confidence sample pool. This will trigger data collection from more similar RF fingerprint sites to improve the statistical confidence. Once enough samples have been accumulated, the confidence score will be recalculated.
[0166] The context gating steps are as follows: Before distributing the globally optimal RF strategy to a specific site, the RF fingerprint vector F(s) of that site is used as the gating condition to calculate the compatibility score compat(a,F(s)) between the strategy action a and the RF fingerprint F(s). The compatibility score is calculated based on the historical successful application records of action a to similar RF fingerprint sites (sim>=delta). When compat(a,F(s)) is lower than the compatibility gating threshold epsilon, the globally optimal strategy is not applied to this site, thereby avoiding the negative effects of the globally optimal strategy in a local heterogeneous environment.
[0167] Please see Figure 7 , Figure 7 This is a comparison chart of the measured results of this application and related technologies. The vertical axis of the chart represents the index value. The chart shows the comparison results of this application and related technologies in roaming latency (ms), 2.4GHz interference rate (%), 5GHz channel utilization rate (%), and 6GHz SNR (dB).
[0168] This solution also includes short-term statistical verification and automatic rollback mechanisms: after configuration execution, in the preset statistical window T... valid (Default 30 to 60 minutes) Continuously collect key experience metrics; if any core metric shows statistically significant deterioration (based on t-test or CUSUM test, significance level alpha=0.05), automatically trigger a rollback to the previous configuration state; and record this execution as a penalty sample (reward value set to R). min Incorporate it into the reinforcement learning training set to prevent the policy from repeatedly triggering the same erroneous decision on sites with the same RF fingerprint type.
[0169] This solution provides a confidence-bounded configuration recommendation system based on radio frequency (RF) environment fingerprints, comprising: an RF environment fingerprint generation engine, including a state mapping entity comprising a CNN subnetwork, a fully connected layer, and an LSTM temporal layer, used to encode multidimensional RF telemetry data into a site-specific RF environment fingerprint vector F(s); a confidence-bounded decision engine, including a reward generation entity, comprising a confidence evaluation subentity, a decision contract generation subentity, and an explosion radius control subentity, used to generate a structured decision contract {A,R,C,S} and perform RF fingerprint similarity screening; and a reinforcement learning core entity, used to reward or punish based on the experience metric signal R. t Continuous update strategy pi(a|s); validation and rollback of entities, used in the statistical window T validInternally, regression tests are performed, and automatic rollback is triggered when metric degradation is detected.
[0170] The reward-generating entity further includes: a state history entity, used to store the recent RF fingerprint state sequence {F(s)} in a first-in-first-out manner. t-n ),...,F(s t The fingerprint matching entity is used to retrieve the historical state record with the highest similarity to the current RF fingerprint F(s) from the state history, and extract the corresponding configuration action and experience index result pair; the confidence calculation entity is used to calculate the success rate of the target action in the historical matching record and calculate the confidence score C by weighting; the contract output entity is used to encapsulate the action A, reasoning path R, confidence C and candidate scope S into a standardized decision contract format output.
[0171] Compared with the related technology CN113302980B, the differences and technical effects of this embodiment are shown in Table 1: Table 1. Comparison of the differences between this application and related technologies:
[0172] This embodiment employs a reinforcement learning-based wireless network optimization framework, utilizing a (state, action, reward, next state) tuple to drive policy updates and designing a reward function based on channel handover QoE-related metrics. This embodiment upgrades state representation from single AP channel handover history to a site-level multi-dimensional RF environment fingerprint vector, introducing a CNN-LSTM joint neural network as the core of the state mapping entity, significantly improving the ability to identify heterogeneous RF environments. This embodiment expands the configuration action space from single channel selection to the entire set of RF configuration parameters (802.11r, RxSOP, FRA, WPA3, data rate baseline), covering all dimensions of wireless network configuration management. This embodiment introduces a confidence-bounded decision contract mechanism, enabling each AI decision to have a quantifiable reliability assessment and a complete auditable inference path, solving the black-box push problem of AI configuration automation. This embodiment introduces a precise blast radius bounding mechanism based on RF fingerprint similarity, limiting the potential impact of configuration change failures to a subset of networks with highly similar RF environments, significantly reducing the operational risks of large-scale automated deployment.
[0173] This embodiment uses site RF fingerprint as the context driver, and the configuration recommendation is highly matched with the actual RF environment. Compared with the static Golden Template solution, the roaming latency is improved by about 93% (470ms→46ms), the co-channel interference is reduced by about 41%, and the accuracy of the configuration is improved.
[0174] In this embodiment, the bounded explosion radius mechanism ensures that the impact of configuration changes is predictable and controllable, and the compatibility gating effectively prevents the global optimal strategy from producing negative effects in local heterogeneous environments, thereby improving the security of the configuration.
[0175] In this embodiment, the decision contract mechanism ensures that each AI decision has a complete and auditable reasoning path of "what to do - why it is safe - where it will affect - how much it will affect", which meets the requirements of enterprise information technology governance for the transparency of AI decision-making and has interpretability.
[0176] This embodiment combines online reinforcement learning closed-loop with automatic rollback and penalty sample feedback, enabling the system to continuously adapt to changes in the RF environment and learn from errors. It can maintain the long-term optimality of the configuration strategy without manual intervention and has self-evolution capabilities.
[0177] This embodiment presents an AI-based adaptive configuration recommendation method and system based on site-specific RF environment fingerprint confidence constraints and scope boundary control. It is applicable to large-scale Wi-Fi enterprise networks, university campus networks, and densely deployed scenarios that include heterogeneous RF environments.
[0178] After implementing the above solution in the university campus network, the specific execution process is described as follows: In a deployment scenario at a university (19 buildings, 930 APs, approximately 14,600 clients), the network health before optimization was as follows: 2.4GHz health: Poor (interference rate >36%, roaming latency ~430ms); 5GHz health: Poor (channel utilization 48%, roaming latency ~564ms); 6GHz: No clients (no SNR data). "Poor" indicates a poor level of performance.
[0179] Data Acquisition: The system continuously collects interference rate, noise floor, channel utilization, client SNR, and time series heatmap from 930 APs, and updates telemetry snapshots every 5 minutes. The data collector performs data transformation and storage.
[0180] Fingerprint generation: CNN in the state mapping entity extracts the spatial pattern of the RF environment of each building, and LSTM integrates the temporal pattern to generate an independent RF environment fingerprint vector F(s) (dimension 512) for each building. The fingerprint database identifies four main RF environment types: administrative building (low interference, low density), teaching building (high interference, high density), laboratory (medium interference, IoT hybrid type), and stadium (low regular load, peak pulse type).
[0181] Decision contract generation: The RL agent recommends enabling Fast BSS Transition for the teaching building cluster, generating a decision contract: Action A = {(802.11r, enable)}; Inference path R records 20 historical similar RF fingerprint sites (average sim = 0.91), showing an average improvement of approximately 92% in roaming latency after enabling 802.11r; Confidence C = 0.91, higher than θ. high =0.85, enter the automatic execution process; the candidate scope S contains all 16 5GHz networks. The above enable indicates that it is enabled.
[0182] Explosion radius control: The explosion radius control sub-entity calculates the RF fingerprint similarity of each network and identifies three networks with a similarity lower than 0.85 (these three networks contain a large number of older 802.11a devices, and the RF fingerprint features show high device type entropy values). These networks are then excluded, and the effective scope S is determined. eff =13 networks.
[0183] Execution and Validation: After enabling Fast BSS Transition on 13 networks, the system monitored roaming latency sequences over 60 minutes. A t-test confirmed that roaming latency decreased from approximately 470ms to approximately 46ms (an improvement of approximately 93%), with a statistically significant difference (p<0.001). No associated failure events were observed, triggering the reward signal h2(dLat). t =h2(424ms), update strategy.
[0184] Final optimization results: 5GHz health improved to Great (channel utilization 32%, roaming latency ~44ms); 6GHz health improved to Great (SNR 34dB, channel utilization 21%); 2.4GHz interference rate decreased from >36% to 6% (an improvement of approximately 83%). FRA was enabled on 16 networks and disabled on 1 network, resulting in a reduction of approximately 41% in co-channel interference; WPA3 Transition was enabled on 13 networks and required to be disabled on 0 networks (IoT gating effectively prevented WPA3 from being pushed to incompatible devices).
[0185] This embodiment prevents local harm from global policies through context gating, as follows: For WPA3 Transition (i.e., WPA3) Transition The system, based on the RF fingerprint compatibility scoring mechanism, identified the RF fingerprint F(s) of the laboratory area. lab In the data, the device type distribution characteristics (analysis of the associated request frame capability bits in WLAN event data) show that older devices (802.11n and below devices that do not support WPA3) account for more than 40%.
[0186] Compatibility rating compat(WPA3) Transition ,F(s lab The threshold value (epsilon=0.35) is below the compatibility gating threshold (epsilon=0.5), triggering context gating. WPA3 Transition is not pushed to the lab SSID (Service Set Identifier), effectively mitigating the risk of association failure for approximately 600 IoT sensor devices. Final results: WPA3 Transition was enabled on 13 networks, and disabled on 0 networks due to context gating (no rollback required after gating). The 6GHz client SNR improved by approximately 100% (from no client to 34dB strong signal coverage).
[0187] This application provides a network configuration system based on radio frequency environmental fingerprinting, comprising: The fingerprint generation module is used to acquire the radio frequency telemetry data of the current wireless network access point and generate the current radio frequency environment fingerprint vector based on the radio frequency telemetry data. A decision module is used to generate decision information based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity; wherein, the decision information includes multiple configuration description fields, and the configuration description fields include at least a target configuration action, a confidence score, and a candidate scope; the confidence score is the confidence of the target configuration action, and the candidate scope is a set of wireless network access points that have applied the target configuration action; The scope determination module is used to determine the reference radio frequency environment fingerprint vector when each wireless network access point in the candidate scope performs the target configuration action if the confidence score is greater than the first confidence threshold, and select the wireless network access point as the effective scope from the candidate scope based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector. The configuration application module is used to apply the target configuration action within the effective scope.
[0188] This embodiment generates a current radio frequency environment fingerprint vector based on the radio frequency telemetry data of the current wireless network access point. It then combines this fingerprint vector with the current policy of the reinforcement learning core entity to generate decision information. This decision information includes a target configuration action, a confidence score, and candidate scopes. The confidence score represents the confidence level of the target configuration action, and the candidate scopes are the set of wireless network access points where the target configuration action has been applied. When the confidence score is greater than a first confidence threshold, this embodiment selects a wireless network access point from the candidate scopes as an effective scope based on the similarity of the radio frequency environment fingerprint vectors, and then applies the target configuration action within the effective scope. This process uses the radio frequency environment fingerprint as the basis for configuration decisions. By limiting the confidence level and candidate scopes, it achieves precise matching between the configuration action and the actual radio frequency environment, and strictly limits the scope of configuration changes to a consistent and verified range. Therefore, this embodiment can accurately configure wireless network access points based on the radio frequency environment, improving the security and reliability of wireless network configuration.
[0189] Furthermore, the process by which the scope determination module selects a wireless network access point as an effective scope from the candidate scopes based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector includes: Determine the fingerprint similarity between each of the reference radio frequency environment fingerprint vectors and the current radio frequency environment fingerprint vector; Based on the fingerprint similarity, all wireless network access points in the candidate scope are divided into a first type of wireless network access point, a second type of wireless network access point, and a third type of wireless network access point; wherein, the fingerprint similarity corresponding to the first type of wireless network access point is greater than a first fingerprint similarity threshold, the fingerprint similarity corresponding to the second type of wireless network access point is less than or equal to the first fingerprint similarity threshold and greater than a second fingerprint similarity threshold, and the fingerprint similarity corresponding to the third type of wireless network access point is less than or equal to the second fingerprint similarity threshold; Add all first-type wireless network access points to the candidate set; If a manual confirmation message is received, the second type of wireless network access point corresponding to the manual confirmation message is added to the candidate set; Determine whether the number of wireless network access points in the candidate set is greater than the number threshold Y; If so, then the set of the Y wireless network access points with the highest fingerprint similarity in the candidate set is set as the effective scope; If not, then the candidate set is set as the effective scope.
[0190] Furthermore, the process by which the decision module generates decision information based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity includes: The current radio frequency environment fingerprint vector is input into the reinforcement learning core entity so that the reinforcement learning core entity determines the target configuration action based on the current policy. Obtain multiple historical status information entries; wherein each historical status information entry includes the configuration action a executed at time step t. t Application configuration action a t The obtained performance index values and the RF environment fingerprint vector at time step t; Calculate the similarity between the reference radio frequency environment fingerprint vector in the historical state information and the current radio frequency environment fingerprint vector, and add the K historical state information entries with the highest similarity to the historical matching set; Configure action a in the historical matching set t The historical state information of the actions configured for the target is added to the reference set; Using the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector as weights, the confidence score is obtained by weighted statistical analysis of the cases in the reference set where the performance index value meets the preset performance conditions after the target configuration action is performed. Set the set of wireless network access points corresponding to each historical state information in the reference set as the candidate scope; Construct the decision information including the target configuration action, the confidence score, and the candidate scope.
[0191] Furthermore, the decision module is also configured to generate summary information for the historical state information in the historical matching set and add the summary information to the decision information; it is also configured to send the decision information to the management device if the confidence score is less than or equal to the first confidence threshold and greater than the second confidence threshold; it is also configured to determine the effective scope and apply the target configuration action within the effective scope if an authorization instruction for the decision information is received from the management device.
[0192] Furthermore, it also includes: The confidence recalculation module is used to add the current radio frequency environment fingerprint vector to the sample pool if the confidence score is less than the second confidence threshold, and to obtain radio frequency telemetry data and configuration execution results of other wireless network access points with a similarity higher than a preset collection threshold to the current radio frequency environment fingerprint vector, so as to increase historical status information; it is also used to recalculate the confidence score of the target configuration action when the newly added historical status information is greater than a preset accumulation quantity.
[0193] Furthermore, it also includes: The policy update module is used to obtain performance metrics after applying the target configuration action within the effective scope; wherein, the performance metrics include signal-to-noise ratio change, roaming duration change, channel utilization offset, number of coverage holes, association failure rate, and number of abnormal devices; the channel utilization offset is used to describe the deviation between the actual channel utilization and the target channel utilization, and the number of abnormal devices is used to describe the number of abnormal wireless network access points after applying the target configuration action; it is also used to calculate a reward function value based on the performance metrics, and normalize the reward function value using the L2 norm of the current radio frequency environment fingerprint vector; it is also used to update the current policy of the reinforcement learning core entity based on the normalized reward function value.
[0194] Furthermore, it also includes: The policy update module is used to collect performance metrics before and after applying the target configuration action within a preset time window after applying the target configuration action within the effective scope; it is used to determine whether the performance metrics before and after applying the target configuration action meet the anomaly judgment conditions; if so, it performs a rollback operation to restore the configuration to the state before applying the target configuration action and modifies the reward function value corresponding to the target configuration action to a preset penalty value; it is used to set the current radio frequency environment fingerprint vector, the target configuration action, and the modified reward function value as penalty samples, so as to update the current policy of the reinforcement learning core entity based on the penalty samples.
[0195] Since the embodiments of the system part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the system part, and they will not be repeated here.
[0196] This application also provides a storage medium on which a computer program is stored, which, when executed, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0197] This application also provides an electronic device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the steps provided in the above embodiments. Of course, the electronic device may also include various network interfaces, power supplies, and other components.
[0198] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
[0199] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A network configuration method based on radio frequency environmental fingerprinting, characterized in that, include: Acquire the radio frequency telemetry data of the current wireless network access point, and generate the current radio frequency environment fingerprint vector based on the radio frequency telemetry data; Decision information is generated based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity; wherein, the decision information includes multiple configuration description fields, and the configuration description fields include at least the target configuration action, confidence score and candidate scope; the confidence score is the confidence of the target configuration action, and the candidate scope is the set of wireless network access points that have applied the target configuration action; If the confidence score is greater than the first confidence threshold, then the reference radio frequency environment fingerprint vector of each wireless network access point in the candidate scope when performing the target configuration action is determined, and wireless network access points are selected from the candidate scope as effective scopes based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector. Apply the target configuration action within the effective scope.
2. The network configuration method based on radio frequency environment fingerprinting according to claim 1, characterized in that, Selecting wireless network access points as effective scopes from the candidate scopes based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector includes: Determine the fingerprint similarity between each of the reference radio frequency environment fingerprint vectors and the current radio frequency environment fingerprint vector; Based on the fingerprint similarity, all wireless network access points in the candidate scope are divided into a first type of wireless network access point, a second type of wireless network access point, and a third type of wireless network access point; wherein, the fingerprint similarity corresponding to the first type of wireless network access point is greater than a first fingerprint similarity threshold, the fingerprint similarity corresponding to the second type of wireless network access point is less than or equal to the first fingerprint similarity threshold and greater than a second fingerprint similarity threshold, and the fingerprint similarity corresponding to the third type of wireless network access point is less than or equal to the second fingerprint similarity threshold; Add all first-type wireless network access points to the candidate set; If a manual confirmation message is received, the second type of wireless network access point corresponding to the manual confirmation message is added to the candidate set; Determine whether the number of wireless network access points in the candidate set is greater than the number threshold Y; If so, then the set of the Y wireless network access points with the highest fingerprint similarity in the candidate set is set as the effective scope; If not, then the candidate set is set as the effective scope.
3. The network configuration method based on radio frequency environment fingerprinting according to claim 1, characterized in that, Decision information is generated based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity, including: The current radio frequency environment fingerprint vector is input into the reinforcement learning core entity so that the reinforcement learning core entity determines the target configuration action based on the current policy. Obtain multiple historical status information entries; wherein each historical status information entry includes the configuration action a executed at time step t. t Application configuration action a t The obtained performance index values and the RF environment fingerprint vector at time step t; Calculate the similarity between the reference radio frequency environment fingerprint vector in the historical state information and the current radio frequency environment fingerprint vector, and add the K historical state information entries with the highest similarity to the historical matching set; Configure action a in the historical matching set t The historical state information of the actions configured for the target is added to the reference set; Using the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector as weights, the confidence score is obtained by weighted statistical analysis of the cases in the reference set where the performance index value meets the preset performance conditions after the target configuration action is performed. Set the set of wireless network access points corresponding to each historical state information in the reference set as the candidate scope; Construct the decision information including the target configuration action, the confidence score, and the candidate scope.
4. The network configuration method based on radio frequency environment fingerprinting according to claim 3, characterized in that, Also includes: Generate summary information for the historical state information in the historical matching set, and add the summary information to the decision information; If the confidence score is less than or equal to the first confidence threshold and greater than the second confidence threshold, then the decision information is sent to the management device. If an authorization instruction for the decision information is received from the management device, the effective scope is determined, and the target configuration action is applied within the effective scope.
5. The network configuration method based on radio frequency environment fingerprinting according to claim 4, characterized in that, Also includes: If the confidence score is less than the second confidence threshold, the current radio frequency environment fingerprint vector is added to the sample pool, and radio frequency telemetry data and configuration execution results of other wireless network access points with a similarity higher than the preset collection threshold to the current radio frequency environment fingerprint vector are obtained in order to increase historical status information; When the newly added historical state information exceeds the preset accumulated amount, the confidence score of the target configuration action is recalculated.
6. The network configuration method based on radio frequency environment fingerprinting according to claim 1, characterized in that, After applying the target configuration action within the effective scope, the method further includes: Obtain performance metrics after applying the target configuration action; wherein, the performance metrics include signal-to-noise ratio change, roaming duration change, channel utilization offset, number of coverage holes, association failure rate, and number of abnormal devices; the channel utilization offset is used to describe the degree of deviation between the actual channel utilization and the target channel utilization, and the number of abnormal devices is used to describe the number of wireless network access points that are abnormal after applying the target configuration action; The reward function value is calculated based on the performance index, and the reward function value is normalized using the L2 norm of the current radio frequency environment fingerprint vector; The current policy of the core entity in the reinforcement learning is updated based on the normalized reward function value.
7. The network configuration method based on radio frequency environment fingerprinting according to claim 6, characterized in that, After applying the target configuration action within the effective scope, the method further includes: Within a preset time window, collect performance metrics before and after applying the target configuration action; Determine whether the performance indicators before and after applying the target configuration action meet the anomaly judgment conditions; If so, a rollback operation is performed to restore the configuration to its state before the target configuration action was applied, and the reward function value corresponding to the target configuration action is modified to a preset penalty value; The current radio frequency environment fingerprint vector, the target configuration action, and the modified reward function value are set as penalty samples so that the current policy of the reinforcement learning core entity is updated based on the penalty samples.
8. A network configuration system based on radio frequency environmental fingerprinting, characterized in that, include: The fingerprint generation module is used to acquire the radio frequency telemetry data of the current wireless network access point and generate the current radio frequency environment fingerprint vector based on the radio frequency telemetry data. A decision module is used to generate decision information based on the current radio frequency environment fingerprint vector and the current policy of the reinforcement learning core entity; wherein, the decision information includes multiple configuration description fields, and the configuration description fields include at least a target configuration action, a confidence score, and a candidate scope; the confidence score is the confidence of the target configuration action, and the candidate scope is a set of wireless network access points that have applied the target configuration action; The scope determination module is used to determine the reference radio frequency environment fingerprint vector when each wireless network access point in the candidate scope performs the target configuration action if the confidence score is greater than the first confidence threshold, and select the wireless network access point as the effective scope from the candidate scope based on the similarity between the reference radio frequency environment fingerprint vector and the current radio frequency environment fingerprint vector. The configuration application module is used to apply the target configuration action within the effective scope.
9. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor, when calling the computer program in the memory, implements the steps of the network configuration method based on radio frequency environmental fingerprinting as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores computer-executable instructions, which, when loaded and executed by a processor, implement the steps of the network configuration method based on radio frequency environmental fingerprinting as described in any one of claims 1 to 7.
Citation Information
Patent Citations
A channel allocation method, system, electronic device, and storage medium
CN113302980B