Intelligent Optimization Method and System for Multi-Gateway Communication Links

Through the optimization of gateway status monitoring through a hierarchical dynamic detection system and artificial intelligence model, combined with performance prediction and hashing algorithm to optimize load allocation, the detection rigidity and load imbalance in gateway communication optimization in the existing technology is solved, and the performance and stability of the system are improved.

CN119743389BActive Publication Date: 2025-07-18HANGZHOU TUYA INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510239632.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-18
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

In the prior art, gateway communication optimization has problems such as rigid detection mechanism, oversimplified gateway selection strategy, unbalanced load allocation and insufficient reliability of equipment migration, especially in complex and changeable practical application environments, resulting in poor system performance.

Method used

The gateway status monitoring is carried out using a hierarchical dynamic detection system and an artificial intelligence model, gateway evaluation is carried out through a performance prediction model and a comprehensive rating mechanism, load balancing optimization is carried out in combination with an improved consistent hashing algorithm and local sensitive hashing mechanism, and device migration strategies are formulated to ensure system stability.

Benefits of technology

It improves the accuracy and efficiency of fault detection, realizes accurate evaluation of gateway performance, optimizes load allocation, improves the overall performance and stability of the system, and ensures business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743389B_ABST
    Figure CN119743389B_ABST
Patent Text Reader

Abstract

This application provides a method and system for intelligent optimization of multi-gateway communication links. The method obtains real-time status data of gateways by constructing a hierarchical dynamic detection system, and uses an artificial intelligence model to optimize dynamic scheduling strategies; constructs a performance prediction model based on in-depth detection results for gateway rating; realizes load balancing optimization by adopting an improved consistent hashing algorithm and a locality-sensitive hashing mechanism; and finally formulates a migration strategy and performs device migration verification. This application realizes the intelligent optimization of gateway communication links and improves the overall performance and stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things communication technologies, and particularly to a method and system for intelligent optimization of multi-gateway communication links. Background Art

[0002] The rapid development of Internet of Things technology has promoted the wide application of various intelligent devices. Among them, as the core node connecting terminal devices and cloud applications, the performance and stability of the gateway directly affect the operation effect of the entire Internet of Things system.

[0003] Currently, common gateway communication optimization technologies mainly include gateway status detection based on fixed time intervals and simple load balancing mechanisms. For example, some systems detect the gateway status at fixed time intervals and allocate devices based on a simple principle of proximity; while other systems use polling for load balancing.

[0004] More advanced solutions introduce a gateway self-check mechanism and device migration function, by regularly detecting gateway performance metrics and reallocating devices when necessary. This solution maintains system stability by triggering device migration when a decrease in communication quality is detected through real-time monitoring of the gateway.

[0005] However, the existing technologies have problems such as rigid detection mechanisms, overly simple gateway selection strategies, unbalanced load distribution, and insufficient reliability of device migration. Especially in complex and changeable actual application environments, the fixed-frequency detection method can neither detect faults in a timely manner nor cause waste of computing resources; while the simple gateway selection strategy cannot fully consider gateway performance, environmental factors, and long-term stability, resulting in low overall system performance. Summary of the Invention

[0006] In view of this, the embodiments of this application provide a method and system for intelligent optimization of multi-gateway communication links, which solve the problems of rigid detection mechanisms, overly simple gateway selection strategies, unbalanced load distribution, and insufficient reliability of device migration in the existing technologies.

[0007] The embodiments of this application provide a method for intelligent optimization of multi-gateway communication links, including:

[0008] Obtaining real-time gateway status data, where the real-time gateway status data includes gateway load, signal strength, and response time, constructing a hierarchical dynamic detection system, and optimizing dynamic scheduling strategies through an artificial intelligence model to obtain a deep detection result set;

[0009] Based on the deep detection result set, constructing a performance prediction model, performing comprehensive gateway rating calculations, and obtaining a gateway reliability score table;

[0010] Based on the gateway reliability scoring table, initialize the hash ring structure through an improved consistent hashing algorithm, construct a virtual node optimization allocation scheme, and optimize the allocation strategy of devices with physically close locations through a locality-sensitive hashing mechanism to obtain an optimal device allocation scheme;

[0011] Based on the optimal device allocation scheme, formulate a device migration strategy, execute the device migration and conduct verification, and output a migration verification report.

[0012] In some of the embodiments, construct the hierarchical dynamic detection system, optimize the dynamic scheduling strategy through an artificial intelligence model, and obtain a deep detection result set, including:

[0013] Construct a two-layer detection framework including a basic detection layer and a deep detection layer, and generate a detection framework configuration parameter set based on the real-time status data of the gateway;

[0014] Based on the detection framework configuration parameter set and historical detection data, construct a reinforcement learning environment, define a gateway state space and an action space of detection behaviors, and use a reward function of detection effect and resource consumption and the action space to obtain a reinforcement learning environment parameter set;

[0015] Based on the reinforcement learning environment parameter set, train a detection strategy model through a deep Q-learning algorithm to generate the deep detection result set.

[0016] In some of the embodiments, before using the reward function of detection effect and resource consumption and the action space to obtain the reinforcement learning environment parameter set, the method further includes:

[0017] Set the reward value and penalty value of the reward function, where a first preset reward value is given when a fault is detected within a preset time, and a second preset reward value is given when the detection resource utilization rate is maintained within a preset range;

[0018] A first preset penalty value is given when there is waste of detection resources, and a second preset penalty value is given when a fault fails to be detected in time.

[0019] In some of the embodiments, training the detection strategy model through the deep Q-learning algorithm to generate the deep detection result set includes:

[0020] Construct an inductive logic learning environment and construct an initial logic rule base including basic predicates and background knowledge;

[0021] Construct a fusion learning framework of inductive logic programming and reinforcement learning, and configure an Actor-Critic network structure and a logic rule learner;

[0022] Introduce an answer set reasoning mechanism to guide the exploration process of the reinforcement learning, and generate an action candidate set based on logical rules for reasoning.

[0023] Based on the action candidate set, discover the first decision-making pattern through meta-rule learning, including: collecting decision-making experience through the interaction between the Actor-Critic network and the environment during the experience accumulation stage; inducing the second decision rule through inductive logic programming technology during the rule induction stage based on the decision-making experience; evaluating the second decision rule during the knowledge integration stage based on the second decision rule.

[0024] Integrate the learned detection strategy, the logical rule base, and the verification data to generate the depth detection result set.

[0025] In some embodiments, build a performance prediction model, perform gateway comprehensive rating calculation, and obtain a gateway reliability score table, including:

[0026] Based on the depth detection result set, build an evaluation system by analyzing hardware metrics, communication quality, and environmental parameters to obtain a gateway performance prediction model parameter set.

[0027] Based on the gateway performance prediction model parameter set, historical stability, and environmental adaptability factors, perform dynamic evaluation of performance indicators to obtain a gateway rating result set.

[0028] Based on the gateway rating result set, formulate a gateway selection strategy through a multi-objective optimization algorithm to generate the gateway reliability score table.

[0029] In some embodiments, performing the dynamic evaluation of the performance indicators includes:

[0030] Evaluate the fluctuation degree of the gateway performance by calculating the coefficient of variation of key indicators, count the number of failures and the average mean time to recovery in the past 30 days, and evaluate the performance fluctuation of the gateway under different load conditions to obtain a historical stability score.

[0031] Evaluate the stability of the gateway in the current temperature environment, the communication quality under electromagnetic interference conditions, and the location adaptability in the current network topology structure, and obtain an environmental adaptability score through the fuzzy comprehensive evaluation method.

[0032] In some embodiments, initialize the hash ring structure through an improved consistent hashing algorithm, construct a virtual node optimization allocation scheme, and optimize the allocation strategy of devices with physically close positions through a locality-sensitive hashing mechanism to obtain an optimal device allocation scheme, including:

[0033] Initialize the hash ring structure based on the reliability score table of the gateway, and map the gateway nodes to the hash space through an improved hash algorithm to obtain an initial hash ring mapping relationship set;

[0034] Based on the initial hash ring mapping relationship set, dynamically adjust the number of virtual nodes of each gateway according to the reliability score of the gateway to obtain a virtual node allocation scheme;

[0035] Based on the virtual node allocation scheme, optimize the physical location information of the device through the locality-sensitive hashing mechanism to generate the optimal device allocation scheme.

[0036] In some embodiments, formulating the device migration strategy includes:

[0037] Based on the optimal device allocation scheme, analyze the business importance, service level, and migration risk of the device to formulate a device migration strategy plan;

[0038] Among them, based on the optimal device allocation scheme, formulate a device migration strategy, execute the device migration and perform verification, and output a migration verification report, including:

[0039] Based on the device migration strategy plan, execute the device migration through a fault detection and rollback mechanism to obtain a migration execution result set;

[0040] Based on the migration execution result set, perform verification through performance testing, stability evaluation, and business continuity check to generate the migration verification report.

[0041] In some embodiments, constructing the hierarchical dynamic detection system and constructing the performance prediction model based on the deep detection result set includes:

[0042] Deploy local detection agents at gateway nodes and global coordination agents at core nodes to form a multi-agent network topology;

[0043] Based on the multi-agent network topology, maintain the stability of the multi-agent network through a lease-based heartbeat mechanism to construct an agent collaboration framework;

[0044] Based on the agent collaboration framework, implement local decision-making through the Actor-Critic architecture and optimize the global detection strategy through the meta-learning method to obtain the deep detection result set;

[0045] Receive the deep detection result set, construct a local evaluation model for each gateway agent, and integrate the evaluation results of each local detection agent through a swarm intelligence algorithm to form a gateway group evaluation model;

[0046] Based on the gateway group evaluation model, calculate the marginal contribution of each gateway to the overall service quality through the coalition game model, dynamically adjust the load distribution weight, and generate a gateway reliability score table.

[0047] An embodiment of the present application further provides a multi-gateway communication link intelligent optimization system, which includes a hierarchical dynamic detection module, a gateway evaluation module, a load balancing optimization module, and a migration verification module, where:

[0048] The hierarchical dynamic detection module is used to obtain gateway real-time status data, where the gateway real-time status data includes gateway load, signal strength, and response time, construct a hierarchical dynamic detection system, optimize the dynamic scheduling strategy through an artificial intelligence model, and obtain a deep detection result set;

[0049] The gateway evaluation module is used to construct a performance prediction model based on the deep detection result set, perform comprehensive gateway rating calculation, and obtain a gateway reliability score table;

[0050] The load balancing optimization module is used to initialize the hash ring structure through an improved consistent hashing algorithm based on the gateway reliability score table, construct a virtual node optimization allocation scheme, and optimize the allocation strategy of devices with physically close positions through the locality-sensitive hashing mechanism to obtain an optimal device allocation scheme;

[0051] The migration verification module is used to formulate a device migration strategy based on the optimal device allocation scheme, perform device migration and verification, and output a migration verification report.

[0052] An embodiment of the present application further provides a computer device, which includes:

[0053] At least one processor; and,

[0054] A memory communicatively connected to the at least one processor; where,

[0055] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the above multi-gateway communication link intelligent optimization method.

[0056] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method of the above multi-gateway communication link intelligent optimization method.

[0057] An embodiment of the present application further provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method of the above multi-gateway communication link intelligent optimization method are implemented.

[0058] The embodiments of the present application have the following technical effects:

[0059] 1. Through the hierarchical dynamic detection system and the artificial intelligence model, the intelligent monitoring of the gateway status is realized, and the accuracy and efficiency of fault detection are improved;

[0060] 2. Based on the performance prediction model and the comprehensive rating mechanism, the accurate evaluation of the gateway performance is realized, providing a reliable decision-making basis for load balancing;

[0061] 3. By adopting the improved consistent hashing algorithm and the locality-sensitive hashing mechanism, the balanced distribution and nearby access of the load are realized, improving the overall performance of the system;

[0062] 4. Through the reliable device migration mechanism, the continuity of the service and the stability of the system are ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the described drawings are only a part of the embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative efforts.

[0064] Figure 1 It is a schematic flowchart of the intelligent optimization method for multi-gateway communication links provided by the embodiments of the present application;

[0065] Figure 2 It is a schematic diagram of the method for hierarchical dynamic detection provided by the embodiments of the present application;

[0066] Figure 3 It is a schematic diagram of the method for the gateway evaluation model provided by the embodiments of the present application;

[0067] Figure 4 It is a schematic diagram of the load balancing optimization method provided by the embodiments of the present application;

[0068] Figure 5 It is a structural block diagram of the intelligent optimization system for multi-gateway communication links provided by the embodiments of the present application;

[0069] Figure 6 It is a schematic diagram of the structure of the computer device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] To make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not used to limit this application. In addition, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0071] As Figure 1 shown, an embodiment of this application provides a method for intelligent optimization of multi-gateway communication links, including the following steps:

[0072] S1: Obtain real-time gateway status data, where the real-time gateway status data includes gateway load, signal strength, and response time, construct a hierarchical dynamic detection system, and optimize the dynamic scheduling strategy through an artificial intelligence model to obtain a deep detection result set;

[0073] First, the system obtains real-time gateway status data from the actual operating environment. This data includes multiple key performance indicators: gateway load (including CPU usage rate, memory occupancy rate, network bandwidth utilization rate, etc.), signal strength (signal strength value in dBm), and response time (delay time for the gateway to process requests). This data is regularly collected and reported through the monitoring module built into the gateway.

[0074] Based on the obtained real-time status data, the system constructs a hierarchical dynamic detection system. This system adopts a two-layer architecture design, including a basic detection layer and a deep detection layer. The basic detection layer conducts routine monitoring with a basic cycle of 5 minutes, mainly focusing on the basic operating status of the gateway; the deep detection layer is activated when specific conditions are triggered to perform more detailed performance analysis. This hierarchical design ensures the continuity of basic monitoring and avoids resource waste caused by over-detection.

[0075] To achieve intelligent scheduling of detection resources, the system introduces an artificial intelligence model for dynamic optimization. Specifically, the system constructs a reinforcement learning environment based on Deep Q-Learning. In this environment, the state space includes real-time status information of the gateway, such as key indicators like CPU usage rate (0 - 100%), memory occupancy rate (0 - 100%), network bandwidth utilization rate (0 - 100%), the current number of connected devices, average response time (ms), signal strength (dBm), etc.; the action space includes operations such as adjusting the basic detection frequency (which can be adjusted within the range of 3 - 10 minutes), triggering deep detection (yes / no), and adjusting the detection index range (expanding / narrowing).

[0076] The system designs a refined reward mechanism to guide the optimization of the detection strategy. When the system successfully discovers a fault within the preset time, it gives a positive reward of +10 points; when the detection resource utilization rate is maintained within a reasonable range, it gives a basic positive reward of +1 point; when there is waste of detection resources, it gives a negative reward of -5 points; when a fault is not discovered in time, it gives a severe negative reward of -20 points. This reward mechanism effectively guides the reinforcement learning agent to learn the optimal detection strategy.

[0077] During the training process, the system uses a deep neural network with a four-layer structure as the Q network: the input layer dimension is consistent with the state space, two hidden layers (each with 128 neurons), and the output layer dimension corresponds to the action space. At the same time, the system implements an experience replay pool with a capacity of 10,000 to store the transition samples (state, action, reward, next state) during the training process, and performs batch learning by randomly sampling 32 samples, improving the stability of the training.

[0078] To further enhance the training effect, the system adopts the target network technology and synchronizes the network parameters every 100 training steps. In terms of the exploration strategy, the system adopts the ε-greedy strategy, with the initial ε value of 0.9, which gradually decreases to 0.1 as the training progresses, achieving a balance between exploration and exploitation. When the fluctuation of the average reward value is less than 5% within 100 consecutive episodes, it is considered that the model training converges.

[0079] Finally, the trained model can autonomously decide the detection strategy according to the real-time state of the gateway and output a deep detection result set containing the detailed performance indicators of each gateway. This result set not only contains performance data but also includes the optimal detection strategies for different scenarios, providing comprehensive data support for the subsequent gateway evaluation.

[0080] This dynamic detection method based on artificial intelligence can better balance the detection effect and resource consumption compared with the traditional fixed-cycle detection method, improving the accuracy and efficiency of fault detection while reducing the waste of system resources. Especially in the scenario where the gateway load changes frequently, this method can adjust the detection strategy in time to ensure the stable operation of the system.

[0081] Specifically, as Figure 2 shown, S1 specifically includes:

[0082] S1.1: Construct a two-layer detection framework including a basic detection layer and a deep detection layer, and generate the detection framework configuration parameter set based on the gateway real-time state data;

[0083] In S1.1, the system constructs a two - layer detection framework including a basic detection layer and a deep detection layer. The basic detection layer performs regular monitoring with a basic cycle of 5 minutes, mainly responsible for collecting and analyzing the basic operation metrics of the gateway, including real - time status data such as CPU load, memory usage rate, network bandwidth utilization rate, etc. The basic detection layer adopts a lightweight detection mechanism, regularly collecting data through the built - in monitoring module to ensure the minimum impact on the normal operation of the gateway.

[0084] The deep detection layer is triggered under specific conditions, such as when the basic detection layer discovers abnormal metrics, or when the system predicts that there may be a performance bottleneck. The deep detection layer will perform a more comprehensive performance analysis, including detailed communication quality assessment, network topology analysis, device connection status inspection, etc. This hierarchical design not only ensures the continuity of basic monitoring but also avoids resource waste caused by over - detection. Based on the collected real - time status data of the gateway, the system generates a set of detection framework configuration parameters, which contains configuration information such as detection period, trigger conditions, detection index range, etc.

[0085] S1.2: Construct a reinforcement learning environment based on the detection framework configuration parameter set and historical detection data, define the gateway state space and the action space of detection behaviors, and use the reward function of detection effect and resource consumption and the action space to obtain the reinforcement learning environment parameter set;

[0086] In S1.2, the system constructs a reinforcement learning environment based on the detection framework configuration parameter set and historical detection data. First, the system defines a complete gateway state space, which contains multiple key dimensions: CPU usage rate (0 - 100%), memory occupancy rate (0 - 100%), network bandwidth utilization rate (0 - 100%), the number of currently connected devices, average response time (ms), signal strength (dBm), etc. These state variables jointly describe the working state of the gateway at a certain moment.

[0087] The system also designs the action space of detection behaviors, including: adjusting the basic detection frequency (which can be adjusted within the range of 3 - 10 minutes), triggering deep detection (yes / no), adjusting the detection index range (expanding / narrowing), etc. These actions directly affect the execution effect of the detection strategy. To evaluate the effects of these actions, the system implements a reward function based on detection effect and resource consumption. A significant positive reward (+10 points) is given when a fault is successfully detected within the preset time, and a basic positive reward (+1 point) is given when the detection resource utilization rate is kept within a reasonable range. On the contrary, a negative reward (-5 points) is given when detection resources are wasted, and a severe negative reward (-20 points) is given when a fault fails to be detected in time.

[0088] The following is a detailed description of steps S1.2.1 and S1.2.2:

[0089] In the process of constructing the reinforcement learning environment, the system first designed a refined reward mechanism to guide the learning and optimization of the detection strategy by setting specific reward values and penalty values. This reward mechanism mainly considers two aspects: the effectiveness of fault detection and the efficiency of resource utilization.

[0090] In S1.2.1, the system set two types of positive reward values. The first preset reward value (+10 points) is used to reward the situation of successfully detecting a fault within the preset time. The "preset time" here is dynamically set according to the severity of different types of faults. For example, for faults that seriously affect the business, the preset time may be set to 5 minutes; for minor performance degradation, the preset time may be extended to 30 minutes. The second preset reward value (+1 point) is used to reward the situation where the detection resource utilization rate remains within the preset range. This preset range is usually set as the CPU usage rate not exceeding 20% and the memory occupancy not exceeding 30%, ensuring that the detection activity will not significantly affect the normal operation of the gateway.

[0091] Specifically, when the system successfully detects a serious fault that the gateway CPU usage rate suddenly soars above 90% within 5 minutes, it will obtain the first preset reward value of +10 points; when the system's resource occupancy always remains within a reasonable range (such as the CPU usage rate maintaining at about 15%) during the execution of the detection task, it will obtain the second preset reward value of +1 point. This differential reward design not only emphasizes the importance of fault detection but also encourages the system to maintain high efficiency in resource utilization.

[0092] In S1.2.2, the system set two types of penalty values to constrain the detection behavior. The first preset penalty value (-5 points) is used to penalize the situation of waste of detection resources. For example, when the system still frequently performs in-depth detection when the gateway load is low and stable, or maintains a high detection frequency when it is unnecessary, this penalty value will be triggered. The second preset penalty value (-20 points) is used to penalize the situation of failing to detect a fault in time, which is the most severe penalty. When the system fails to detect an important fault of the gateway within the preset time, resulting in business interruption or serious performance degradation, this penalty value will be triggered.

[0093] For example, if the system performs an in-depth detection every 1 minute unnecessarily during the normal operation of the gateway, resulting in the detection resource utilization rate exceeding 50%, it will receive the first preset penalty value of -5 points; if the system fails to detect an abnormal important service process of the gateway within 10 minutes, resulting in the devices connected unable to communicate normally, it will receive the second preset penalty value of -20 points. The design of this penalty mechanism particularly emphasizes the penalty for failing to detect a fault in time because this situation may cause more serious business impacts.

[0094] Through this mechanism design with clear rewards and punishments, the system can effectively guide the reinforcement learning agent to form a reasonable detection strategy, ensuring the detection effect while avoiding resource waste. In particular, by imposing the severest punishment on the failure to detect a fault in a timely manner, it is ensured that the system will prioritize the timeliness and accuracy of detection, which is of great significance for maintaining the stable operation of the gateway. At the same time, the punishment for resource waste also ensures that the system will adjust the detection strategy in a timely manner to avoid the performance overhead caused by over-detection.

[0095] In addition, during the process of constructing the reinforcement learning environment, the system first designs a fine-grained reward mechanism to guide the learning and optimization of the detection strategy by setting specific reward values and punishment values. This reward mechanism mainly considers two aspects: the effectiveness of fault detection and the efficiency of resource use.

[0096] Specifically, before executing S1.2, it also includes:

[0097] S1.2.1: Set the reward value and punishment value of the reward function. Among them, when a fault is detected within the preset time, a first preset reward value is given, and when the detection resource utilization rate is maintained within the preset range, a second preset reward value is given;

[0098] In S1.2.1, the system sets two types of positive reward values. The first preset reward value (+10 points) is used to reward the situation where a fault is successfully detected within the preset time. Here, the "preset time" is dynamically set according to the severity of different types of faults. For example, for faults that seriously affect the business, the preset time may be set to 5 minutes; for minor performance degradation, the preset time may be extended to 30 minutes. The second preset reward value (+1 point) is used to reward the situation where the detection resource utilization rate is maintained within the preset range. This preset range is usually set as the CPU usage rate not exceeding 20% and the memory occupancy not exceeding 30%, ensuring that the detection activity will not significantly affect the normal operation of the gateway.

[0099] Specifically, when the system successfully detects a serious fault that the CPU usage rate of the gateway suddenly soars to more than 90% within 5 minutes, it will obtain the first preset reward value of +10 points; when the resource occupancy of the system always remains within a reasonable range (such as the CPU usage rate is maintained at about 15%) during the execution of the detection task, it will obtain the second preset reward value of +1 point. This differential reward design not only emphasizes the importance of fault detection but also encourages the system to maintain high efficiency in resource use.

[0100] S1.2.2: Give the first preset punishment value when detecting resource waste, and give the second preset punishment value when failing to detect a fault in a timely manner.

[0101] In S1.2.2, the system sets up two types of penalty values to constrain the detection behavior. The first preset penalty value (-5 points) is used to punish the situation of waste of detection resources. For example, when the system frequently performs in-depth detection under the condition of low and stable gateway load, or maintains a high detection frequency when it is unnecessary, this penalty value will be triggered. The second preset penalty value (-20 points) is used to punish the situation of failure to detect faults in time, which is the most severe penalty. When the system fails to detect important faults of the gateway within the preset time, resulting in service interruption or serious performance degradation, this penalty value will be triggered.

[0102] For example, if the system unnecessarily performs an in-depth detection every 1 minute during the normal operation of the gateway, resulting in the detection resource utilization rate exceeding 50%, it will receive the first preset penalty value of -5 points; if the system fails to detect the abnormal situation of the important service process of the gateway within 10 minutes, resulting in the devices connected unable to communicate normally, it will receive the second preset penalty value of -20 points. The design of this penalty mechanism particularly emphasizes the penalty for failure to detect faults in time, because this situation may cause more serious business impacts.

[0103] Through this mechanism design with clear rewards and punishments, the system can effectively guide the reinforcement learning agent to form a reasonable detection strategy, avoiding resource waste while ensuring the detection effect. In particular, by imposing the most severe penalty on failure to detect faults in time, it ensures that the system will prioritize ensuring the timeliness and accuracy of detection, which is of great significance for maintaining the stable operation of the gateway. At the same time, the penalty for resource waste also ensures that the system will adjust the detection strategy in a timely manner, avoiding the performance overhead caused by over-detection.

[0104] S1.3: Based on the set of reinforcement learning environment parameters, a detection strategy model is trained through the deep Q-learning algorithm to generate the set of in-depth detection results;

[0105] In S1.3, the system is trained through the deep Q-learning algorithm based on the constructed set of reinforcement learning environment parameters. The system uses a four-layer deep neural network as the Q-network, including: an input layer corresponding to the dimension of the state space, two hidden layers each containing 128 neurons, and an output layer corresponding to the dimension of the action space. To improve the stability of training, the system implements an experience replay pool with a capacity of 10,000, which is used to store the transition samples (state, action, reward, next state) during the training process, and batch learning is performed by randomly sampling 32 samples.

[0106] The system also adopts the target network technology to further improve the training stability, and parameter synchronization is performed every 100 training steps. In terms of the exploration strategy, the system uses the ε-greedy strategy, with the initial ε value set to 0.9, which gradually decreases to 0.1 as the training progresses, achieving a balance between exploration and exploitation. When the fluctuation of the average reward value of the model within 100 consecutive episodes is less than 5%, the training is considered to have converged.

[0107] Finally, the trained model can autonomously decide the detection strategy according to the real-time state of the gateway. The system integrates information such as the trained detection strategy model and verification data into the deep detection result set. This result set not only contains the detailed performance indicators of each gateway, but also includes the optimal detection strategies for different scenarios, providing comprehensive data support for subsequent gateway evaluation. This detection method based on deep reinforcement learning can dynamically adjust the detection strategy according to the actual operating state of the gateway, improving the detection efficiency while reducing the consumption of system resources.

[0108] In addition, in one specific embodiment, S1.3 may further include:

[0109] S1.3.1: Construct an inductive logic learning environment and build an initial logic rule base containing basic predicates and background knowledge;

[0110] In S1.3.1, the system first constructs an inductive logic learning environment. This environment contains a set of basic predicates describing the gateway state and behavior, such as hasHighLoad / 1 indicating high gateway load, needsInspection / 1 indicating the need for in-depth detection, etc. At the same time, the system also establishes a background knowledge base containing information such as gateway topology relationships and historical performance laws. Based on these predicates and background knowledge, the system constructs an initial logic rule base for each local detection agent. These rule bases provide a basic knowledge framework for the subsequent learning process.

[0111] S1.3.2: Construct a fusion learning framework for inductive logic programming and reinforcement learning, and configure the Actor-Critic network structure and the logic rule learner;

[0112] In S1.3.2, the system constructs an innovative fusion learning framework for inductive logic programming and reinforcement learning. In this framework, each detection agent is equipped with both an Actor-Critic network structure and a logic rule learner. The Actor-Critic network is responsible for generating specific detection actions, while the logic rule learner is responsible for inducing high-level decision rules from successful interaction experiences. This dual learning mechanism can simultaneously obtain low-level action strategies and high-level decision rules.

[0113] S1.3.3: Introduce an answer set reasoning mechanism to guide the exploration process of the reinforcement learning, and generate the action candidate set through reasoning based on logical rules;

[0114] In S1.3.3, the system introduces an Answer Set Programming (ASP) mechanism to guide the exploration process of reinforcement learning. When the agent needs to make a detection decision, it first uses the learned logical rules to reason and generate a set of possible action candidates. For example, the system may infer through rule reasoning, and the reasoning code is as follows:

[0115] prolog

[0116] needsInspection(X) :hasHighLoad(X), lowResponseTime(X), notrecentlyChecked(X).

[0117] adjustFrequency(X, high) :hasHighLoad(X), stablePerformance(X).

[0118] These candidate actions are used as the preferred options for reinforcement learning exploration, significantly reducing the probability of ineffective exploration.

[0119] S1.3.4: Based on the action candidate set, discover the first decision pattern through meta-rule learning, including: collecting decision-making experience through the interaction between the Actor-Critic network and the environment during the experience accumulation stage; based on the decision-making experience, inducing the second decision rule through the inductive logic programming technology during the rule induction stage; evaluating the second decision rule during the knowledge integration stage;

[0120] In S1.3.4, the system discovers decision patterns through meta-rule learning based on the generated action candidate set. This process is divided into three stages: First, during the experience accumulation stage, the Actor-Critic network interacts with the environment to collect successful decision-making experience; then, during the rule induction stage, the inductive logic programming technology is used to induce new decision rules from these experiences; finally, during the knowledge integration stage, the effectiveness and generality of these new rules are evaluated.

[0121] Specifically, during the rule induction stage, the system can extract more abstract decision rules from specific interaction experiences. For example, the system may induce from multiple successful detection experiences that when the gateway load continuously exceeds 90% and the response time increases by more than 50%, a deep detection should be triggered immediately. These rules not only include specific triggering conditions but also include corresponding time windows and threshold settings.

[0122] S1.3.5: Integrate the learned detection strategy, the logic rule base, and the verification data to generate the deep detection result set.

[0123] Integrate the learned detection strategy, the logic rule base, and the relevant verification data to generate a complete deep detection result set. This result set is a multi-level knowledge base that contains a complete strategy system from low-level detection actions to high-level decision rules. In this way, the system can not only make accurate detection decisions but also explain the reasons for these decisions, improving the interpretability and maintainability of the system.

[0124] This method of combining inductive logic programming and reinforcement learning has unique advantages: it not only retains the optimization ability of reinforcement learning in the continuous action space but also obtains the advantages of logic programming in knowledge representation and reasoning. This enables the system to better adapt to complex gateway detection scenarios and generate interpretable decision rules. Especially in scenarios that require quick responses, rule-based reasoning can provide immediate decision suggestions, while reinforcement learning can optimize these rules through continuous learning.

[0125] S2: As Figure 3 shown, based on the deep detection result set, construct a performance prediction model, perform comprehensive gateway rating calculation, and obtain a gateway reliability score table;

[0126] The following is a detailed description of step S2:

[0127] After obtaining the deep detection result set, the system begins to construct a gateway performance prediction model. This model uses the method of ensemble learning, combining the advantages of multiple base models to provide comprehensive performance prediction. Specifically, the system uses three base models, namely random forest, XGBoost, and LightGBM, to construct an ensemble predictor, and each model is responsible for predicting different aspects of performance.

[0128] The random forest is mainly used to handle the prediction of hardware metrics, including the fluctuation range of CPU usage, the stability of memory occupancy, the peak performance of data throughput, etc. Due to the insensitivity of the random forest to outliers, it can provide stable hardware performance prediction results. For example, when the system detects that the fluctuation range of a certain gateway's CPU usage in the past 24 hours is abnormal, the random forest model can accurately predict the possible future load trend.

[0129] XGBoost is responsible for predicting communication quality-related metrics, including the attenuation characteristics of signal strength, the distribution of end-to-end latency, and the trend of packet loss rate. XGBoost's excellent feature combination ability enables it to capture complex communication patterns. For example, by analyzing historical data, XGBoost can predict the changes in communication quality at different times and under different load conditions.

[0130] LightGBM, on the other hand, focuses on the impact assessment of environmental parameters, including factors such as temperature and humidity effects, electromagnetic interference levels, and physical location constraints. Its efficient leaf growth strategy is particularly suitable for handling such multi-category features. For example, LightGBM can evaluate the impact of different environmental conditions on the gateway performance and predict possible performance bottlenecks.

[0131] During the model training process, the system uses five-fold cross-validation to evaluate the model performance and determines the optimal hyperparameter combination through grid search. For the random forest, the system optimizes the number of trees (100 - 500) and the maximum depth (10 - 30); for XGBoost, it focuses on adjusting the learning rate (0.01 - 0.1) and the minimum child weight (1 - 5); for LightGBM, it pays attention to the number of leaves (31 - 127) and the feature sampling ratio (0.6 - 0.8).

[0132] Based on the trained models, the system performs the comprehensive rating calculation for the gateways. This process uses the Analytic Hierarchy Process (AHP) to determine the weights of each evaluation dimension: the weight of the performance prediction result is 0.4, the weight of historical stability is 0.35, and the weight of environmental adaptability is 0.25. The system evaluates each gateway from these three dimensions and finally generates a comprehensive score ranging from 0 to 100.

[0133] When conducting the comprehensive rating, the system also considers historical stability factors, including metrics such as the number of failures and the average fault recovery time in the past 30 days. At the same time, the environmental adaptability assessment takes into account factors such as the stability of the gateway in the current temperature environment, the communication quality under electromagnetic interference conditions, and the location adaptability in the current network topology.

[0134] Finally, the system integrates the prediction results, rating results, and related parameters into a gateway reliability score table. This score table not only includes the current status score of each gateway but also information such as future performance prediction, applicable scenario suggestions, and load capacity assessment. For example, for gateways with a score above 90, the system will recommend that they can undertake more device connections; for gateways with a score below 70, the system will give specific optimization suggestions or recommend reducing their load.

[0135] This evaluation method based on multi-model integration can comprehensively and accurately evaluate the performance status and potential risks of the gateway, providing a reliable decision-making basis for subsequent load balancing optimization. Especially in a complex network environment, this multi-dimensional evaluation method can help the system make more accurate load distribution decisions.

[0136] S2 specifically includes:

[0137] S2.1: Based on the deep detection result set, construct an evaluation system by analyzing hardware metrics, communication quality, and environmental parameters to obtain the gateway performance prediction model parameter set;

[0138] In S2.1, the system constructs a multi-dimensional evaluation system based on the deep detection result set. In terms of hardware metrics, the system mainly focuses on indicators such as the fluctuation range of CPU usage, the stability of memory occupancy, and the peak performance of data throughput; in terms of communication quality, the system analyzes the attenuation characteristics of signal strength, the distribution of end-to-end delay, and the change trend of packet loss rate; in terms of environmental parameters, factors such as temperature and humidity effects, electromagnetic interference levels, and physical location constraints are considered. The system uses an ensemble learning method, applying three models, namely Random Forest, XGBoost, and LightGBM, to the prediction of these different types of indicators respectively.

[0139] Through five-fold cross-validation and grid search, the system optimizes the hyperparameters of each model. For the Random Forest model, the prediction accuracy of hardware metrics is optimized by adjusting the number of trees (100 - 500) and the maximum depth (10 - 30); for the XGBoost model, the learning rate (0.01 - 0.1) and the minimum child weight (1 - 5) are mainly optimized to improve the accuracy of communication quality prediction; for the LightGBM model, the prediction effect of environmental parameters is optimized by adjusting the number of leaves (31 - 127) and the feature sampling ratio (0.6 - 0.8). Finally, the system integrates these optimized model parameters, feature weights, prediction thresholds, and other information into the gateway performance prediction model parameter set.

[0140] S2.2: Based on the gateway performance prediction model parameter set, historical stability, and environmental adaptability factors, perform dynamic evaluation of performance indicators to obtain the gateway rating result set;

[0141] In S2.2, the system performs dynamic evaluation of performance indicators based on the prediction model parameter set. First, the system calculates the coefficient of variation of key indicators to evaluate the fluctuation degree of gateway performance, counts the number of failures and the average mean time to recovery in the past 30 days, and evaluates the performance fluctuation of the gateway under different load conditions, so as to obtain the historical stability score. For example, when the number of failures of a certain gateway in the past 30 days is less than 3 times and the average mean time to recovery is less than 5 minutes, its historical stability score will be relatively high.

[0142] Meanwhile, the system evaluates the environmental adaptability of the gateway, including its stability in the current temperature environment (e.g., evaluating the performance of the gateway in a high-temperature environment above 35°C), the communication quality under electromagnetic interference conditions (such as evaluating the signal stability in an industrial environment), and the location adaptability in the current network topology (such as evaluating the connection reliability of the gateway at the edge location). The system uses the fuzzy comprehensive evaluation method to comprehensively obtain the environmental adaptability score. Finally, the system integrates these evaluation results into the gateway rating result set.

[0143] Among them, the dynamic evaluation of the performance indicators in S2.2 specifically includes:

[0144] S2.2.1: Evaluate the fluctuation degree of the gateway performance by calculating the coefficient of variation of key indicators, count the number of failures and the average time to recover from failures in the past 30 days, and evaluate the performance fluctuation of the gateway under different load conditions to obtain the historical stability score;

[0145] S2.2.2: Evaluate the stability of the gateway in the current temperature environment, the communication quality under electromagnetic interference conditions, and the location adaptability in the current network topology, and obtain the environmental adaptability score through the fuzzy comprehensive evaluation method.

[0146] Exemplarily, in S2.2.1, the historical stability of gateway A in a certain industrial park is evaluated. First, the system calculates the coefficient of variation of key indicators. For example, the CPU usage data of gateway A in the past 7 days: the average value is 45%, the standard deviation is 9%, then the coefficient of variation is 0.2 (standard deviation / average value). The same method is also applied to indicators such as memory usage (coefficient of variation 0.15) and network throughput (coefficient of variation 0.25). These coefficients of variation reflect the stability of the gateway performance, and the smaller the coefficient of variation, the more stable the performance.

[0147] Next, the system counts the failure situations of gateway A in the past 30 days: a total of 3 failures occurred, 2 of which were short-term response delays caused by sudden load increases (recovering in 3 minutes and 4 minutes respectively), and 1 was a planned restart due to system updates (taking 5 minutes). The average time to recover from failures is 4 minutes, which is lower than the system-set warning line of 10 minutes, showing good performance.

[0148] In the performance evaluation under different load conditions, the system simulated three load scenarios: light load (the number of connected devices does not exceed 30% of the total capacity), medium load (the number of connected devices is between 30% - 70% of the total capacity), and heavy load (the number of connected devices exceeds 70% of the total capacity). Gateway A performed stably under light and medium load conditions, with the CPU usage rate fluctuating by no more than 10%; however, under heavy load conditions, once the number of connected devices exceeded 85% of the total capacity, the response time increased sharply, and the CPU usage rate fluctuated up to 30%. Based on these data, the system gave a historical stability score of 85 points.

[0149] In S2.2.2, the system conducted an environmental adaptability evaluation on Gateway A. In terms of temperature adaptability, since Gateway A is installed in an air-conditioned computer room, the ambient temperature remains within the range of 18 - 25°C all year round, far lower than the critical value of 40°C. Monitoring data shows that even when the air conditioner has a short-term failure in summer and the room temperature rises to 32°C, the gateway performance can still remain stable. Therefore, it obtained a temperature adaptability score of 92 points.

[0150] In terms of electromagnetic interference evaluation, since Gateway A is located near an industrial workshop, it is often subject to electromagnetic interference from various mechanical equipment. Through analysis, it is found that during normal working hours (8:00 - 20:00 every day), the signal strength averages at -65 dBm, and the signal quality is good; however, when large equipment starts (about 3 - 4 times a day), the signal strength will briefly drop to -75 dBm, and the communication quality slightly decreases but is still within the acceptable range. Therefore, it obtained an electromagnetic interference adaptability score of 78 points.

[0151] In terms of network topology location adaptability, Gateway A is located at the core of the campus network, and the network hops to upstream and downstream devices do not exceed 2 hops, and there are redundant link protections. Tests show that even when the main link fails, the standby link can complete the switch within 200 ms to ensure business continuity. Therefore, it obtained a location adaptability score of 88 points.

[0152] The system uses the fuzzy comprehensive evaluation method to calculate the final environmental adaptability score based on temperature adaptability (weight 0.3), electromagnetic interference adaptability (weight 0.4), and location adaptability (weight 0.3): 92×0.3 + 78×0.4 + 88×0.3 = 85.2 points.

[0153] This multi-dimensional evaluation method can comprehensively reflect the adaptability of the gateway in the actual operating environment, providing a reliable basis for gateway selection and load distribution. Especially in a complex industrial environment, the environmental adaptability score can help the system better predict and respond to possible performance fluctuations.

[0154] S2.3: Based on the gateway rating result set, formulate a gateway selection strategy through the multi-objective optimization algorithm, and generate the gateway reliability score table.

[0155] In S2.3, the system formulates a gateway selection strategy through a multi-objective optimization algorithm. Specifically, the system sets three core optimization objectives: maximizing the overall service quality of the system, minimizing the load imbalance degree, and minimizing the device migration cost. The system uses an improved NSGA-III (Non-dominated Sorting Genetic Algorithm III) to solve this multi-objective optimization problem. The search efficiency is improved by introducing adaptive crossover and mutation operators, the quality of the solution is enhanced by introducing a local search strategy, and the diversity of the solution is improved by introducing a reference point adaptive adjustment mechanism.

[0156] For example, when the system needs to select an access gateway for a new device, it will consider simultaneously: the ratings of current gateways (gateways with scores above 90 can be given priority), the load balancing situation (to avoid excessive load on a certain gateway), and the potential migration cost (preferably select a gateway with a relatively close physical location). The system selects an optimal compromise solution from the Pareto optimal solution set, and integrates this selection strategy, along with information such as the reliability scores, applicable scenarios, and load suggestions of each gateway, into the final gateway reliability score table. This score table not only contains static score data but also includes dynamic adjustment suggestions, providing comprehensive decision-making support for subsequent load balancing.

[0157] S3: As Figure 4 shown, based on the gateway reliability score table, initialize the hash ring structure through an improved consistent hashing algorithm, construct a virtual node optimization allocation scheme, and optimize the allocation strategy for devices with close physical locations through a locality-sensitive hashing mechanism to obtain the optimal device allocation scheme;

[0158] After obtaining the gateway reliability score table, the system first constructs a consistent hash ring structure. The system selects MurmurHash3 as the basic hash function, which has excellent randomness and high computational performance. The system sets the hash value range from 0 to 2^32 - 1 to form a hash ring that is connected end to end. To solve the problem that node distribution may be uneven in the traditional consistent hashing algorithm, the system introduces the concept of a balance factor. By calculating the hash space size between adjacent gateway nodes, the hash value of the node is dynamically adjusted to ensure that the difference in the hash space size responsible for each node is within 20%.

[0159] To cope with possible failures of gateway nodes, the system implements a backup partition strategy on the hash ring. For each primary partition, the system selects two adjacent partitions in the clockwise direction on the hash ring as its backup partitions and establishes a primary-backup mapping relationship. When the gateway where the primary partition is located fails, the system can quickly transfer the load to the backup partition to ensure the continuity of the service.

[0160] In terms of virtual node allocation, the system has established a dynamic adjustment model for virtual nodes. Based on the reliability scores of gateways, this model uses a non-linear mapping function to calculate the basic number of virtual nodes for each gateway: gateways with scores above 90 receive 100 - 150 virtual nodes, those with scores between 80 - 90 receive 50 - 100 virtual nodes, those with scores between 70 - 80 receive 30 - 50 virtual nodes, and those with scores below 70 only retain 10 - 30 virtual nodes. By monitoring the performance metrics of gateways in real-time, including CPU usage, memory occupancy, network throughput, etc., the system dynamically adjusts the number of virtual nodes.

[0161] To ensure the uniformity of virtual node distribution, the system uses an improved consistent hashing algorithm for virtual node placement. The system calculates the "pressure coefficient" (the density and load conditions of surrounding nodes) for each location and selects the location with a lower pressure coefficient to place virtual nodes, avoiding local aggregation phenomena.

[0162] In terms of optimizing device allocation, the system introduces a Locality-Sensitive Hashing (LSH) mechanism. For each device, the system collects information such as its physical coordinates (latitude and longitude or indoor coordinates), network topology location, communication latency, etc., to construct a multi-dimensional feature vector. The system uses a p-stable distribution (such as the Cauchy distribution) to construct an LSH function family and improves the query efficiency through a multi-level LSH index structure.

[0163] When a gateway needs to be selected for a device, the system first quickly locates the location bucket to which the device belongs through LSH, and then selects a suitable gateway node within the hash ring area corresponding to that bucket. The selection process not only considers the proximity of the physical location but also the current load status and performance score of the gateway. To handle boundary cases, the system has implemented a cross-bucket load balancing mechanism. When the load of the gateways within a certain location bucket is too high, some devices can be allocated to the gateways of adjacent buckets.

[0164] For example, assume that there is a new device in an industrial park that needs to access the network. The system first determines through LSH that the device is located in the southeast corner area of the park (location bucket 1). There are three gateway nodes on the hash ring in this area: Gateway A (score 92, current load 60%), Gateway B (score 85, current load 75%), and Gateway C (score 88, current load 45%). Although Gateway A has the highest score, considering that it already bears a relatively high load, the system finally selects the gateway C with a lower load as the access point for this device. This allocation strategy that comprehensively considers location proximity and load balancing not only ensures communication efficiency but also avoids the problem of overloading a single gateway.

[0165] Through this optimized solution that combines consistent hashing and locality-sensitive hashing, the system can, while ensuring load balancing, assign devices to gateways with physically close locations as much as possible, thereby improving the overall network performance and reliability. Especially in scenarios where devices frequently join and leave dynamically, this solution can effectively reduce the overhead of network reconstruction and improve the scalability and stability of the system.

[0166] In another embodiment, a multi-agent approach can also be adopted to construct the hierarchical dynamic detection system and construct the performance prediction model based on the deep detection result set, specifically including:

[0167] A1: Deploy local detection agents at the gateway nodes and global coordination agents at the core nodes to form a multi-agent network topology;

[0168] In A1, the system deploys local detection agents (Local Detection Agent, LDA) at each gateway node and global coordination agents (Global Coordination Agent, GCA) at the core nodes. Each LDA is responsible for monitoring and managing the status of a single gateway, including collecting performance data, making local decisions, etc. For example, there are 5 gateway nodes deployed in an industrial park, and each gateway node is installed with an LDA, while a GCA is deployed on the core server in the central control room. These agents are connected to each other through the network to form a hierarchical multi-agent network topology structure.

[0169] A2: Based on the multi-agent network topology, maintain the stability of the multi-agent network through a lease-based heartbeat mechanism and construct the agent collaboration framework;

[0170] The system constructs a lease-based heartbeat mechanism to maintain the stability of the multi-agent network. Each LDA needs to send a heartbeat message to the GCA regularly (such as every 30 seconds), including its current status and brief performance data. The GCA will assign a lease to each LDA, and the validity period is usually three times the heartbeat interval (such as 90 seconds). If the GCA does not receive the heartbeat message of a certain LDA before the lease expires, it will trigger the corresponding fault handling mechanism. At the same time, the system uses the Paxos algorithm to ensure the consistency of distributed decisions and implements a message middleware based on the publish / subscribe mode.

[0171] A3: Based on the agent collaboration framework, implement local decisions through the Actor-Critic architecture, optimize the global detection strategy through the meta-learning method, and obtain the deep detection result set;

[0172] A distributed learning mechanism is implemented based on the intelligent agent collaboration framework. Each LDA adopts the Actor-Critic architecture to achieve local decision-making capabilities: the Actor network is responsible for generating detection actions (such as adjusting the detection frequency, selecting detection metrics, etc.), and the Critic network is responsible for evaluating the value of these actions. The GCA optimizes the global detection strategy through meta-learning methods. It can extract common knowledge from the experiences of each LDA to form a more general strategy pattern. For example, a successful strategy adopted by an LDA when dealing with a sudden increase in gateway load can be quickly mastered by other LDAs through the meta-learning of the GCA.

[0173] A4: Receive the deep detection result set, construct a local evaluation model for each gateway agent, and integrate the evaluation results of each local detection agent through a swarm intelligence algorithm to form a gateway group evaluation model;

[0174] A local evaluation model is constructed for each gateway agent. These models evaluate the local gateway from three dimensions: performance dimension (processing capacity, response time, etc.), stability dimension (failure rate, recovery ability, etc.), and adaptability dimension (load elasticity, environmental adaptability, etc.). Through the swarm intelligence algorithm, the system integrates the evaluation results of each LDA to form a complete gateway group evaluation model. For example, when multiple gateways in a certain area detect a similar performance degradation pattern, the swarm intelligence algorithm can quickly identify possible system-level problems.

[0175] A5: Based on the gateway group evaluation model, calculate the marginal contribution of each gateway to the overall service quality through a coalition game model, dynamically adjust the load distribution weights, and generate the gateway reliability score table.

[0176] Specifically, a coalition game model is introduced to optimize the load distribution. By calculating the marginal contribution (Shapley value) of each gateway to the overall service quality, the system can more accurately evaluate the importance of each gateway. For example, a gateway located at a critical position in the network may have a high Shapley value due to its special geographical location and network topology relationship, although its independent performance indicators may not be optimal. Based on these calculation results, the system dynamically adjusts the load distribution weights and finally generates a gateway reliability score table containing detailed scoring and weight information.

[0177] This multi-agent based implementation scheme significantly improves the scalability and adaptability of the system through a distributed architecture design and collaborative learning mechanism. Especially in the scenario of large-scale gateway deployment, the multi-agent architecture can better handle complex coordination and optimization problems and provide a more flexible and efficient solution.

[0178] Among them, S3 specifically includes:

[0179] S3.1: Initialize the hash ring structure based on the gateway reliability score table, and map the gateway nodes to the hash space through the improved hash algorithm to obtain the initial hash ring mapping relationship set;

[0180] In S3.1, the system first initializes the hash ring structure based on the gateway reliability score table. The system selects MurmurHash3 as the basic hash function, which performs well in terms of computing performance and randomness. The hash value range is set from 0 to 2^32 - 1, forming a complete hash ring. To solve the problem of uneven node distribution in traditional consistent hashing, the system introduces a balance factor mechanism.

[0181] For example, there are four gateway nodes (A, B, C, D) in an industrial park, and their initial hash values are 0x2000000, 0x4000000, 0x8000000, and 0xE000000 respectively. The system detects that the hash space (0x4000000 to 0x8000000) between nodes B and C is significantly larger than other intervals. At this time, through balance factor adjustment, the system appropriately shifts the hash value of node C to the left to 0x6000000, so that the difference in the size of the hash space between adjacent nodes is controlled within 20%.

[0182] At the same time, the system implements a backup partition strategy. For each primary partition, two adjacent partitions are selected in the clockwise direction on the hash ring as backups. For example, the primary partition backup of gateway A is assigned to gateways B and C. When A fails, the system can transfer its load to these two backup gateways within 200 ms.

[0183] S3.2: Based on the initial hash ring mapping relationship set, dynamically adjust the number of virtual nodes of each gateway according to the reliability score of the gateway to obtain the virtual node allocation plan;

[0184] In S3.2, the system constructs a dynamic adjustment model for virtual nodes based on the initial hash ring mapping relationship set. Taking gateway A as an example, its reliability score is 92 points, and the system initially assigns 120 virtual nodes to it. The system determines the number of virtual nodes through a non-linear mapping function: more than 90 points are assigned 100 - 150 nodes, 80 - 90 points are assigned 50 - 100 nodes, 70 - 80 points are assigned 30 - 50 nodes, and less than 70 points are only assigned 10 - 30 nodes.

[0185] The system monitors the performance metrics of gateway A in real time: when it observes that its CPU usage rate continuously exceeds 80%, the system gradually reduces the number of its virtual nodes from 120 to 80; when the load drops below 30%, the number of virtual nodes is appropriately increased. This dynamic adjustment ensures the flexibility of load distribution. To avoid local aggregation of virtual nodes, the system calculates the "pressure coefficient" at each location. For example, if there are already many nodes and high load in a certain area, new virtual nodes will be placed in areas with less pressure.

[0186] S3.3: Based on the virtual node allocation scheme, optimize the physical location information of the device through the locality-sensitive hashing mechanism to generate the optimal device allocation scheme.

[0187] In S3.3, the system optimizes device allocation through the locality-sensitive hashing (LSH) mechanism. Suppose a new temperature sensor needs to be connected to the network, and its physical coordinates are (120.5, 30.2), and the network latency requirement is less than 100 ms. The system first constructs the feature vector of the device, including information such as location coordinates and network requirements. The LSH function family constructed using the p-stable distribution maps this feature vector to specific location buckets.

[0188] The system finds that the device belongs to the "East Area Industrial Park" location bucket, and there are three candidate gateways (E, F, G) in this area. Although the physical location of gateway E is the closest, its current load has reached 85%; although the load of gateway F is relatively low (45%), its communication latency is close to 150 ms; finally, the system selects gateway G with a load of 60% and a latency of 80 ms as the optimal access point. To handle possible sudden increases in load, the system also reserves gateway H in adjacent location buckets as an alternative solution.

[0189] This allocation strategy that combines physical location and performance metrics not only ensures communication efficiency but also maintains the balance of the overall load. For example, in a scenario of batch device access, the system successfully distributes 100 new devices to 5 different gateways, which not only ensures the principle of nearby access (average communication latency controlled within 50 ms) but also avoids overloading a single gateway (the highest load does not exceed 75%).

[0190] S4: Based on the optimal device allocation scheme, formulate a device migration strategy, execute device migration and conduct verification, and output a migration verification report;

[0191] In S4, based on the optimal device allocation scheme, analyze the business importance, service level, and migration risk of the device to formulate the device migration strategy plan.

[0192] Among them, S4.1 also includes:

[0193] S4.1.1: Based on the device migration strategy plan, perform device migration through a fault detection and rollback mechanism to obtain the migration execution result set;

[0194] S4.1.2: Based on the migration execution result set, conduct verification through performance testing, stability evaluation, and business continuity check, and generate the migration verification report.

[0195] The following uses an example of an IoT system in an industrial park to illustrate the device migration and verification process:

[0196] Suppose there are three main gateways (Gateway-A, Gateway-B, and Gateway-C) and a newly deployed gateway (Gateway-D) in the park. Due to the performance degradation of Gateway-A, the system needs to migrate some of its devices to the newly deployed Gateway-D. Gateway-A is currently connected to 100 IoT devices, including 30 high-priority production line monitoring devices, 50 medium-priority environmental monitoring devices, and 20 low-priority asset tracking devices.

[0197] Based on the analysis of the business importance of the devices, the system divides the migration strategy into three stages:

[0198] The first stage (pre-migration preparation): Select to perform the migration during the daily maintenance period of the production line (2:00 - 4:00 in the early morning), when the impact on the business is minimal. The system has prepared a detailed rollback plan, including device configuration backup, network connection parameter archiving, etc.

[0199] The second stage (hierarchical migration execution):

[0200] The first round of migration (2:15): Select 5 low-priority asset tracking devices for trial migration;

[0201] The second round of migration (2:45): Migrate 20 environmental monitoring devices;

[0202] The third round of migration (3:15): Migrate the remaining 25 environmental monitoring devices and 10 production line monitoring devices;

[0203] The fourth round of migration (3:45): Complete the migration of the remaining devices;

[0204] The third stage (emergency preparation): Reserve a 15-minute rollback window for each group of devices.

[0205] Taking the 20 environmental monitoring devices in the second round of migration as an example, the system executed the following migration process (S4.1.1):

[0206] 1. Record the baseline performance data before migration: data collection frequency (60 data points per minute), average response time (120 ms), data accuracy rate (99.95%);

[0207] 2. Execute the migration program:

[0208] 2:45:00 Notify Gateway-A and Gateway-D to enter the migration preparation state;

[0209] 2:45:05 Start the device configuration migration and update the gateway connection parameters of 20 devices;

[0210] 2:45:15 Switch the gateway connection of each device one by one, with a 3-second switching time reserved for each device;

[0211] 2:45:35 Complete the connection switching of all devices;

[0212] 2:45:40 Conduct a preliminary connectivity verification;

[0213] During the migration process, the system monitors key metrics in real-time:

[0214] Network connection status: Keep 100% of the devices online;

[0215] Data transmission delay: Slightly increased from the original 120 ms to 125 ms, within an acceptable range;

[0216] Data integrity: No data loss during migration;

[0217] During the verification phase (S4.1.2), taking one of the environmental monitoring devices (temperature and humidity sensor TH-001) as an example, the system conducted a comprehensive verification:

[0218] Performance verification:

[0219] Data collection: Maintain a data collection frequency of 60 data points per minute;

[0220] Response time: Optimized from the original 120 ms to 115 ms;

[0221] Data accuracy rate: Maintain at the level of 99.95%;

[0222] Signal strength: Improved from the original -75 dBm to -65 dBm;

[0223] Stability verification:

[0224] Execute continuous monitoring for 24 hours to verify the performance of the device at different time periods;

[0225] Simulate network fluctuation conditions to verify the automatic reconnection ability of the device;

[0226] Test data caching mechanism to ensure that temporary network interruptions do not result in data loss;

[0227] Business continuity verification:

[0228] Check the data integration with the environmental control system;

[0229] Verify the normal operation of the alarm triggering mechanism;

[0230] Confirm the integrity and accessibility of historical data;

[0231] Finally, the system generated a detailed migration verification report, recording:

[0232] Performance comparison data before and after migration;

[0233] Records of temporary performance fluctuations during the migration process;

[0234] Detailed results of various verification tests;

[0235] Potential problems found and optimization suggestions;

[0236] The report shows that this migration not only successfully ensured business continuity but also achieved an improvement in device performance through the optimized deployment of Gateway-D, especially in terms of signal strength and response time. This successful case provides valuable experience for the subsequent optimization and upgrade of other gateways.

[0237] As Figure 5 shown, the embodiment of the present application further provides a multi-gateway communication link intelligent optimization system, which includes a hierarchical dynamic detection module, a gateway evaluation module, a load balancing optimization module, and a migration verification module, where:

[0238] The hierarchical dynamic detection module is used to obtain real-time gateway status data, where the real-time gateway status data includes gateway load, signal strength, and response time, construct a hierarchical dynamic detection system, optimize the dynamic scheduling strategy through an artificial intelligence model, and obtain a deep detection result set;

[0239] The gateway evaluation module is used to construct a performance prediction model based on the deep detection result set, perform comprehensive gateway rating calculation, and obtain a gateway reliability score table;

[0240] The load balancing optimization module is used to initialize the hash ring structure based on the gateway reliability score table through an improved consistent hashing algorithm, construct a virtual node optimization allocation scheme, and optimize the allocation strategy of devices with physically close locations through a locality-sensitive hashing mechanism to obtain an optimal device allocation scheme;

[0241] A migration verification module, which is used to formulate a device migration strategy based on the optimal device allocation scheme, execute device migration and perform verification, and output a migration verification report.

[0242] Each module in the above system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0243] In one embodiment, a computer device is further provided, as Figure 6 shown. This computer device is the multi-gateway communication link intelligent optimization system mentioned in the above method embodiment. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface.

[0244] Among them, the processor of this computer device is used to provide computing and control capabilities, and can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation. The processor may include one or more processors, for example, including one or more central processing units (central processing unit, CPU). In the case where the processor is a single CPU, the CPU can be a single-core CPU or a multi-core CPU. The processor may further include one or more dedicated processors, and the dedicated processors may include GPU, FPGA, etc., for acceleration processing. The processor is used to call the program code and data in the memory and execute the steps in the above method embodiment. For specific details, refer to the description in the method embodiment, which will not be elaborated here.

[0245] The memory of the computer device includes, but is not limited to, non-volatile storage media and internal memory. The non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media.

[0246] The input / output interface of the computer device is used for the processor to exchange information with external devices.

[0247] The communication interface of the computer device is used to communicate with an external terminal through a network connection.

[0248] When the computer program is executed by the processor, it implements the above method.

[0249] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the division of each unit / module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings, or direct couplings, or communication connections can be through some interfaces, and the indirect couplings or communication connections of systems or units can be in electrical, mechanical, or other forms.

[0250] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0251] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape, a magnetic disk, or an optical medium, such as a digital versatile disc (DVD), or a semiconductor medium, such as a solid state disk (SSD), etc.

[0252] The system of the embodiments of the present disclosure can execute the methods provided by the embodiments of the present disclosure, and the implementation principles are similar. The actions performed by each module in the system of the embodiments of the present disclosure correspond to the steps in the methods of the embodiments of the present disclosure. For the detailed function descriptions of the modules of the system, reference can be specifically made to the descriptions in the corresponding methods shown above, and details are not described herein again.

[0253] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art in the technical field of the present disclosure, without departing from the technical concept of the solution of the present disclosure, other similar implementation means based on the technical idea of the present disclosure also belong to the protection scope of the embodiments of the present disclosure.

[0254] The above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art in the technical field of the present application can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present application, and these modifications or substitutions should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An intelligent optimization method for multi-gateway communication links, characterized in that, The method includes the following steps: Obtain real-time gateway status data, where the real-time gateway status data includes gateway load, signal strength, and response time, construct a hierarchical dynamic detection system, optimize the dynamic scheduling strategy through an artificial intelligence model, and obtain a deep detection result set; Based on the deep detection result set, construct a performance prediction model, perform comprehensive gateway rating calculation, and obtain a gateway reliability score table; Based on the gateway reliability score table, initialize the hash ring structure through an improved consistent hashing algorithm, construct a virtual node optimization allocation scheme, and optimize the allocation strategy of devices with physically close locations through a locality-sensitive hashing mechanism to obtain an optimal device allocation scheme; Based on the optimal device allocation scheme, formulate a device migration strategy, perform device migration and verification, and output a migration verification report; Among them, constructing the hierarchical dynamic detection system, optimizing the dynamic scheduling strategy through an artificial intelligence model, and obtaining the deep detection result set includes: Construct a two-layer detection framework including a basic detection layer and a deep detection layer, and generate a detection framework configuration parameter set based on the real-time gateway status data; Based on the detection framework configuration parameter set and historical detection data, construct a reinforcement learning environment, define the gateway state space and the action space of detection behaviors, and use the reward function of detection effect and resource consumption and the action space to obtain the reinforcement learning environment parameter set; Based on the reinforcement learning environment parameter set, train a detection strategy model through a deep Q-learning algorithm to generate the deep detection result set; Among them, initializing the hash ring structure through an improved consistent hashing algorithm, constructing a virtual node optimization allocation scheme, and optimizing the allocation strategy of devices with physically close locations through a locality-sensitive hashing mechanism to obtain an optimal device allocation scheme includes: Initialize the hash ring structure based on the gateway reliability score table, map the gateway nodes to the hash space through the improved hashing algorithm to obtain the initial hash ring mapping relationship set; Based on the initial hash ring mapping relationship set, dynamically adjust the number of virtual nodes of each gateway according to the reliability score of the gateway to obtain the virtual node allocation scheme; Based on the virtual node allocation scheme, optimize the physical location information of the devices through the locality-sensitive hashing mechanism to generate the optimal device allocation scheme.

2. The method according to claim 1, characterized in that, Before obtaining the reinforcement learning environment parameter set using the reward function of detection effect and resource consumption and the action space, the method further includes: Set the reward value and penalty value of the reward function, where a first preset reward value is given when a fault is detected within a preset time, and a second preset reward value is given when the detection resource utilization rate is maintained within a preset range; A first preset penalty value is given when there is waste of detection resources, and a second preset penalty value is given when a fault fails to be detected in time.

3. The method according to claim 1, characterized in that Training the detection strategy model through the deep Q-learning algorithm to generate the deep detection result set includes: Construct an inductive logic learning environment and construct an initial logic rule base including basic predicates and background knowledge; Construct a fusion learning framework that combines inductive logic programming and reinforcement learning, and configure the Actor-Critic network structure and the logic rule learner; Introduce an answer set reasoning mechanism to guide the exploration process of the reinforcement learning, and generate an action candidate set based on logical rules; Based on the action candidate set, discover the first decision pattern through meta-rule learning, including: collecting decision-making experiences through the interaction between the Actor-Critic network and the environment in the experience accumulation stage; inducing the second decision rule through the inductive logic programming technology in the rule induction stage based on the decision-making experiences; evaluating the second decision rule in the knowledge integration stage based on the second decision rule; Integrate the learned detection strategy, the logic rule library, and the verification data to generate the deep detection result set.

4. The method according to claim 1, characterized in that, Construct a performance prediction model, perform gateway comprehensive rating calculation, and obtain a gateway reliability scoring table, including: Based on the deep detection result set, construct an evaluation system by analyzing hardware metrics, communication quality, and environmental parameters to obtain a gateway performance prediction model parameter set; Based on the gateway performance prediction model parameter set, historical stability, and environmental adaptability factors, perform dynamic evaluation of performance indicators to obtain the gateway rating result set; Based on the gateway rating result set, formulate a gateway selection strategy through a multi-objective optimization algorithm to generate the gateway reliability scoring table.

5. The method according to claim 4, characterized in that, Perform the dynamic evaluation of the performance indicators, including: Evaluate the fluctuation degree of the gateway performance by calculating the coefficient of variation of key indicators, count the number of faults and the average fault recovery time in the past 30 days, and evaluate the performance fluctuation of the gateway under different load conditions to obtain a historical stability score; Evaluate the stability of the gateway in the current temperature environment, the communication quality under electromagnetic interference conditions, and the position adaptability in the current network topology, and obtain an environmental adaptability score through the fuzzy comprehensive evaluation method.

6. The method according to claim 1, wherein Formulate the device migration strategy, including: Based on the optimal device allocation plan, analyze the business importance, service level, and migration risk of the device to formulate the device migration strategy plan; Among them, based on the optimal device allocation plan, formulate a device migration strategy, execute the device migration and conduct verification, and output a migration verification report, including: Based on the device migration strategy plan, execute the device migration through a fault detection and rollback mechanism to obtain the migration execution result set; Based on the migration execution result set, conduct verification through performance testing, stability evaluation, and business continuity inspection to generate the migration verification report.

7. The method according to claim 1, wherein Construct the hierarchical dynamic detection system and construct the performance prediction model based on the deep detection result set, including: Deploy local detection agents at the gateway nodes and global coordination agents at the core nodes to form a multi-agent network topology; Based on the multi-agent network topology, maintain the stability of the multi-agent network through a lease-based heartbeat mechanism to construct the agent collaboration framework; Based on the agent collaboration framework, local decision-making is achieved through the Actor-Critic architecture, and the global detection strategy is optimized through the meta-learning method to obtain the deep detection result set; Receive the deep detection result set, construct a local evaluation model for each gateway agent, and integrate the evaluation results of each local detection agent through the swarm intelligence algorithm to form a gateway group evaluation model; Based on the gateway group evaluation model, calculate the marginal contribution of each gateway to the overall service quality through the coalition game model, dynamically adjust the load distribution weight, and generate the gateway reliability score table.

8. An intelligent optimization system for multi-gateway communication links, characterized in that, The system includes a hierarchical dynamic detection module, a gateway evaluation module, a load balancing optimization module, and a migration verification module, where: The hierarchical dynamic detection module is used to obtain the real-time status data of the gateway. The real-time status data of the gateway includes gateway load, signal strength, and response time. Construct a hierarchical dynamic detection system, and optimize the dynamic scheduling strategy through the artificial intelligence model to obtain the deep detection result set; The gateway evaluation module is used to construct a performance prediction model based on the deep detection result set, perform comprehensive rating calculation of the gateway, and obtain the gateway reliability score table; The load balancing optimization module is used to initialize the hash ring structure through the improved consistent hashing algorithm based on the gateway reliability score table, construct a virtual node optimization allocation scheme, and optimize the allocation strategy of devices with physically close positions through the locality-sensitive hashing mechanism to obtain the optimal device allocation scheme; The migration verification module is used to formulate a device migration strategy based on the optimal device allocation scheme, perform device migration and verification, and output a migration verification report; Among them, constructing the hierarchical dynamic detection system, optimizing the dynamic scheduling strategy through the artificial intelligence model, and obtaining the deep detection result set includes: Construct a two-layer detection framework including a basic detection layer and a deep detection layer, and generate the detection framework configuration parameter set based on the real-time status data of the gateway; Based on the detection framework configuration parameter set and historical detection data, construct a reinforcement learning environment, define the gateway state space and the action space of the detection behavior, and use the reward function of the detection effect and resource consumption and the action space to obtain the reinforcement learning environment parameter set; Based on the reinforcement learning environment parameter set, train a detection strategy model through the deep Q-learning algorithm to generate the deep detection result set; Among them, initializing the hash ring structure through the improved consistent hashing algorithm, constructing a virtual node optimization allocation scheme, and optimizing the allocation strategy of devices with physically close positions through the locality-sensitive hashing mechanism to obtain the optimal device allocation scheme includes: Initialize the hash ring structure based on the gateway reliability score table, and map the gateway nodes to the hash space through the improved hashing algorithm to obtain the initial hash ring mapping relationship set; Based on the initial hash ring mapping relationship set, dynamically adjust the number of virtual nodes of each gateway according to the reliability score of the gateway to obtain the virtual node allocation scheme; Based on the virtual node allocation scheme, optimize the physical location information of the device through the locality-sensitive hashing mechanism to generate the optimal device allocation scheme.

Citation Information

Patent Citations

  • Time delay guarantee intelligent routing method and system of TSN

    CN118487988A