Industrial wireless network resource scheduling method based on multi-weight quality assessment model
By adopting a multi-weight quality evaluation model and an improved Q-learning algorithm in industrial wireless networks, time slot selection is optimized, and the problem of lack of targeted time slot selection and insufficient transmission quality evaluation in the prior art is solved, and data transmission with high reliability and low latency is achieved.
Patent Information
- Application Number
- CN202411839740.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-12-13
AI Technical Summary
现有工业无线网络资源调度方法在时隙选择上缺乏针对性,无法有效评估传输质量,导致网络在高干扰环境下表现不佳,传输失败率高,网络延迟大。
The resource scheduling method based on the multi-weight quality evaluation model is adopted, combined with the improved Q-learning algorithm, time slot selection in the TSCH network is optimized. Through edge computing, network state analysis and policy optimization can be processed, the computational pressure and energy consumption of sensor nodes can be reduced, and dynamic adjustment of time slots and channels can be achieved through multi-weight quality evaluation model and Pareto optimization.
It realizes high reliability and end-to-end low latency of wireless sensor node data transmission in complex and changeable industrial network environments, significantly improving transmission efficiency and network adaptability and reducing energy consumption.
Smart Images

Figure CN119316944B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet of Things resource scheduling, and particularly to an industrial wireless network resource scheduling method based on a multi-weight quality evaluation model. Background Art
[0002] In the industrial Internet of Things (IIoT) environment, communication between devices needs to have the characteristics of high reliability, low latency, and low energy consumption to support complex automation control and monitoring applications, such as intelligent manufacturing and process control. The network scale in the industrial environment is large and highly dynamic, posing a huge challenge to resource management. To address these challenges, the IETF 6TiSCH working group proposed the 6TiSCH network, which is a network architecture based on IEEE 802.15.4 TSCH (Time Slotted Channel Hopping) and IPv6, aiming to provide efficient multi-hop wireless network connections for low-power Internet of Things (IoT) devices.
[0003] The 6TiSCH network combines the time synchronization and channel hopping mechanisms of TSCH with 6LoWPAN and RPL (IPv6 routing protocol) to provide a standardized scheduling mechanism for low-power, multi-hop networks. In the 6TiSCH network, a reliable and efficient scheduling mechanism is essential, which determines when devices communicate and directly affects the energy efficiency, latency, and reliability of the network. The scheduling algorithm ensures that each node exchanges data with neighbor nodes at an appropriate time by reasonably allocating time slots and channels, maximizing resource utilization, reducing transmission conflicts, and extending the battery life of nodes.
[0004] The 6TiSCH network divides the communication time into fixed-length time slots through TSCH technology and hops between multiple channels, forming the basic unit in scheduling - the "cell". Each cell consists of a slot offset and a channel offset. The slot offset defines the transmission time of the node, and the channel offset is used to determine the frequency of data transmission. Existing scheduling schemes, although providing lightweight scheduling methods, still have the following deficiencies: (1) Randomly select time slots from the candidate pool for scheduling. This randomness ignores the transmission performance differences of each time slot, resulting in poor performance of the network in a high-interference environment, high transmission failure rate, and large network latency. (2) Lack of intelligent evaluation of the transmission quality of time slot-channel combinations. There are significant differences in the transmission quality of different time slot-channel combinations, which leads to inefficient utilization of some time slots, increasing the number of retransmissions and latency. Summary of the Invention
[0005] To overcome the deficiencies of existing methods, the purpose of the present invention is to provide an industrial wireless network resource scheduling method based on a multi-weight quality evaluation model, which combines an improved Q-learning algorithm to optimize the time slot selection in a TSCH network, and solves the problems of lack of pertinence in time slot selection and inability to effectively evaluate transmission quality in existing resource scheduling methods. This method processes complex network state analysis and policy optimization through edge computing, reduces the computing pressure and energy consumption of sensor nodes, and achieves efficient resource scheduling. At the same time, by using a multi-weight quality evaluation model and Pareto optimization, dynamic adjustment of time slots and channels is realized, ensuring high reliability of data transmission and end-to-end low latency of wireless sensor nodes in a complex and changeable industrial network environment. The present invention uses a 6TiSCH simulator to conduct multi-faceted evaluations of quantitative and qualitative indicators, verifying the significant advantages of the solution in terms of transmission efficiency, network adaptability, and energy consumption.
[0006] The present invention is realized through the following technical solutions: An industrial wireless network resource scheduling method based on a multi-weight quality evaluation model, including:
[0007] Step S1A: Internal nodes in the network apply a sliding window mechanism to collect relevant data on the recent network state, including the expected number of transmissions, queue length, packet retransmission rate, and receive-forward distance, and perform data preprocessing to obtain the average value of the expected number of transmissions, the median value of the queue length, the truncated average value of the retransmission rate, and the receive-forward distance, and use them as inputs to the multi-weight quality evaluation model to calculate the transmission quality of time slots in the candidate pool; The multi-weight quality evaluation model is expressed as:
[0008] ;
[0009] Wherein, is the transmission quality of the time slot, are respectively the average value of the expected number of transmissions , the median value of the queue length , the truncated average value of the retransmission rate , and the receive-forward distance weights; Use the Q-learning algorithm of Pareto optimization and adaptive exploration to dynamically adjust the weights;
[0010] Step S2A: Based on the transmission quality of the time slot output by the multi-weight quality evaluation model, combined with the 6P protocol, dynamically manage the time slot resources of the node, including dynamically releasing and supplementing time slots, maintaining and optimizing the time slot pool, selecting the most suitable channel combination through a pseudo-random channel selection mechanism, and the node uses the selected channel for data transmission and updates the scheduling table.
[0011] More preferably, the most suitable channel combination is selected through a pseudo-random channel selection mechanism, specifically: by combining the setting of the channel record table and the active channel set, all channels in the network are divided into available channels and the active channel set, where the active channel set includes the channels that have transmitted in several time slots recorded in the channel record table and their adjacent channels before and after, and the available channels are the remaining channels after the division; for multiple consecutive channel blocks obtained by the channel division, the largest channel block is selected; then the linear congruence formula is applied to generate a random number, and the channel with the lowest probability of collision and interference is selected from the largest channel block.
[0012] More preferably, the node uses the selected channel for data transmission and updates the scheduling table, including: the node uses the selected channel for data transmission, and after the transmission is completed, stores the channel information used in the channel recorder for reference when selecting channels subsequently; and through the 6P protocol, the node sends the relevant data back to the parent node to update the scheduling table, thereby further dynamically releasing and supplementing time slots, maintaining the dynamic update of the time slot candidate pool, and continuously optimizing the time slot selection strategy according to the new transmission data.
[0013] Further, the dynamic adjustment of weights using the Q-learning algorithm with Pareto optimization and adaptive exploration includes the following steps:
[0014] Step S1B: Collect the real-time state data of the network in the edge routing node, including the packet delivery ratio PDR, traffic load TL, and delay; the real-time state data is used as the input of the Q-learning algorithm, initializes the key parameters of Q-learning according to the current network state, and constructs a Q-table; the Q-table is used to store the Q-values of each state-action pair, and the Q-values will be used as the decision basis to guide the routing node on how to allocate resources and adjust the network strategy; design a reward function to quantify the performance of the network and support multi-objective optimization; by introducing Pareto optimization, determine the preference weight vector of each objective in the reward function to represent the priorities of different optimization objectives, ensure that the priorities of different objectives are reasonably considered, and achieve multi-objective optimization of the reward function.
[0015] Step S2B: Update the values in the Q-table according to the output of the reward function in step S1B; adopt an adaptive exploration strategy for the action selection control in each iteration after updating the Q-values, balance between exploration and exploitation by monitoring the fluctuations of network delay, packet delivery ratio, and traffic load, dynamically adjust the exploration rate, and perform a convergence test on the update iteration of the Q-values to monitor whether the change in Q-values between two consecutive iterations is less than a given threshold.
[0016] The beneficial effects of the present invention are as follows: Compared with the prior art, by adopting the edge computing concept, combining a multi-weight quality evaluation model and an improved Q-learning algorithm, the present invention optimizes the time slot selection in the TSCH network, solves the problems of lack of pertinence in time slot selection, lack of transmission quality evaluation, and high node computing pressure in the existing scheduling methods. A multi-weight quality evaluation model based on edge computing is proposed, which comprehensively considers multiple parameters such as the expected number of transmissions, the number of retransmissions, the queue length, and the time slot distance factor, and intelligently evaluates the transmission performance of each time slot-channel combination. This method ensures that nodes can preferentially select time slots with high transmission success rate and low latency for data packet transmission, while reducing the computing overhead of nodes. In addition, by introducing an improved Q-learning algorithm through the edge router, the system can dynamically adjust the weight parameters in the multi-weight quality evaluation model using Pareto optimization according to the real-time changes of the network state, and adjust the exploration rate of Q-learning through an adaptive exploration strategy, thereby achieving dynamic network adaptability. This mechanism enhances the system's response ability to network load and interference, ensures efficient and reliable wireless transmission, and meets the strict requirements for sensor node data interaction in the industrial Internet of Things environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Structural diagram of an industrial wireless network resource scheduling method based on a multi-weight quality evaluation model of the present invention;
[0018] Figure 2 Structural diagram of the method for realizing time slot-channel selection of the edge computing-driven multi-weight quality evaluation model of the present invention;
[0019] Figure 3 Structural diagram of the Q-learning algorithm with Pareto optimization and adaptive exploration of the present invention;
[0020] Figure 4 Structural diagram of the pseudo-random channel selection of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0022] Such as Figure 1As shown in the figure, this embodiment provides an industrial wireless network resource scheduling method based on a multi-weight quality evaluation model, aiming to solve the randomness of slot selection in existing resource scheduling and the lack of slot quality evaluation. At the same time, it ensures the high reliability of data transmission of wireless sensor nodes and the low end-to-end latency of data transmission in a complex and changeable industrial network environment. The present invention proposes a unique multi-weight quality evaluation model: in the process of selecting and evaluating time slots - channels, a multi-weight index calculation scheme is invented, and relevant data of the recent network state is used as data input and preprocessed; after preprocessing, the sliding window mechanism is applied to select the performance indicators that best reflect the recent performance of the candidate time slots of the current node; the historical data is processed by methods such as arithmetic mean, median, and truncated mean to obtain robust statistical features, and the receiving and forwarding distance is comprehensively considered to form the calculation input of the multi-weight quality evaluation model, so as to obtain the time slot with the highest quality score in the candidate pool, and according to the output, the nodes are dynamically released and supplemented with time slots in combination with the 6P protocol, the time slot pool is maintained and dynamically optimized, and then the pseudo-random channel selection method is used to select channels with low interference and low conflict to form a time slot-channel combination for task transmission.
[0023] In the process of dynamically adjusting the weight parameters of the multi-weight quality evaluation model: a Q-learning algorithm of Pareto optimization and adaptive exploration is invented. By collecting the basic performance parameters in the network at the edge router, combining the dynamic adjustment of the reward function by Pareto optimization and the balance between exploration and exploitation of the adaptive exploration strategy to evaluate the overall network state, the weight parameters of the multi-weight quality evaluation model most suitable for the current network state are calculated. These weight parameters will be fed back to the internal nodes in the network through specific data packets (EB packets) to assist them in time slot-channel selection.
[0024] In one embodiment, an industrial wireless network resource scheduling method based on a multi-weight quality evaluation model includes using the multi-weight quality evaluation model for time slot-channel quality evaluation and selection and using the Q-learning algorithm of Pareto optimization and adaptive exploration for dynamic adjustment of weights.
[0025] The multi-weight quality evaluation model for time slot-channel quality evaluation and selection includes the following steps:
[0026] Step S1A: The internal network nodes apply the sliding window mechanism to collect relevant data on the recent network status, including the expected number of transmissions, queue length, packet retransmission rate, and reception-forwarding distance, and perform data preprocessing to reduce the computational overhead while retaining the key features of the data, obtaining the average value of the expected number of transmissions, the median value of the queue length, the truncated average value of the retransmission rate, and the reception-forwarding distance, and using these as the input to the multi-weight quality evaluation model to calculate the transmission quality of the time slots in the candidate pool. These data are used for the calculation of the multi-weight quality evaluation model to ensure that the multi-weight quality evaluation model can accurately evaluate the transmission quality of each time slot-channel combination, providing high-quality decision support for resource scheduling.
[0027] Further, in the embodiment, the process is described as:
[0028] The internal network nodes collect relevant transmission performance data (expected number of transmissions, queue length, i.e., the number of data packets in the transmission queue, packet retransmission rate) within the recent sliding window time, regard it as the first preprocessing, and use the processed data for the calculation of the multi-weight quality evaluation model. In addition, consider the reception-forwarding distance between the node's received time slot and the forwarded time slot, and comprehensively use the above data to ensure that the multi-weight quality evaluation model can evaluate the true quality of each time slot, providing high-quality decision support for scheduling, expressed as:
[0029] (1);
[0030] (2);
[0031] (3);
[0032] (4);
[0033] Among them, represents the number of transmissions required to transmit a data packet, i.e., the expected number of transmissions, where and are the success rates of forward and reverse transmissions respectively. The lower the ETX value, the better the link quality and the fewer the number of retransmissions required for data transmission. Therefore, ETX is an important indicator reflecting the transmission stability of the time slot; is the node's periodic monitoring of the number of data packets in its transmission queue, is the number of packets currently waiting to be transmitted, i.e., the queue length. The queue length is the result of direct counting, indicating the current load; The node statistics the packet retransmission rate within the same time slot. represents the number of retransmissions, i.e., the number of times the packet is resent in order to successfully transmit the data packet within the same time slot. Represents the total number of transmission attempts, including all initial transmissions and all retransmissions. Represents the distance between a certain time slot and the receiving time slot, that is, the receive-forward distance. The closer the distance, the faster the time slot can be used for data transmission, reducing waiting time and delay. Therefore, SDF is an important parameter for optimizing delay. In the multi-weight quality evaluation model, the weight: for delay-sensitive scenarios, increase the weight and preferentially select time slots with good time slot adjacency for transmission, thereby reducing transmission delay. Represents the position of the time slot in the candidate time slot pool. Represents the position of the time slot when the node signal is received. Represents the length of the time slot frame.
[0034] For the data obtained above, perform a second data processing to generate the data input for the function in the multi-weight quality evaluation model. Use the arithmetic mean of the data collected by the sliding window to obtain the average value of the expected number of transmissions. , smooth the fluctuations of the transmission success rate, and avoid being misled by short-term link fluctuations; use the median of the data collected by the sliding window to obtain the median value of the queue length. , avoid affecting the model evaluation due to occasional load peaks; use the truncated mean of the data collected by the sliding window to obtain the truncated average value of the retransmission rate. , avoid the influence of extreme retransmission times on the evaluation of the multi-weight quality evaluation model, and enhance the stability of the multi-weight quality evaluation model. It is expressed as:
[0035] (5);
[0036] (6);
[0037] (7);
[0038] Among them, is the size of the sliding window (such as the last 10 cycles). indicates that the retransmission rate is data after truncation processing, removing extreme values (such as abnormally large or small values) in the window. is the th transmission value, QL i is the th transmission time of the queue length. () is the median function, indicating taking the median of a set of numerical values. is the number of remaining data points after truncation (data after removing the highest and lowest 10%). is the th transmission retransmission rate;
[0039] Take the processed data above as the input of the multi-weight quality evaluation model, and calculate the transmission quality of the time slots in the candidate pool, expressed as:
[0040] (8);
[0041] Among them, is the transmission quality of the time slot; are the weights of the average expected transmission times, the median value of the queue length, the truncated average retransmission rate, and the receive-forward distance respectively. The Q-learning algorithm of Pareto optimization and adaptive exploration is used for dynamic adjustment of the weights to balance the influence of different metrics;
[0042] In summary, regarding the internal node, the transmission quality of the relevant time slots can be obtained through a processing flow with less computational effort, and further combined with this, the 6P protocol is used to dynamically manage the time slot resources of the node, including dynamically releasing and supplementing time slots, maintaining and optimizing the time slot pool, so as to improve the network transmission efficiency. Through the pseudo-random channel selection mechanism, a low-interference channel is selected and combined with the previously output time slots, so as to ensure the stability and reliability of the transmission. Further, in step S1A, time slots with higher transmission quality have been selected. For the selection of channels, the pseudo-random channel selection mechanism is used to select low-interference and low-conflict channels to form a time slot-channel combination for task transmission.
[0043] Step S2A: Based on the transmission quality of the time slots output by the multi-weight quality evaluation model, the 6P protocol is used to dynamically manage the time slot resources of the node, including dynamically releasing and supplementing time slots, maintaining and optimizing the time slot pool, selecting the most suitable channel combination through the pseudo-random channel selection mechanism, and the node uses the selected channel for data transmission and updates the scheduling table.
[0044] Specifically, the selection of the most suitable channel combination through the pseudo-random channel selection mechanism is specifically: by combining the setting of the channel record table and the active channel set, all channels in the network are divided into available channels and the active channel set. Among them, the active channel set has the channels that have transmitted several time slots recorded in the channel record table and their adjacent channels before and after, and the available channels are the remaining channels after the division; for multiple consecutive channel blocks obtained by the channel division, the largest channel block is selected from them; then the linear congruence formula is applied to generate a random number, and the channel with the lowest probability of conflict and interference is selected from the largest channel block.
[0045] More specifically, step S2A includes:
[0046] Step S2A1: First, there is a group of available channels in the system that form a channel candidate pool, denoted as , respectively represent the 1st, 2nd, …, nth channels. For example, the 16 available channels in the network are Each node maintains a channel record table , which has a small storage space and is only used to record the channels used by the node in the past several transmissions. At time, based on the historical channel records in the channel record table , the system sets an active channel set , including the recently used channels and their adjacent channels, to avoid these channels being reused in a short period of time. The formula for setting the insulation area is:
[0047] (9);
[0048] Among them, represents the (r - 1)-th channel, represents the r-th channel, represents the (r + 1)-th channel;
[0049] The setting of the active channel set aims to reduce the interference problems caused by the frequent or adjacent use of channels. Next, the system generates a new available channel candidate pool , which is obtained by excluding the channels in the active channel set from the channel candidate pool:
[0050] (10);
[0051] Then, the available channel candidate pool is divided into several channel blocks according to the channel continuity , the channel block number i = {1, 2,..., m}, m is the number of channel blocks, where each channel block contains several consecutive channels. All channel blocks can be expressed as:
[0052] (11);
[0053] Among these channel blocks, the system selects the largest channel block containing the most channels , and the purpose of selecting the largest channel block is to ensure that the channels are as evenly distributed as possible within the available channel range and reduce the possibility of interference:
[0054] (12);
[0055] Step S2A2: To select a specific channel in the largest channel block, a linear congruence formula is used, and random numbers are generated based on the seed value to determine the channel. The seed value is generated by the function , where the function takes the current time slot number as the input to ensure the randomness and dynamics of channel selection:
[0056] (13);
[0057] (14);
[0058] Where the time slot number: the unique number of the current time slot, used to provide the time series information for each transmission, ensuring that each time slot has a unique seed value; For calculating the output random number, For the previously generated random number, the random number initialized here ; For the multiplier, ensuring the distribution characteristics of the random number, c is the increment, M is the modulus. The hash function is used to convert the input (time slot number) into a value with better randomness. Using the hash function to generate the seed value can ensure the randomness of the seed value, while avoiding the periodicity and repeatability of the pseudo-random numbers that may be caused by directly using the number. The random number generated by the linear congruence formula Is used for indexing in the maximum channel block, and the finally selected channel is:
[0059] (15);
[0060] (16);
[0061] Where k is the channel index, Is the selected channel, Represents the th channel of the maximum channel block.
[0062] Step S2A3: After selecting the channel , the node will use this channel to transmit data in the current time slot t, and after the transmission is completed, the channel information used will be stored in the channel recorder for reference when selecting channels later. And through the 6P protocol, the node will return the relevant data to the parent node, update the scheduling table, thereby further dynamically releasing and supplementing time slots, maintaining the dynamic update of the time slot candidate pool, and continuously optimizing the time slot selection strategy according to the new transmission data. The update of the channel recorder follows the following rules:
[0063] (17);
[0064] Where, Is the updated channel record table, Is the original channel.
[0065] The pseudo-random channel selection mechanism reduces channel usage conflicts and channel interference by avoiding scheduling the recently used channels as much as possible. It combines the setting of the active channel set and the linear congruence formula, and adopts the strategy of dividing channels into blocks. First, there is a set of available channels in the system, which constitutes the channel candidate pool. For example, there are 16 available channels in the network. Further, according to the historical channel records, the system will set an active channel set, including the recently used channels and their adjacent channels, to avoid these channels being reused in a short time. Then, the available channel candidate pool is divided into several channel blocks according to the channel continuity, and the linear congruence formula is used to randomly select channels in the largest channel block, which greatly reduces the probability of conflicts and interference generated by each channel selection, so as to ensure efficient, safe and anti-interference time slot channel allocation in a complex industrial wireless network environment.
[0066] Furthermore, the Q-learning algorithm with Pareto optimization and adaptive exploration is used to dynamically adjust the weights, including the following steps:
[0067] Step S1B: Collect the real-time state data of the network in the edge routing node, including information such as packet delivery ratio (PDR), traffic load (TL), delay, etc.; these real-time state data are used as the input of the Q-learning algorithm, initialize the key parameters of Q-learning according to the current network state, and construct the Q-table structure. The Q-table is used to store the expected rewards (Q-values) of each state-action pair, and these Q-values will be used as the decision basis to guide the routing node how to allocate resources and adjust the network strategy. In addition, a reward function needs to be designed, which will be used to quantify the performance of the network and support multi-objective optimization. By introducing Pareto optimization, determine the preference weight vector of each objective in the reward function to represent the priorities of different optimization objectives, ensure that the priorities of different objectives are reasonably considered, and achieve multi-objective optimization of the reward function.
[0068] Step S2B: Update the values in the Q-table according to the output of the reward function in step S1B. The Q-value reflects the expected reward of selecting a certain action in a specific network state. After updating the Q-value, the action selection control for each iteration adopts an adaptive exploration strategy. By monitoring the state fluctuations such as network delay, packet delivery ratio and traffic load, balance between exploration and exploitation, and dynamically adjust the exploration rate . The greater the network fluctuation, the greater the exploration rate , and the higher the exploration frequency; when the network is stable, reduce the exploration rate , make more use of the existing optimal strategy, and conduct a convergence test on the update iteration of the Q-value to monitor whether the change of the Q-value between two consecutive iterations is less than a given threshold.
[0069] More specifically, step S1B includes:
[0070] Step S1B1: Initialize the Q-table for storing the Q-values of each state-action combination, and set the learning rate α, discount factor γ, exploration rate and other parameters and initialize the Q-table.
[0071] Step S1B2: Collect the network state according to the edge router and preprocess the data to provide the basic input for Q-learning. Among them, the packet transmission rate is the core index for evaluating the network link quality, which is calculated by monitoring the number of successfully sent and received packets; the traffic load is used to reflect the utilization of the buffer (queue) of the node, indicating the load status of the network. By monitoring the number of packets in the node queue, the system can judge the network congestion situation; the delay reflects the time from sending to receiving the data, which is the key index for evaluating the network transmission efficiency and is especially suitable for time-sensitive application scenarios. Each node records the timestamp when sending a packet and calculates the total transmission delay when receiving the acknowledgment packet, and uses the average delay of the past several data transmission cycles to calculate the delay performance of the time slot to ensure its effectiveness in time-sensitive applications. Expressed as:
[0072] (18);
[0073] (19);
[0074] (20);
[0075] Where: is the packet transmission rate of the j-th node, and the sliding window method is adopted, for example, the average packet transmission rate in the past sliding window size = 10 time slots; represents the traffic load of the j-th node. When the traffic load reaches or exceeds a certain threshold (such as 80% of the maximum queue capacity), the node sends a boolean alarm signal = +1 indicates that the current load is high; otherwise, send = 0, indicating that the load is normal. This scheme reduces the communication overhead of regular monitoring and only triggers an alarm when the node load is close to saturation; represents the delay of the j-th node, and the average delay of the past several data transmission cycles is used to calculate the delay performance of the time slot to ensure its effectiveness in time-sensitive applications.
[0076] According to the selected three key indicators: packet transmission rate, traffic load, and delay, a comprehensive reward function can be constructed. Further, by introducing the multi-objective optimization idea through the design of the Pareto optimization method, the reward function is dynamically adjusted to improve the adaptability to complex scenarios and the multi-objective balance ability. First, the input of the reward function includes the current network state and the action selected by the current policy The change amount of the target value, as well as a dynamic or fixed preference weight vector , which is used to represent the priorities of each optimization objective. The calculation process takes constructing a multi-dimensional reward vector as the core. For each action , calculate its multi-objective performance and generate a reward vector , where each dimension represents the optimization effect of one objective. For example, the change amount of the data packet transmission rate , the change amount of the traffic load , the change amount of the delay . The reward vectors of all actions form a solution set .
[0077] (21);
[0078] (22);
[0079] (23);
[0080] (24);
[0081] Among them, is the PDR value at the current moment, is the PDR value at the previous moment, , is the value at the current moment, is the delay at the previous moment, is the delay at the current moment, is the reward vector solution set, , , are the target rewards represented by the 1st, 2nd, and nth components respectively.
[0082] Subsequently, the solution set is screened by the Pareto optimization method, eliminating the inferior solutions dominated by other solutions and retaining the non-dominated solutions that achieve a balance among all objectives, forming the Pareto front . The judgment criterion for non-dominated solutions is that if a solution satisfies not being completely superior to other solutions in all objectives, it belongs to a non-dominated solution. Then, the final solution can be selected from the Pareto front, which can be based on strategies such as random selection, preference weight weighted selection, or ideal solution distance selection. The preference weight weighted selection strategy calculates the weighted score of each Pareto solution, and then selects the solution with the highest score as the optimal solution ; the ideal solution distance selection calculates the Euclidean distance between each Pareto solution and the ideal solution , select the solution with the minimum distance. The output result is the selected reward vector and its corresponding action , and this reward vector achieves an optimal balance among multiple objectives.
[0083] (25);
[0084] (26);
[0085] (27);
[0086] (28);
[0087] wherein, is a certain solution, is the ideal solution, is the preference weight of the index, is the Pareto solution of the index;
[0088] The selected reward vector is used to update the Q-value table after scalarization, and the scalarization method is obtained by using the scoring method based on the distance from the ideal solution:
[0089] (29);
[0090] wherein, is the reward value of the action at time t.
[0091] Step S2B is specifically as follows:
[0092] Step S2B1: According to the setting of the multi-objective optimization of the reward function in step S1B, and according to formulas (25)-(29), the obtained reward value , substitute it into the following formula to obtain the update of the Q-value:
[0093] (30);
[0094] wherein: is the state at time t under which the action is taken; is the reward value of the action at time t, represents the Q-value (expected return) of taking the action under the state . represents the next state after the execution of the action . represents the Q-value of the optimal action that may be taken under the state at time t ; is the learning rate, which is used for the trade - off between new and old information. is the discount factor, which is used to represent the influence degree of future rewards on the current action. Through the above update, the Q - value can reflect the comprehensive contribution of the current action to multi - objective optimization.
[0095] Further, regarding the action control strategy in the Q - learning iteration process, an adaptive exploration strategy is adopted. According to the network status data collected by the edge router, considering the situation of abnormal fluctuations in its appearance status, the exploration rate is dynamically set for the action. According to the fluctuations of the network status, the exploration rate is dynamically adjusted, so as to achieve an effective balance between exploration and exploitation in the Q - learning algorithm: when the network fluctuates greatly, the exploration rate is increased. When the network is relatively stable, the exploration rate is decreased, and the current strategy is used more.
[0096] (31);
[0097] (32);
[0098] Among them, is the exploration rate at the T - th iteration, which determines the trade - off between exploration and exploitation in the Q - learning algorithm; is the maximum exploration rate, is the minimum exploration rate, is the decay rate, which controls the speed of the exploration rate decrease; is the delay at time t, is the delay at time t - 1, is the packet transmission rate at time t, is the packet transmission rate at time t - 1, is the traffic load at time t, is the traffic load at time t - 1, , are the weight parameters of the delay, packet transmission rate and traffic load fluctuations respectively, which control the contribution of each network metric to the fluctuation metric. The maximum exploration rate indicates that the exploration possibility should be increased when the network status fluctuates greatly; the minimum exploration rate ensures that a certain degree of exploration is retained at any time, even when the network status is relatively stable. The function is used to measure the fluctuation of the network status at time t, and its result is used to dynamically adjust the exploration rate.
[0099] : represents the network delay at the current time t and the previous time . The change of network delay can reflect the real - time performance of the network.
[0100] : represents the packet transmission rate at the current time and the previous time, and its change reflects the reliability of the network.
[0101] Indicates the network load at the current moment and the previous moment. The fluctuation of the network load reflects the change in the amount of data transmitted by the node.
[0102] Step S2B2: For the exploration rate dynamically adjusted above, perform action selection in the action space, that is, select a weight combination, and according to the selected action Calculate the reward function to evaluate the effect of this action, and use it to update the Q value in the Q table. And gradually optimize the weights of the multi-weight quality evaluation model; by adjusting the exploration rate And perform convergence judgment to ensure that the algorithm finds the optimal weight combination during the continuous learning process . Indicates The optimal value of Indicates The optimal value of Indicates The optimal value of Indicates The optimal value of
[0103] Action space A: Define the action space A as a combination of multiple weight parameters:
[0104] (33);
[0105] When performing action selection, generate a random number . If ζ < , then perform exploration (randomly select an action combination). Otherwise, perform exploitation (select the action combination with the highest current Q value). The action selection formula for exploration and exploitation:
[0106] (34);
[0107] (35);
[0108] Among them, Is a random action, Is the optimal action, that is, the action corresponding to the maximum Q value, Is the state at time t-1 Under the condition, take the action The Q value of Is a given threshold. When the change in the Q value between two consecutive iterations is less than a given threshold , it can be considered that a certain stable state has been reached, that is, convergence. This means that the policy has tended to be stable.
[0109] Such as Figure 2As shown, in this embodiment, first, key performance parameters of each candidate time slot are collected at the internal node, including the expected number of transmissions, the transmission queue length, the packet retransmission rate, etc., to help evaluate the time slot quality subsequently. These performance parameters are the basic data for subsequent multi-weight quality evaluation and directly determine the quality of time slot transmission. Next, the collected performance parameter data is preprocessed to remove outliers and smooth the data, so as to ensure that the subsequent evaluation results are more accurate and robust, eliminate the influence of short-term network fluctuations, and improve the reliability of evaluating the time slot performance. Then, relevant statistical data processing methods are used to process the above data. For each candidate time slot, the distance between it and the receiving unit is directly calculated. The receive-forward distance is an important parameter used to evaluate the spatial distance between the time slot and the receiving unit. Generally speaking, when the distance is farther, the interference is less. In the next step, weights obtained by the Q-learning algorithm of Pareto optimization and adaptive exploration need to be assigned to each performance metric. The preprocessed performance parameters and the corresponding weights are substituted into the quality evaluation function to calculate the comprehensive quality score of each candidate time slot. This quality evaluation function is the core of the entire solution, and it uses different performance metrics to comprehensively calculate the final score of each time slot. According to this score, the quality of the time slots in the candidate pool is detected. If the quality is abnormal, such as a significant decrease or frequent changes in quality, then the time slots in the candidate pool need to be adjusted. By dynamically adjusting the candidate pool, the overall quality of the candidate time slots can be ensured to maintain at a high level. Based on the candidate pool quality detection, the time slots in the candidate pool are maintained, including replacing low-quality time slots or adding new high-quality time slots to ensure the quality of the entire candidate pool, thereby ensuring the reliability and stability of the time slots in subsequent scheduling. Finally, according to the comprehensive score calculated by the quality evaluation function, the time slot with the highest score is selected as the priority transmission time slot. This process ensures that in the current network state, the selected time slot can optimally meet the transmission requirements, with a higher transmission success rate and lower latency. For example, node A needs to select an optimal time slot from 3 candidate time slots S 1 , S 2 , S 3 for data transmission. First, node A collects the transmission performance parameters of each time slot, including the expected number of transmissions, the transmission queue length, the packet retransmission rate, and the receive-forward distance. To ensure the accuracy and robustness of the data, node A preprocesses the collected data to remove outliers and eliminate the influence of accidental fluctuations. After the preprocessing is completed, node A uses different methods to process the expected number of transmissions, the transmission queue length, and the packet retransmission rate in the performance metrics to obtain more robust data. For the expected number of transmissions (ETX), node A uses the arithmetic mean method and calculates the average value using the data of the last 5 times. For example, for time slot S 1, the last 5 data of ETX are [1.1, 1.2, 1.3, 1.2, 1.1], and the average value is 1.18. For the transmission queue length, node A processes it using the median. For example, for time slot S 1 , the last 5 queue length data are [3, 4, 3, 5, 3], and the median is 3. For the packet retransmission rate, node A uses the truncated mean method, that is, removing the highest and lowest values and then averaging the remaining values. For example, for time slot S 1 , the last 5 data of the retransmission rate are [0.05, 0.04, 0.06, 0.05, 0.04]. After removing the highest 0.06 and the lowest 0.04, the remaining data are [0.05, 0.05, 0.04], and its truncated mean is 0.047. At the same time, for the receive-forward distance (SDF), node A directly uses the logical distance between the current candidate time slot and the receiving unit. For example, for time slot S 1 , the SDF of it is 2, indicating that the logical distance between it and the RX receiving unit is 2 channels. Next, node A determines the weights for evaluating each metric. The weights of the four metrics are assumed to be: w 1 = 0.3, w 2 = 0.2, w 3 = 0.2, w 4 = 0.3, corresponding to ETX, transmission queue length, retransmission rate, and receive-forward distance respectively. These weights are used to reflect the importance of different metrics in the comprehensive quality evaluation. Node A then substitutes the preprocessed performance parameters and weights into the quality evaluation function to calculate the comprehensive quality score for each time slot. For example, for time slot S 1 , the score is:
[0110] During the evaluation process, node A also needs to detect and maintain the quality of the time slots in the candidate pool. If it is detected that the quality evaluation results of some time slots are continuously lower than a certain threshold, the system will perform candidate pool quality maintenance, such as replacing low-quality time slots and adding new high-quality time slots to maintain the overall quality of the candidate pool. Finally, node A selects the time slot with the highest score as the preferred time slot for data transmission according to the comprehensive score of each time slot. In this example, the time slot with the highest score is time slot S 3 , , so node A selects time slot S 3Data transmission is carried out. Through steps such as data collection, preprocessing, selection of different data processing methods (average value, median, truncated mean), calculation of receiving and forwarding distance, weight allocation, quality assessment, and candidate pool maintenance, it is ensured that node A can efficiently select the best time slot for data transmission in a complex industrial network environment, achieving the effects of high efficiency, reliability, and adaptability to dynamic changes.
[0111] As Figure 3 shown, in one embodiment, the entire process of the Pareto optimization and adaptive exploration Q-learning algorithm starts when the node receives a transmission task, and the scheduling mechanism is initiated through the collection of TSCH network performance parameters. These network performance parameters include network latency: reflecting the time for data packets to be transmitted in the network; data packet transmission rate: measuring the proportion of successfully transmitted data packets in the total sent data packets, reflecting the stability of the transmission link; network load: reflecting the change in the communication volume between nodes in the network. These data are collected by the edge router using a sliding window mechanism to accumulate historical data and monitor changes in the network state. The collected data is preprocessed to ensure data accuracy and stability by filtering out outliers and using methods such as truncated mean. The processed sliding window data provides a continuous performance trend for the model, ensuring that scheduling decisions are based on the actual network state. By processing these complex computing tasks on the edge side, the computing pressure on the sensor nodes is significantly reduced, saving energy consumption, and thus making the network nodes more suitable for resource-constrained industrial Internet of Things environments.
[0112] Based on the collected data, the edge router initializes the Q-table, state space, and key parameters. The Q-table is used to record the value estimates of each state-action pair to provide a basis for subsequent action selection. The initial Q-values are usually set to small random values for more exploration in the initial stage. When choosing an action, we adopt an adaptive exploration strategy to balance exploration and exploitation, dynamically adjusting the balance between exploration and exploitation in different network states to ensure the efficient use of network resources and at the same time minimize energy consumption and latency. By monitoring state fluctuations such as network latency, data packet transmission rate, and load, the ε value is dynamically adjusted. The greater the network fluctuation, the greater the ε value and the increased exploration frequency; when the network is stable, the ε value is reduced to make more use of the existing optimal strategy. Action space A: Define the action space A as a combination of multiple weight parameters: generate a random number . If ζ < , then exploration is carried out (randomly select an action combination). Otherwise, exploitation is carried out (select the action combination with the highest current Q value). The action selection of exploration and exploitation adopts formulas (31)-(34).
[0113] To optimize the scheduling strategy, we adopted Pareto optimization to dynamically adjust the weights in the reward function. Specifically, the edge router dynamically adjusts the priorities of these objectives according to the current network latency, packet transmission rate, and load. For example, it gives priority to reducing latency under high load and reduces the energy consumption priority when the node energy consumption is high, in order to achieve an optimal balance in different network environments. First, the reward function is designed as a multi-dimensional reward vector to preserve the independence of each metric. The constructed reward vector is = , , and the reward vectors of all actions form a solution set . Subsequently, the solution set is screened by the Pareto optimization method to eliminate the inferior solutions dominated by other solutions and retain the non-dominated solutions that achieve a balance among all objectives, forming the Pareto front . The criterion for judging non-dominated solutions is that if a solution satisfies not being completely superior to other solutions in all objectives, it belongs to the non-dominated solution. Then, the final solution is selected from the Pareto front, which can be based on strategies such as random selection, weighted selection with preference weights, or distance selection from the ideal solution. The weighted selection strategy calculates the weighted score of each Pareto solution: , and then selects the solution with the highest score as the optimal solution: ; The distance selection from the ideal solution selects the solution with the smallest Euclidean distance from each Pareto solution to the ideal solution . The output result is the selected reward vector and its corresponding action , and this reward vector achieves an optimal balance among multiple objectives. The selected reward is used to update the Q-value table after scalarization, and the scalarization method uses a scoring method based on the distance from the ideal solution. Finally, the Q-value update follows the classic Q-learning formula, that is, formula (30). Through the above update, the Q-value can reflect the comprehensive contribution of the current action to multi-objective optimization.
[0114] The key to the entire process lies in adopting the idea of edge computing, concentrating complex network state analysis and optimization calculations on the edge router for processing, while sensor nodes only need to perform simple local calculations, thus effectively reducing the computational burden and energy consumption of the nodes. Through the multi-weight quality evaluation model driven by the edge router and the improved Q-learning algorithm, changes in network performance (such as latency, packet transmission rate, and load) can be monitored in real time and used for dynamic optimization of the selection of time slot-channel combinations and the maintenance of the candidate pool. The edge router makes optimal scheduling decisions based on this comprehensive information and then distributes the results to each node for execution. The sliding window mechanism is used to collect and preprocess network state data in real time, enabling the edge router to efficiently evaluate the network condition and make intelligent decisions. Sensor nodes only need to operate according to the optimal strategy provided by the edge router, ensuring that the system can maintain low latency, high reliability, and low energy consumption in a dynamic environment. For channel allocation after time slot selection, this scheme adopts the method of randomly selecting channels to maintain the security and anti-interference ability of the TSCH network. In this way, the channel selection is not fixed, effectively preventing channel eavesdropping and attacks, further improving the security and robustness of the network. At the same time, through the 6P protocol to dynamically manage time slot resources, the edge router continuously optimizes the maintenance of the candidate pool. Combining with the multi-weight quality evaluation model, the system can adaptively improve the transmission performance and achieve efficient and flexible resource scheduling.
[0115] Figure 4 Shows how the time slot offset and channel offset affect the selection of the final channel.
[0116] Currently, node A records the channels it has recently used at time t as . These channels have been recently used in data transmission and will not be used for this selection to reduce possible interference. Based on the historical records in the recorder, node A takes the recently used channels and adjacent channels as the recently used area: The recently used area includes channels . Exclude the channels in the recently used area from the channel candidate pool to obtain the remaining available channel candidate pool:
[0117] ;
[0118] Further divide the available channel candidate pool into several channel blocks according to continuity:
[0119] ;
[0120] Compare the sizes of each channel block and select the channel block with the largest number of channels: The largest channel block contains 10 channels and is the block with the largest number of channels. Next, use the current time slot number (assuming the time slot number is (generate a seed value). Assume the hash function converts the time slot number 42 into a seed value .
[0121] Select LCG parameters:
[0122] a = 1664525, c = 1013904223, m = 2 32 = 4294967296;
[0123] Use the LCG formula to generate the next pseudo - random number:
[0124] ;
[0125] Now we need to select a channel in the maximum channel block. The maximum channel block contains 10 channels: Take the modulus of the generated random number with the size of the channel block to get the index of the channel: . According to the index , select the channel in the maximum channel block: . So, the finally selected channel is . Node A uses the selected channel for data transmission. After the transmission is completed, update the channel recorder and record the latest used channel in it:
[0126] ;
[0127] The present invention proposes an industrial wireless network resource scheduling method based on a multi-weight quality evaluation model, effectively solving the problems of randomness in time slot selection and lack of transmission quality evaluation in existing resource scheduling, while ensuring high reliability of data transmission of wireless sensor nodes and low end-to-end latency in a complex and changeable industrial network environment. First, each sensor node collects the performance data of time slots in the candidate pool through a sliding window mechanism and performs reasonable preprocessing to ensure the accuracy and robustness of the data input into the multi-weight quality evaluation model. To relieve the computing pressure on the nodes, the edge router undertakes the complex quality evaluation calculation tasks. The edge router uses the multi-weight quality evaluation model to comprehensively score each time slot and generates dynamic weight parameters applicable to the current network environment. This method not only improves the accuracy of evaluation, but also significantly saves the computing resources and energy consumption of the internal nodes of the network, and extends the working life of the nodes. Then, the edge router feeds back the evaluation results and weight parameters to each node, and the node combines the 6P protocol to dynamically manage the time slot pool, including releasing and supplementing time slots, to ensure the optimal utilization of resources. Subsequently, the node selects the most suitable time slot for data transmission according to the output of the multi-weight quality evaluation model, and combines the TSCH channel hopping mechanism to select the optimal channel, further improving the stability and reliability of network transmission. By using edge computing to centralize complex computing tasks on the edge router, the present invention effectively reduces the computing burden and energy consumption loss of sensor nodes, while ensuring the efficiency and reliability of transmission tasks in a complex network environment, meeting the requirements for the efficiency and stability of data transmission in the industrial Internet of Things environment.
[0128] The present invention has been described in detail with reference to the embodiments accompanied by drawings. Those of ordinary skill in the art can make various variations to the present invention according to the above description. Therefore, certain details in the embodiments should not constitute a limitation to the present invention, and the present invention will take the scope defined by the appended claims as the protection scope.
Claims
1. An industrial wireless network resource scheduling method based on a multi-weight quality assessment model, characterized in that: include: Step S1A: The internal nodes of the network apply a sliding window mechanism to collect relevant data of the recent network status, including the expected number of transmissions, queue length, packet retransmission rate, and receiving forwarding distance, and perform data preprocessing to obtain the average value of the expected number of transmissions, the median value of the queue length, the truncated average value of the retransmission rate, and the receiving forwarding distance, and use them as inputs of the multi-weight quality assessment model to calculate the transmission quality of the time slots in the candidate pool; the multi-weight quality assessment model is expressed as: ; in, is the transmission quality of the time slot, The average expected number of transmissions , median queue length , retransmission rate truncated mean And receiving forwarding distance The weights are adjusted dynamically using the Pareto optimization and adaptive exploration Q learning algorithm; Step S2A: Based on the transmission quality of the time slot output by the multi-weight quality assessment model, the 6P protocol is combined to dynamically manage the time slot resources of the node, including dynamically releasing and replenishing time slots, maintaining and optimizing the time slot pool, and selecting the most suitable channel combination through a pseudo-random channel selection mechanism. The node uses the selected channel for data transmission and updates the scheduling table; The most suitable channel combination is selected by the pseudo-random channel selection mechanism, specifically: by combining the channel record table and the setting of the active channel set, all channels in the network are divided into available channels and active channel sets, wherein the active channel set is the channel that has been transmitted in several time slots recorded in the channel record table and its adjacent channels before and after, and the available channel is the remaining channel after the division; for multiple continuous channel blocks divided by the channel, the largest channel block is selected; and then a linear congruential formula is applied to generate random numbers, and a channel with the smallest probability of conflict and interference is selected from the largest channel block.
2. The industrial wireless network resource scheduling method based on a multi-weight quality assessment model as claimed in claim 1, characterized in that: The node uses the selected channel for data transmission and updates the scheduling table, including: the node uses the selected channel for data transmission, and after the transmission is completed, stores the used channel information in the channel recorder for reference during subsequent channel selection; and through the 6P protocol, the node transmits the relevant data back to the parent node, updates the scheduling table, thereby further dynamically releasing and supplementing time slots, maintaining the dynamic update of the time slot candidate pool, and continuously optimizing the time slot selection strategy according to the new transmission data.
3. The industrial wireless network resource scheduling method based on a multi-weight quality assessment model according to claim 1, characterized in that: The Q learning algorithm using Pareto optimization and adaptive exploration is used to dynamically adjust the weights, including the following steps: Step S1B: The edge routing node collects real-time status data of the network, including packet transmission rate PDR, traffic load TL, and delay; the real-time status data is used as the input of the Q learning algorithm, and the key parameters of Q learning are initialized according to the current network state, and a Q table is constructed; the Q table is used to store the Q value of each state-action pair, and the Q value will be used as a decision-making basis to guide the routing node on how to allocate resources and adjust the network strategy; a reward function is designed to quantify the performance of the network and support multi-objective optimization; by introducing Pareto optimization, the preference weight vector of each target in the reward function is determined to represent the priority of different optimization targets, to ensure that the priority of different targets is reasonably considered, and to achieve multi-objective optimization of the reward function; Step S2B: Update the value in the Q table according to the output of the reward function in step S1B; adopt an adaptive exploration strategy for the action selection control of each iteration after updating the Q value, balance between exploration and utilization by monitoring the network delay, packet transmission rate and traffic load fluctuations, dynamically adjust the exploration rate, and perform convergence tests on the updated iterations of the Q value to monitor whether the change in the Q value between two consecutive iterations is less than a given threshold.
4. The industrial wireless network resource scheduling method based on a multi-weight quality assessment model according to claim 1, characterized in that: Step S2A includes: Step S2A1: First, there is a set of available channels in the system to form a channel candidate pool, denoted as , Represents the 1st, 2nd, ..., nth channels respectively. Each node maintains a channel record table ,exist At this time, according to the channel record table The system will set an active channel set for the historical channel records in , including the most recently used channel and its adjacent channels; the formula for setting the isolation zone is: (9); in, represents the r-1th channel, represents the rth channel, represents the r+1th channel; Next, the system generates a new pool of available channel candidates , obtained by excluding the active channel set channels from the channel candidate pool: (10); Then, the available channel candidate pool Divide into several channel blocks according to the continuity of the channel , channel block number i={1,2,…,m}, m is the number of channel blocks, where each channel block Contains several continuous channels; all channel blocks are represented as: (11); The system selects the largest channel block containing the most channels : (12); Step S2A2: In order to select a specific channel in the largest channel block, a linear congruential formula is used and based on the seed value Generate a random number to determine the channel; seed value By function Generate, where the function Taking the current time slot number as input: (13); (14); in, To calculate the output random number, The random number generated last time, the random number initialized here ; is the multiplier, c is the increment, and M is the modulus; the hash function is used to convert the time slot number into a value with good randomness; the random number generated by the linear congruential formula Used to index in the largest channel block, the final selected channel is: (15); (16); Where k is the channel index, For the selected channel, The largest channel block channels; Step S2A3: After the channel is selected After that, the node will use the channel for data transmission in the current time slot t, and after the transmission is completed, the channel information used will be stored in the channel recorder for reference in subsequent channel selection; and through the 6P protocol, the node will return the relevant data to the parent node and update the scheduling table, so as to further dynamically release and supplement time slots, maintain the dynamic update of the time slot candidate pool, and continuously optimize the time slot selection strategy according to the new transmission data; the update of the channel recorder follows the following rules: (17); in, is the updated channel record table, The original channel.
5. The industrial wireless network resource scheduling method based on a multi-weight quality assessment model as described in claim 3 is characterized in that: Step S1B includes: Step S1B1: Initialize the Q table to store the Q value of each state and action combination, set the learning rate α, discount factor γ, exploration rate ϵ and other parameters and initialize the Q table; Step S1B2: Collect network status and pre-process data based on edge routers to provide basic input for Q learning; including packet transmission rate, traffic load and delay; construct a comprehensive reward function based on packet transmission rate, traffic load and delay; introduce multi-objective optimization ideas by combining the design of Pareto optimization method, and dynamically adjust the reward function.
6. The industrial wireless network resource scheduling method based on a multi-weight quality assessment model as claimed in claim 5, characterized in that: The step S1B2 comprises: First, the input of the reward function includes the current network state , the action selected by the current strategy , the change in the target value, and the dynamic or fixed preference weight vector , used to indicate the priority of each optimization goal; the calculation process is centered on constructing a multidimensional reward vector. , calculate its multi-objective performance and generate a reward vector , where each dimension represents the optimization effect of an objective; the reward vectors of all actions constitute the solution set ; Subsequently, the solution set is screened by the Pareto optimization method, the inferior solutions dominated by other solutions are eliminated, and the non-dominated solutions that achieve a balance between all objectives are retained to form the Pareto frontier. ; Next, the final solution is selected from the Pareto frontier based on random selection, preference weighted selection, or ideal solution distance selection strategy; the preference weighted selection strategy calculates the weighted score of each Pareto solution Then choose the solution with the highest score as the optimal solution ; The ideal solution distance is selected by calculating the distance between each Pareto solution and the ideal solution Euclidean distance , select the solution with the smallest distance; the output result is the selected reward vector and its corresponding actions , the reward vector achieves the optimal balance among multiple objectives; the selected reward vector is used to update the Q value table after scalarization.
7. The industrial wireless network resource scheduling method based on a multi-weight quality assessment model as claimed in claim 6, characterized in that: The step S2B is specifically as follows: Step S2B1: According to the setting of multi-objective optimization of reward function in step S1B, the reward value is obtained, and the reward value is used to update the Q value; the action control strategy of the Q learning iteration process adopts an adaptive exploration strategy, and according to the network status data collected by the edge router, the exploration rate is dynamically set in consideration of the abnormal fluctuation of the status; Step S2B2: For the dynamically adjusted exploration rate, select actions in the action space, that is, select weight combinations according to the selected actions. Calculate the reward function to evaluate the effect of the action, which is used to update the Q value in the Q table; and gradually optimize the weights of the multi-weight quality assessment model; by adjusting the exploration rate ϵ and making convergence judgments, ensure that the Q learning algorithm finds the optimal weight combination in the process of continuous learning ; express The optimal value of express The optimal value of express The optimal value of express The optimal value of .
Citation Information
Patent Citations
Static link scheduling method aiming at 6TiSCH multi-hop wireless network
CN107277884A
Dynamic spectrum pool construction system and method of industrial wireless network
CN108882246A