Protocol adaptation parameter determination method based on deep reinforcement learning
Through the weighted dual-delay DDPG network based on deep reinforcement learning, the problems of low efficiency and high cost when devices access the network are solved, fast and efficient protocol adaptation is achieved, and the efficiency of device access to the network and the protocol matching rate are improved.
Patent Information
- Application Number
- CN202510947508.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-10
AI Technical Summary
In scenarios where device protocol information is unknown, existing technologies have low efficiency and high cost when devices access the network. In addition, the existing plug-and-play method cannot adapt to protocol diversity scenarios, and the protocol matching process requires multiple feedbacks, which reduces the adaptation efficiency of device protocols.
A weighted double-delay DDPG network based on deep reinforcement learning is adopted. By building a device information library and protocol library, and utilizing the Actor-Critic framework, similarity weights are output, similar device groups are constructed, and protocol groups are verified to improve protocol adaptation efficiency.
It achieves fast and efficient device access to the network, improves protocol adaptation efficiency and matching rate, reduces the cost of manual development, and expands the diversity of small sample training and the smoothness of network training.
Smart Images

Figure CN120768962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet of Things, and in particular to a method for determining protocol adaptation parameters based on deep reinforcement learning. Background Art
[0002] Devices are a crucial source of IoT data. With the deployment of massive heterogeneous devices, the number of connected devices is exploding. Faced with a vast array of diverse protocols, and in scenarios where device protocol information is unknown, efficiently generating the best-matching protocol for each device and rapidly connecting them to the network is crucial to the development of the IoT.
[0003] Existing dynamic device access methods rely on professionals to manually develop and select device protocols, which can be inefficient and costly when the number of protocols is large. Existing plug-and-play device access methods, relying on standardized protocols, are not adaptable to scenarios with diverse protocols. While existing device protocol generation methods can predict protocol types, they require multiple feedback cycles to find the optimal weight during protocol matching, reducing the efficiency of device protocol adaptation. Summary of the Invention
[0004] In view of the above-mentioned technical defects, the present invention provides a method for determining protocol adaptation parameters based on deep reinforcement learning. The method of the present invention can deterministically output the similarity weight used to calculate the similarity of basic information based on the basic information of the device, avoiding multiple similarity optimization processes, improving the protocol adaptation efficiency, and accelerating the speed of device access to the network.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] A method for determining protocol adaptation parameters based on deep reinforcement learning, comprising the following steps:
[0007] S1. Build a device information database; device information includes basic device information and protocol identification;
[0008] S2. Constructing a weighted double-delay DDPG network including an Actor network and a Critic network; the weighted double-delay DDPG network outputs a similarity weight for determining the similarity of basic information according to the device information;
[0009] The interaction between the weighted double-delay DDPG network and the environment includes protocol group construction; each device protocol generation task in the protocol group construction is an agent, and the agent performs actions based on rewards and states;
[0010] The status includes the basic information of the device and the similarity weight;
[0011] The actions include increasing the similarity weight and decreasing the similarity weight;
[0012] The Actor network includes an online strategy network and a target strategy network; the online strategy network determines the action according to the current state, and the target strategy network determines the next action according to the next state;
[0013] S3. Constructing a similar device group based on the similarity weights;
[0014] S4. Construct a protocol group based on the similar device group; verify the protocol text in the protocol group to obtain an adaptation protocol.
[0015] Preferably, the Critic network includes an online Q network and a target Q network; the online Q network and the target Q network respectively output their respective minimum values; the two minimum values are weightedly combined as the target Q value; the weighted double-delay DDPG network outputs the similarity weight according to the device information and the target Q value.
[0016] Preferably, the Actor network updates parameters of the online policy network and the target policy network according to the target Q value.
[0017] Preferably, the reward is the sum of the protocol matching rate R and the hit rate H; the protocol matching rate R is defined as:
[0018]
[0019] The device protocol is defined as P = [P id ,b],P id represents the protocol identifier, b represents the basic data of the device; * represents the number of correct elements in the generated b, and λ represents the number of real elements in b; the hit rate H is defined as:
[0020]
[0021] Where σ represents the number of protocol verifications during the protocol adaptation process.
[0022] Preferably, a gradient descent algorithm is used to minimize the error, and the parameters of the online Q network are calculated and updated; after the online Q network is updated, the target Q value evaluated by it is calculated; a gradient ascent algorithm is used to maximize the target Q value, and the parameters of the online policy network are calculated and generated.
[0023] Preferably, during the training of the weighted double-delay DDPG network, random noise is added to the action.
[0024] As preferred, the weighted double-delay DDPG network further comprises a memory replay mechanism and a mini-batch sample; in the memory replay mechanism, experience tuples (s t ,a t ,r t ,s t+1 ) of each interaction are stored in a dataset; wherein s t is a current state, a t is a current action, s t+1 is a next state, and a t+1 is a next action; the mini-batch sample is formed by randomly extracting a subset from the dataset, for gradient updating.
[0025] Compared with the prior art, the beneficial effects of the present application are embodied in:
[0026] 1. When the device protocol is missing, the traditional way is to develop the protocol manually, which is low in efficiency and high in cost. Different from the traditional technical method, the present application dynamically generates the best adaptive protocol by using the deep reinforcement learning method, thereby improving the efficiency of device access.
[0027] 2. In the protocol adaptation, the protocol adaptation parameters of the traditional technology are manually set. Different from the traditional technical method, the present application determines the protocol adaptation parameters by using the deep reinforcement learning method, thereby improving the parameter determination efficiency and shortening the protocol adaptation time.
[0028] 3. Different from the traditional deep reinforcement learning method based on DDPG, the deep reinforcement learning algorithm used in the present application is a weighted double-delay deep deterministic policy gradient algorithm, which outputs the minimum value of two Q networks to the loss function in a weighted form, thereby improving the smoothness of network training and expanding the diversity of small sample training.
[0029] 4. The traditional protocol verification only focuses on the protocol matching rate, resulting in a decrease in the protocol hit rate. Different from the traditional technical method, the present application combines the hit rate and the matching rate of the protocol during the protocol verification, thereby improving the overall performance of the protocol adaptation in terms of efficiency and matching. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 is a method framework diagram of embodiment 1 of the present application;
[0031] Figure 2 is a device protocol adaptation parameter optimization network diagram of embodiment 1 of the present application. DETAILED DESCRIPTION
[0032] In order to make the technical means, creative features, purposes and effects of the application easy to understand, the present application will be further described in conjunction with specific drawings. However, the present application is not limited to the following embodiments.
[0033] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them. They are not used to limit the conditions under which the present invention can be implemented. Therefore, they have no substantive technical significance. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention.
[0034] Aiming at the shortcomings of existing methods, this paper proposes a method for determining protocol adaptation parameters based on deep reinforcement learning. The framework of this method is as follows Figure 1 As shown in the figure, the device information library stores information about a large number of devices. Each device information includes basic device information and supported protocol identifiers. The protocol library stores protocol text corresponding to the protocol identifiers supported by the devices in the device information library. After a device is connected to an edge computing node via a digital sensor network (such as RS485, CAN bus, etc.), the edge computing node obtains the device's basic information and sends it to the protocol adapter. After receiving the device's basic information, the deep reinforcement learning (DRL) agent in the protocol adapter outputs a deterministic weight for the device. Using this weight, the protocol adapter calculates the similarity between the basic information of the connected device and all devices in the device information library, and uses the top 20% of devices with the highest similarity to form a similar device group. Based on the protocol identifiers supported by the devices in the similar device group, the protocol text corresponding to the protocol identifier is extracted from the protocol library to form a protocol group. The protocol group is then sent to the edge computing node, which verifies each protocol in the protocol group and obtains the most suitable protocol.
[0035] During the protocol adaptation process, the similarity weight required to calculate the basic information similarity directly affects the construction of the similar device group, and thus affects the protocol adaptation efficiency and protocol matching rate. In order to improve the adaptation efficiency and protocol matching rate of the device protocol, the present invention proposes a protocol adaptation parameter determination method based on deep reinforcement learning. The state space of the weighted similarity weight is a continuous state, and the action space is the adjustment of the weighted state value, both of which are continuous. The Deep Deterministic Policy Gradient (DDPG) algorithm can handle the problems of continuous state space and action space, but there is a problem of overestimation of Q value. In order to overcome the problems of overestimation of Q value and continuous state and action space, the present invention designs an improved deep reinforcement learning method based on DDPG-weighted double-delay DDPG algorithm, which is used to determine the similarity weight required for basic information similarity calculation. After each time the best matching protocol is obtained, the protocol adapter performs incremental training on the DRL agent to gradually optimize the similarity weight output.
[0036] Example 1:
[0037] A method for determining protocol adaptation parameters based on deep reinforcement learning, comprising the following steps:
[0038] Step S1: Build a device information model to lay the foundation for protocol adaptation.
[0039] The device model in the device information library is defined as d = [i, P], where i represents the basic information of the device and P represents the device protocol.
[0040] The basic information of the device is defined as in Represents a set of positive integers, and β is the length of the basic information.
[0041] The device protocol is defined as P = [P id ,b], where P id Represents the protocol identifier, and b represents the basic data of the device. The basic data of the device is defined as λ represents the amount of basic data of the device. Basic information includes manufacturer, type, model, etc.
[0042] Device d a With device d b The basic information similarity is defined as:
[0043]
[0044] Where s j Indicates i j where k is the number of basic information elements involved in the similarity calculation, and w is the similarity weight used in the similarity calculation. As can be seen from the definition of similarity, the similarity weight directly determines the similarity value between devices, which in turn affects the construction of similar device groups.
[0045] Step S2: construct a weighted double-delay DDPG network and output similarity weights based on device information.
[0046] Step S21: Weighted double-delay DDPG network design
[0047] During training, the traditional DDPG approach uses the minimum value of the online network as the output of the Q network. Because DDPG requires a large number of training samples, it is inefficient when the number of samples is small. To address this, a weighted double-delayed DDPG approach was proposed. During network training, the weighted sum of the minimum values of the online network and the target network is used as the output of the Q network. By appropriately adjusting the weighting, the stability and performance of network training are improved.
[0048] Weighted double-delay DDPG is based on the Actor-Critic framework and consists of four parts: Actor network, Critic network, memory replay, and mini-batch samples, as shown in Figure 2
[0049] In the weighted double-delay DDPG, the interaction with the environment contains two operations: similar device group construction, protocol group construction, and protocol verification. During the training process, each device protocol generation task is regarded as an agent, which determines the action (a t+1 ) according to the reward (r t ) and the next state (s * ) obtained from the environment.
[0050] State space. The state space is defined as where Here, i represents the device basic information element, and w represents the weight value for similarity calculation. The value of i is an integer, and the value of w is a continuous floating-point data.
[0051] Action space. The action space is defined as where I represents increasing the similarity weight, and D represents decreasing the similarity weight. Since the basic information of the access device is fixed, it is not operated on.
[0052] Reward r. Since the goal of protocol generation is to maximize the matching rate and hit rate, the reward is defined as the sum of the protocol matching rate R and the hit rate H. The protocol matching rate R is defined as:
[0053]
[0054] In the formula, λ * represents the number of correct elements in the generated b, and λ represents the number of real elements in b. The hit rate H is defined as:
[0055]
[0056] In the formula, σ represents the number of protocol verification times in the protocol adaptation process.
[0057] Actor network, Critic network, memory replay, and mini-batch sample descriptions are as follows:
[0058] Actor network. The Actor network consists of an online policy network and a target policy network. The online policy network determines the action (a t ) according to the current state (s t ), and the target policy network determines the next action (a t+1 ) according to the next state (s t+1). The policy gradient update of the Actor network parameters is provided by the Q value of the Critic network. The parameters of the online policy network and the target policy network are θ P and
[0059] Critic network. Critic network includes two online Q networks and two target Q networks. Online Q network estimation (s t , a t ) Q value, target Q network estimation (s t+1 , a t+1 ) The parameters of the online Q network and the target Q network are and
[0060] Memory replay. In order to break the correlation between data and improve learning efficiency, the experience tuple (s t ,a t ,r t ,s t+1 ) are stored in the dataset and learned by random sampling.
[0061] Mini-batch samples. A mini-batch is formed by randomly sampling a subset from the memory replay for gradient updates.
[0062] Step S22: Weighted double-delay DDPG network training and parameter update
[0063] In weighted double-delay DDPG, the Actor network generates actions using a deterministic policy. The online policy network generates actions according to the policy function μ(s t θ P ) Generate current action a t , the target policy network is based on the policy function Generate the next action a t+1 During the training process, random noise is added to the generated actions in order to smooth the value evaluation and improve the evaluation accuracy. t+1 for:
[0064]
[0065] Where, is random noise, a min is the minimum action value, a max is the maximum action value.
[0066] In the critic network, according to the theory of Deep Q-network (DQN), the outputs of the online Q network and the target Q network are defined as:
[0067]
[0068] In the formula, γ is the discount factor. In order to eliminate the overestimation problem of Q value, the online Q network and the target Q network select their respective minimum values as the output values. The minimum values of the two are defined as:
[0069]
[0070] In order to increase the sample diversity of training data, the present invention designs a weighting mechanism to combine Q and TQ weightedly as the Q value, so the Q value is defined as:
[0071] y=αQ+(1-α)TQ
[0072] Where α is the combination coefficient.
[0073] According to y, the error value (Temporal-Difference, TD) of the Critic network is defined as:
[0074]
[0075] Use the gradient descent algorithm to minimize the error, calculate and update the parameters of the Critic network The loss function of the online Q network is defined as:
[0076]
[0077] Where E represents the expected value of the TD error. Using the loss function, the update of the online Q network parameters can be expressed as:
[0078]
[0079] Where, α Q for Learning rate in the critic network.
[0080] After updating the online Q network, the parameters of the Actor network can be calculated. Use the online policy network to calculate a μ =μ(s t θ P ), use online Q i Network (i=1, 2) evaluation (s t ,a μ )’s Q value. Here, we use an online Q1 network to evaluate a μ The Q value of the evaluation is Maximize using the gradient ascent algorithm And calculate and update θ P The gradient algorithm is defined as:
[0081]
[0082] Where V represents the size of the mini-batch samples.
[0083] The weighted double-delay DDPG updates the target network parameters using a soft update method. The process of updating the target parameters using the online network and the target network is described as follows:
[0084]
[0085] Where τ is the soft update coefficient.
[0086] Step S22: Obtain weights based on weighted double-delay DDPG
[0087] Actor network uses actor function μ(s t θ P ) The state s in the continuous state space S t Mapping to continuous action a t According to the basic information of the device and the initial weight, use function a t =μ(s t θ P ) generates the current similarity weight.
[0088] Step S3: Construct a similar device group based on the obtained weight parameters.
[0089] When constructing similar device groups, the similarity between the access device and each device in the device database is calculated using methods such as edit distance similarity based on the obtained weight parameters. The top 20% of similarity-ranked devices are used to construct similar device groups. Based on the device information model, each device has at least one protocol. The protocols supported by each device in the similar device group are extracted to form a protocol group.
[0090] Step S4: Protocol Verification
[0091] The protocol group is sent to the edge computing node, which verifies each protocol in the protocol group and obtains the best-fitting protocol based on the verification results.
[0092] Example 2:
[0093] To better illustrate the protocol adaptation method, the following is an example of a device access process based on a cloud-edge collaborative architecture in the Internet of Things. The edge computing node is configured as follows: CPU is MT7688AN, main frequency is 580MHz, memory is 128MB, local storage is 32GB, has 4 Ethernet interfaces, 8 RS48 interfaces, 1 CAN interface, 1 4G wireless interface and 1 USB interface. The edge computing node is connected to the device through the RS485 bus and interacts with the Internet of Things platform through the 4G wireless network. The device information library, protocol adapter and device protocol library are located in the Internet of Things platform, and the server deployed by the Internet of Things platform is configured as follows: Aliyun server, CPU frequency is 2.5GHz, memory is 16GB, hard disk is 256GB. The basic information used for similarity is [brand, type, model].
[0094] The basic information of the access device is [Enatel, UPS, SM09Z], and the action value of the similarity weight output by the weighted double-delay DDPG network according to this information is [weight 1 decreases 0.2, weight 2 decreases 0.08, weight 3 increases 0.22]. Since the initial value of the similarity weight is [1, 1, 1], the deterministic weight obtained after the action is [0.8, 0.92, 1.22]. Using this weight to construct a similar device group, the devices in this device group include [Enatel, UPS, SM09L], [Enatel, UPS, SM21], [Enatel, UPS, SM32Ex], [Enatel, UPS, SM31J]. According to the device group, extract the protocol from the protocol library to construct a protocol group, and the details of the protocol group are shown in Table 1:
[0095] Table 1:
[0096]
[0097]
[0098] After verifying the protocol identification and data items of the protocol, the protocol matching rate of protocol 2 is the highest, which is the best matching protocol.
Claims
1. A method for determining protocol adaptation parameters based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Build a device information database; device information includes basic device information and protocol identification; S2. Constructing a weighted double-delay DDPG network including an Actor network and a Critic network; the weighted double-delay DDPG network outputs a similarity weight for determining the similarity of basic information according to the device information; The interaction between the weighted double-delay DDPG network and the environment includes protocol group construction; each device protocol generation task in the protocol group construction is an agent, and the agent performs actions based on rewards and states; The status includes the basic information of the device and the similarity weight; The actions include increasing the similarity weight and decreasing the similarity weight; The Actor network includes an online strategy network and a target strategy network; the online strategy network determines the action according to the current state, and the target strategy network determines the next action according to the next state; S3. Constructing a similar device group based on the similarity weights; S4. Constructing a protocol group based on the similar device group; The protocol texts in the protocol group are verified to obtain the adapted protocol.
2. A method for determining protocol adaptation parameters based on deep reinforcement learning according to claim 1, characterized in that: The critic network includes an online Q network and a target Q network; the online Q network and the target Q network respectively output their own minimum values; and the two minimum values are weightedly combined as the target Q value; The weighted double-delay DDPG network outputs the similarity weight according to the device information and the target Q value.
3. The method for determining protocol adaptation parameters based on deep reinforcement learning according to claim 2, characterized in that: The Actor network updates parameters of the online policy network and the target policy network according to the target Q value.
4. The method for determining protocol adaptation parameters based on deep reinforcement learning according to claim 1, characterized in that: The reward is the sum of the protocol matching rate R and the hit rate H; the protocol matching rate R is defined as: The device protocol is defined as P = [P id ,b],P id represents the protocol identifier, b represents the basic data of the device; * represents the number of correct elements in the generated b, and λ represents the number of real elements in b; the hit rate H is defined as: Where σ represents the number of protocol verifications during the protocol adaptation process.
5. The method for determining protocol adaptation parameters based on deep reinforcement learning according to claim 3, characterized in that: A gradient descent algorithm is used to minimize the error, and the parameters of the online Q network are calculated and updated; after the online Q network is updated, the target Q value evaluated by it is calculated; a gradient ascent algorithm is used to maximize the target Q value, and the parameters of the online policy network are calculated and generated.
6. The method for determining protocol adaptation parameters based on deep reinforcement learning according to claim 1, characterized in that: During the training of the weighted double-delay DDPG network, random noise is added to the action.
7. The method for determining protocol adaptation parameters based on deep reinforcement learning according to claim 1, characterized in that: The weighted double-delay DDPG network also includes a memory replay mechanism and small batch samples; in the memory replay mechanism, the experience tuple (s t ,a t ,r t ,s t+1 ) are stored in the dataset; where s t is the current state, a t is the current action, s t+1 is the next state, a t+1 is the next action; the mini-batch samples are formed by randomly extracting a subset from the dataset for gradient updating.