Communication protocol selection method for power optical communication digital twin system based on deep Q network

By applying the deep Q network (DQN) intelligent communication protocol selection method in the digital twin system, the problems of low transmission efficiency and poor reliability in the power optical communication system are solved, and intelligent communication protocol selection is realized, improving the overall efficiency and reliability of the system.

CN119299532BActive Publication Date: 2025-05-16STATE GRID JILIN ELECTRIC POWER COMPANY LIMITED +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411413123.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-05-16
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

The low transmission efficiency and poor reliability of digital twin systems in power and optical communication systems lead to improper selection of communication protocols and affecting transmission efficiency.

Method used

Using an intelligent communication protocol selection method based on deep Q network (DQN) is adopted, by building a DQN agent network model, using data in the memory for training, optimizing communication protocol selection, and dynamically selecting the optimal communication protocol.

Benefits of technology

It improves the transmission efficiency and reliability of the power optical communication system, realizes dynamic and intelligent communication protocol selection, and is suitable for power system needs in different scenarios and data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119299532B_ABST
    Figure CN119299532B_ABST
Patent Text Reader

Abstract

A communication protocol selection method for a power optical communication digital twin system based on a deep Q network involves power optical communication technology and a digital twin system; the method solves the problems of low transmission efficiency and poor reliability in the existing digital twin communication system. The method obtains power optical communication optical cable data for transmission, selects different communication protocols for different transmission paths of the digital twin system and the form, size and complexity of the required transmission data, calculates the reward feedback value according to the generated delay information, forms a quadruple, and stores it in the memory bank; the DQN intelligent agent network model is trained, and after the training is completed, the results obtained from each interaction are also stored in the memory bank to form a training sample set, and the deep learning optimization algorithm RMSProp algorithm is used to converge the mean square error function of the DQN network, thereby improving the DQN model and improving its accuracy. It enables it to select the most reasonable communication protocol.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power optical communication and digital twin systems, and in particular to a communication protocol selection method for a power optical communication digital twin system based on a deep Q network (DQN). Background Art

[0002] With the continuous development and popularization of information technology, people's demand for high-speed and high-bandwidth communications is becoming more and more urgent. This demand has promoted the modernization and intelligentization of power systems, because modern power systems require reliable, safe and efficient power transmission and monitoring systems. To meet these needs, power optical communication technology has emerged. It not only provides high bandwidth and supports the transmission of large amounts of data, but also has the advantage of low latency. In addition, power optical communication technology also has high security, which further ensures the reliability of the power system. However, with the continuous development and widespread application of this technology, people's requirements for its real-time monitoring function and communication transmission quality are also increasing.

[0003] Digital twin technology is a cutting-edge technology that has emerged in recent years. It is a digital mapping technology for analog entities, which provides a new solution for real-time monitoring of the power optical communication transmission process. Through digital twin technology, real-time detection, simulation and optimization of the power optical communication system can be achieved. However, in the actual application of the digital twin system, information transmission involves the transmission from the physical entity to the digital twin, the interaction between the digital twins, and the transmission process from the digital twin to the database. These transmission processes often face the problem of improper selection of communication protocols, resulting in low transmission efficiency.

[0004] Patent number: 202210830067.2, patent name is method for selecting communication protocol for data transmission, the method proposes a method for communication protocol suitable for data transmission between vehicles and external infrastructure, and determines the weighted proportion of system performance parameters after selecting the protocol, so as to obtain the optimal protocol combination. However, this selection method is mainly applicable to the field of Internet of Vehicles, and it is also unable to optimize weight parameter information such as weight.

[0005] Patent number: 202210122076.6, the patent name is for the selection of management bus communication protocol, the technology proposes a device for system management bus. This device is controlled by detecting trigger events associated with the management bus to select the communication protocol for transmission on the management bus. However, this method only selects the communication protocol through certain trigger events, which cannot guarantee the optimal system performance, and it cannot be applied to other fields, and it cannot optimize the performance.

[0006] Based on the above problems, the present invention proposes a communication protocol selection method for a power optical communication digital twin system based on a deep Q network, aiming to improve the communication efficiency and reliability of the transmission process through intelligent protocol selection. By using the DQN agent network model, different transmission processes are optimized to achieve dynamic and intelligent communication protocol selection, and it is suitable for power system requirements for different scenarios and different transmission data, providing more reliable communication support for the digital transformation of power optical communication systems.

[0007] Through the present invention, the problems of low efficiency and low reliability in the transmission process of the current power optical communication system can be effectively solved, thereby improving the actual application effect of digital twins in the power optical communication system and promoting the digital transformation of the power industry. Summary of the invention

[0008] In order to solve the problems of low transmission efficiency and poor reliability existing in the current digital twin communication system, the present invention provides a communication protocol selection method for an electric power optical communication digital twin system based on a deep Q network.

[0009] A communication protocol selection method for a power optical communication digital twin system based on a deep Q network is implemented by the following steps:

[0010] Step 1: Obtaining power optical communication cable data;

[0011] Step 2: The power optical communication cable data described in step 1 is transmitted in stages. When transmitting data in each stage, a communication protocol is randomly selected to obtain the delay of data transmission in each stage respectively, and then the corresponding reward feedback value and the system state at the next moment are obtained according to the total delay of the transmission data, and the system state at the current moment, the action selected at the current moment, the reward feedback value obtained after the action is selected at the current moment, and the system state at the next moment are stored as a four-tuple in the memory bank;

[0012] Step 3: construct a DQN agent network model, and use the four-tuple data in the memory bank in step 2 to train the network model; the trained network model selects the action with the largest Q value according to the system state as the final selected communication protocol;

[0013] The four-tuple data corresponding to the action with the largest Q value selected by the trained network model is stored in the memory as a test sample set, and the deep learning optimization algorithm RMSProp is used to converge the mean square error function of the network model and optimize the parameters of the network model; in the digital twin system of electric power optical communication, the deep Q network is used to realize the selection of communication protocols for each transmission process.

[0014] Beneficial effects of the present invention:

[0015] The communication protocol selection method described in the present invention collects physical entity data and environmental data through sensors and transmits them in the entire digital twin system. It is transmitted from the physical entity to the digital twin, and at the same time transmitted between digital twins, and the data is uploaded from the digital twin to the database. During the entire data transmission process, the present invention can select appropriate communication protocols for different processes and different data.

[0016] In the method of the present invention, the deep Q network (DQN) is applied to the field of digital twins to optimize the overall transmission process of the twin system. It has the following advantages:

[0017] 1. Since there are multiple digital twin services and selectable communication protocols, this method can select the appropriate communication protocol according to different environments and needs, thereby optimizing transmission performance.

[0018] 2. This method can improve transmission efficiency. Selecting the optimal communication protocol can minimize latency, thereby improving the overall efficiency of data transmission.

[0019] 3. The method described in the present invention can achieve resource optimization and optimize the resource configuration of the entire system by selecting an appropriate communication protocol.

[0020] In short, the DQN intelligent agent model enables the digital twin system to select the appropriate communication protocol according to the data form, size, complexity and different transmission processes, so as to improve the performance, improve the actual application effect of digital twins in the power optical communication system, and promote the digital transformation of the power industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flow chart of the communication protocol selection method of the power optical communication digital twin system based on the deep Q network described in the present invention;

[0022] Figure 2 This is a flowchart of the DQN agent network model;

[0023] Figure 3 This is a graph showing the relationship between the reward value and the number of iterations after a single path protocol selection using the DQN algorithm. DETAILED DESCRIPTION

[0024] Combination Figures 1 to 3 This embodiment describes a method for selecting a communication protocol for a power optical communication digital twin system based on a deep Q network, which is implemented by the following steps:

[0025] Step 1: Obtain data information. From the physical entity, i.e., the power optical communication cable, use sensors to obtain data for transmission, which may include the physical data, performance parameters, and environmental parameters of the optical fiber physical entity.

[0026] In this implementation, in order to improve the transmission quality of the power optical communication system and realize functions such as real-time monitoring, the power optical communication network is detected, simulated and optimized in real time by using a digital twin system. Therefore, it is necessary to select an appropriate modeling method based on information such as the physical entity and network status of the power optical communication system to establish a digital twin system for power optical communication.

[0027] Among them, the digital twin system of the currently applicable power optical communication system generally includes a camera system, a micro-weather station system, a communication room dynamic environment system, and a communication equipment power data monitoring system. In order to construct a digital twin system that can simulate physical entities in the real world, the data transmitted in step 1 can include:

[0028] (1) Physical data of the optical fiber, such as optical fiber material data, such as the proportion of silicon dioxide SiO2 η, optical fiber radius a, optical fiber length L, etc.;

[0029] (2) Fiber performance parameters, such as fiber attenuation coefficient α, fiber refractive index n(r), and fiber cutoff wavelength λ c ;

[0030] The optical fiber attenuation coefficient α refers to the gradual attenuation of optical power as the transmission distance increases. This performance parameter directly affects the length of the optical fiber communication transmission distance. The specific calculation formula is:

[0031]

[0032] Among them, P i is the power of the input optical fiber; P o is the output fiber power.

[0033] The refractive index n(r) of the optical fiber determines the propagation mode and transmission performance of light in the optical fiber. The specific formula of the optical fiber refractive index (parabolic refractive index profile) is:

[0034]

[0035] Where n(r) is the refractive index at a distance r from the central axis of the optical fiber; n1 is the refractive index at the center of the optical fiber; Δ is the relative refractive index difference of the optical fiber; and a is the radius of the optical fiber.

[0036] The cut-off wavelength λ of the optical fiber c , this parameter determines the transmission mode of the optical fiber. The specific calculation formula of the cut-off wavelength of the optical fiber is:

[0037]

[0038] Among them, λ c is the cut-off wavelength of the optical fiber; a is the radius of the optical fiber; n1 is the refractive index at the center of the optical fiber; Δ is the relative refractive index difference of the optical fiber. The specific formula of Δ is:

[0039]

[0040] Where n2 is the refractive index of the outer layer of the optical fiber.

[0041] (3) Parameters of the environment in which the optical fiber is located, such as the ambient temperature T, humidity RH, stress σ on the optical fiber, and other data information.

[0042] Step 2: Transmit data. The entire transmission process is divided into three stages: the data transmission process from the physical entity to the digital twin, the data transmission process between digital twins, and the data transmission process between the digital twin and the database. When the system transmits data at each stage, it will randomly select a communication protocol, obtain the delay of each stage, and add them together to obtain the total delay of the entire transmission process. The total delay is then used to obtain the corresponding reward value, and the state of the next moment is generated at the same time. The current state, selected action, generated reward feedback value and the state of the next moment are formed into a four-tuple and stored in the memory bank.

[0043] In this implementation, in the power optical communication digital twin system, in order to improve the performance and algorithm efficiency, the following strategies will be adopted in this implementation: the transmission efficiency of each transmission process can be optimized separately during the interaction process, and the three processes can be arbitrarily combined in pairs, and the DQN intelligent agent network model can be combined to select the best communication protocol with the lowest latency and the lowest latency of the two, so as to further improve the overall efficiency of the system.

[0044] The communication protocols that can be selected by these three processes are as follows:

[0045] For the transmission process from physical entities to digital twins, sensors, data acquisition devices, etc. are mainly used to obtain data and transmit it to digital twins. Using the transmitted data, a digital model of the physical entity is created, including information on structure, performance, and behavior. The communication transmission of this process requires obtaining data from the physical entity and converting it into digital form for communication. Therefore, the communication protocol for this process will select the following communication protocols: (1) MQTT communication protocol, which has high efficiency, low latency, and bandwidth saving performance, and is particularly suitable for scenarios with unstable wired and wireless network connections; (2) OPC UA communication protocol, which is widely used in industrial automation and the Internet of Things, has good scalability, cross-platform and high security, and is suitable for wired network environments; (3) 5G communication protocol, which is suitable for high bandwidth, low latency, and supports large-scale device connections, and is suitable for high-density wireless network environments.

[0046] For the data transmission process between digital twins, the data transmitted from the physical entity is mainly exchanged and integrated between the digital twins to facilitate comprehensive analysis and service implementation. The communication protocol of this process is required to be as efficient as possible, so the optional communication protocols include the following: (1) HTTP communication protocol, which is easy to implement and suitable for scenarios with low real-time requirements; (2) AMQP communication protocol, digital twins at different sites or regions can use this communication protocol to achieve asynchronous data transmission and task scheduling, with high reliability; (3) DDS communication protocol, which is suitable for real-time data exchange scenarios.

[0047] The data transmission process between the digital twin and the database mainly involves storing the digital twin data in the database for management and analysis. The optional communication protocols for this process include: (1) ODBC communication protocol, which is a standard database access interface with enhanced versatility; (2) RESTAPI communication protocol, which is suitable for remote database access scenarios.

[0048] Step 3: Build a DQN agent network model and complete model training. After multiple iterations, use the four-tuple data in the memory bank to train the DQN agent network model so that after the training, it can select the communication protocol action with the maximum Q value with a higher probability according to the form, size and complexity of the transmitted data. That is, after training, the DQN agent network model can select the communication protocol with the largest Q value for data transmission based on different states based on the greedy algorithm; and use the mean square error formula to use the RMSProp algorithm to optimize the DQN model parameters and improve the DQN agent network model.

[0049] The states, actions, and reward feedback values ​​corresponding to the DQN agent network model in this implementation specifically include:

[0050] Status: the form, size and complexity of the data to be communicated;

[0051] (1) Data forms include digital data and analog data. Digital data includes binary data; analog data refers to continuous signals, such as audio signals, video signals, etc.

[0052] (2) The data size formula is:

[0053] Data size (bits) = data size (bytes) × 8

[0054] (3) Data complexity is usually used to describe time complexity and space complexity, indicating the time and space resources required for execution. The specific formula for time complexity is:

[0055] Time complexity = O(f(n))

[0056] Where f(n) represents a function of the input size n.

[0057] Its specific form is:

[0058] F(t): represents the data format of the data transmitted at time t.

[0059] D(t): represents the data size of the data transmitted at time t.

[0060] C(t): represents the data complexity of data transmission at time t.

[0061] At time t, the state space of the system is:

[0062] s(t)=[F(t),D(t),C(t)]

[0063] Action: The behavior of the system covers four key actions, namely:

[0064] (1) The first action is the path selection method. The system selects single-path transmission or multi-path transmission. In single-path transmission, the selection of communication protocols will be applied to the three processes of physical entity to digital twin, between digital twins, and from digital twin to database. In multi-path transmission, the system will select any two of the above processes and combine them, selecting two communication protocols with relatively small total delay for them, while the remaining process will select the communication protocol independently.

[0065] (2) The second action is to select the communication protocol during the transmission process from the physical entity to the digital twin, among which MQTT, OPC UA, and 5G can be selected at will.

[0066] (3) The third action is to select the communication protocol during the transmission process between digital twins, among which HTTP, AMQP, and DDS can be arbitrarily selected.

[0067] (4) The fourth action is to upload data from the digital twin to the database. In this process, you can select the communication protocol. You can choose ODBC or RESTAPI.

[0068] In this implementation, the action space is specifically expressed as:

[0069] a1(t): a1(t)∈[0,1], indicating the selected path transmission mode, single path transmission or multi-path transmission.

[0070] a2(t): a2(t)∈[0,1,2], indicating the selection of communication protocol during the transmission from physical entity to digital twin.

[0071] a3(t): a3(t)∈[0,1,2], indicating the selection of communication protocol during the transmission process between digital twins.

[0072] a4(t): a4(t)∈[0,1], indicating the selection of communication protocol during the transmission process between the digital twin and the database.

[0073] Then at time t, the action space selected by the system is:

[0074] a(t)=[a1(t),a2(t),a3(t),a4(t)]

[0075] Reward feedback value: By calculating the total delay and other information generated by the entire process after selecting any communication protocol, the reward feedback value of the corresponding method is obtained using the reward feedback value formula.

[0076] In this implementation, for different transmission processes, the delay caused by selecting different communication protocols is used as the criterion for the reward feedback value required in the algorithm. The total delay formula for the entire process is:

[0077] T=t1+t2+t3

[0078] Among them, T is the total delay of the whole process, t1 is the delay caused by the selection of communication protocol for transmission from the physical entity to the digital twin, t2 is the delay caused by the selection of communication protocol for transmission between digital twins, and t3 is the delay caused by the selection of communication protocol for transmission from the digital twin to the database.

[0079] In this implementation, a DQN agent network model is constructed. The DQN algorithm mainly obtains an estimated Q value table and a target Q value table by establishing a Q table, where the Q value represents the expected future reward of the action when the state is s and the action is selected as a. The update formula of the Q value is:

[0080] Q(s,a)=Q(s,a)+α[r(s,a)+γmaxQ(s',a')-Q(s,a)]

[0081] Among them, α represents the learning rate; r(s,a) represents the reward value for state s and action a; γ represents the decay factor, which determines the probability of taking subsequent actions; r(s,a)+γmaxQ(s',a') represents the estimated Q value; Q(s,a) on the right side of the equation represents the original Q value in the Q table. At the same time, the DQN algorithm uses the stored estimated Q value to assign the target Q value, so that the agent can select the appropriate action without storage.

[0082] The DQN agent network model is trained using the training set obtained by the agent interacting with the environment (i.e., the four-tuple stored in the memory bank in step 2). After each interaction between the agent and the environment, a four-tuple [s t ,a t ,r t ,s t+1 ], where s t represents the system state at the current time t, a t represents the action selected at the current time t, r t represents the feedback reward value after selecting the action at the current time t, s t+1 It represents the system state at the next moment t+1. The four-tuple corresponding to each moment is stored in the memory bank. After a certain number of iterations, the DQN model is trained using these data, and the system has a greater probability of selecting the action with the largest Q value.

[0083] In this implementation, the corresponding total delay is obtained according to the different actions selected for each state, thereby calculating the reward feedback value and storing it in the memory bank of the DQN algorithm. After the training is completed, the current system state, including data format, size, and complexity, is directly input into the DQN network model, and the action with the largest Q value is selected based on the ε-greedy greedy algorithm as the final selected communication protocol.

[0084] In this implementation, the ε-greedy greedy algorithm is specifically that the DQN algorithm uses the exploration rate ε to determine the action selected by the agent, randomly selects an action with a probability of ε, and selects the action with the largest Q value with a probability of 1-ε. The specific formula of the ε-greedy greedy algorithm is:

[0085]

[0086] By gradually decreasing ε during the training process, it can explore more in the initial stage of training and finally converge to a better strategy.

[0087] At the same time, the deep reinforcement learning DQN algorithm separates the target network by obtaining the estimated Q value that conforms to the state environment. After a fixed number of iterations, the generated estimated Q value is directly transferred to the target Q table, thereby reducing the possibility of divergence in training and improving the convergence speed.

[0088] As Figure 2 shown, Figure 2 is the algorithm flowchart based on the DQN algorithm;

[0089] Step 3-1: Initialize and define the parameters of the DQN agent network model, including the parameters in the estimation network and the target network, the learning rate α, the decay factor γ, the greedy coefficient ε, etc. Among them, the learning rate α represents a parameter for weight update in the neural network, which is used to control the step size or rate of weight update. The decay factor γ determines the probability of taking subsequent actions. The greedy coefficient ε, also known as the exploration rate, means that in the process of the agent choosing an action, there is a certain probability of choosing the action with the maximum reward, while there is still a certain probability of choosing other actions, which affects the subsequent learning.

[0090] Step 3-2: Define the number of training times K of the DQN agent network model;

[0091] Step 3-3: Obtain the data form F t , the data size D t and the complexity C t as the state s at the current time t t ;

[0092] Step 3-4: Judge whether the interaction times i < K. If so, randomly select an action a t , and execute Step 3-5; otherwise, select an action a based on the ε-greedy greedy strategy t ; Execute Step 3-5;

[0093] Step 3-5: Execute the action to obtain the reward feedback value r t and the next system state s t+1 , forming a quadruple [s t , a t , r t , s t+1 and store it in the memory bank;

[0094] Step 3-6: Judge whether the memory bank is full. If so, overwrite the old stored samples and execute Step 3-7; otherwise, directly store; execute Step 3-8;

[0095] Step 38: Randomly select samples from the memory bank, calculate the estimated Q value and the target Q value, and calculate the mean square error function;

[0096] Step 39: If the mean square error function has not converged, the deep learning optimization algorithm RMSProp is used to optimize the parameters of the DQN model. After a certain number of times, the optimized estimated network parameter structure is assigned to the target network, and the process returns to step 38; if the mean square error function has converged, the process ends.

[0097] Step 4: Optimize parameters. The four-tuple formed after multiple interactions is also stored in the memory bank to form a test sample set. After multiple iterations, the deep learning optimization algorithm RMSProp is used to converge the mean square error function of the network, thereby improving the DQN agent model, optimizing related parameters, and improving the algorithm accuracy.

[0098] In this implementation, after the training of the DQN agent is completed, the four-tuple obtained from each interaction is also stored in the memory bank for cyclic coverage. The difference between the estimated Q value and the target Q value is used to update the parameters of the DQN network using the deep learning optimization algorithm RMSProp algorithm to improve its accuracy. The formula of the mean square error function in the DQN model is:

[0099] L(θ)=E[r(s,a)+γ*maxQ(s',a',θ)-Q(s,a,θ)]

[0100] Among them, θ represents the parameter set in the DQN algorithm; r(s,a) represents the reward feedback value obtained by selecting action a in state s; γ is the decay factor, which determines the probability of taking subsequent actions. The system measures whether the accuracy of the model meets the requirements by the value of the mean square error function, and uses the deep learning optimization algorithm RMSProp to update the parameters of the DQN agent, which helps to reduce the instability of the training process and improve the performance and convergence speed of the algorithm.

[0101] In this implementation, the deep learning optimization algorithm RMSProp algorithm uses different learning rates to optimize each parameter and adaptively adjusts the parameters according to the historical gradient size. It mainly uses exponential weighted average to calculate the moving average of the square gradient of the mean square error function. The specific formula is:

[0102] n=β*n+(1-β)*g^2

[0103] Among them, n is represented as a moving average in the algorithm, which is initially 0; β is the attenuation factor in the optimization algorithm, which controls the influence weight of the historical gradient; g is the gradient of the parameter.

[0104] The learning rate adaptive update formula of the deep learning optimization algorithm RMSProp algorithm is:

[0105]

[0106] Among them, η is the learning rate of updating parameters, and ε is a very small constant to avoid the situation where the denominator is 0. The final update parameter formula is

[0107] x=xg*η′

[0108] Among them, x is the parameter that needs to be updated. The deep learning optimization algorithm RMSProp algorithm uses the exponential decay average of historical gradients to accelerate the convergence of the algorithm, thereby improving the efficiency of the DQN intelligent model and enabling it to more accurately select the appropriate communication protocol.

[0109] like Figure 3 As shown, this embodiment gives a schematic diagram of the simulation results of the digital twin system communication protocol selection method based on the DQN algorithm applied to a single path of power optical communication. In the figure, by adjusting the reward value formula parameters and other methods, it clearly reflects the trend of change. After multiple iterations, a correlation image of the final reward value and the number of iterations was obtained. Through the simulation image, it can be obtained that as the number of iterations increases, the reward value obtained by the system also gradually increases, and finally converges, which means that through the DQN model, the system can select the communication protocol that produces the smallest total delay, which can increase the total revenue of the system and maintain stability, thereby improving the efficiency and performance of the entire system.

[0110] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A communication protocol selection method for a power optical communication digital twin system based on a deep Q network is characterized by: The method is implemented by the following steps: Step 1: Obtaining power optical communication cable data; Step 2: The power optical communication cable data described in step 1 is transmitted in stages. When transmitting data in each stage, a communication protocol is randomly selected to obtain the delay of data transmission in each stage respectively, and then the corresponding reward feedback value and the system state at the next moment are obtained according to the total delay of the transmission data, and the system state at the current moment, the action selected at the current moment, the reward feedback value obtained after the action is selected at the current moment, and the system state at the next moment are stored as a four-tuple in the memory bank; Step 3: construct a DQN agent network model, and use the four-tuple data in the memory bank in step 2 to train the network model; the trained network model selects the action with the largest Q value according to the system state as the final selected communication protocol; The four-tuple data corresponding to the action with the largest Q value selected by the trained network model is stored in the memory bank as a test sample set, and the deep learning optimization algorithm RMSProp is used to converge the mean square error function of the network model and optimize the parameters of the network model; In the digital twin system of electric power optical communication, the deep Q network is used to select the communication protocol for each transmission process; In step 2, the transmission of power optical communication cable data is divided into three stages: the data transmission process from the physical entity to the digital twin, the data transmission process between digital twins, and the data transmission process between the digital twin and the database; The specific implementation process of step three is: Step 31: Initialize the DQN agent network model parameters, including the parameters in the estimation network and the target network, the learning rate α, the decay factor γ, and the greed coefficient ε; Step 32: Define the maximum number of training times K for the DQN agent model; Step 3: The system randomly selects an action according to the current state at time t, obtains a reward feedback value, generates a quadruple and stores it in the memory bank, and then uses the quadruple data in the memory bank to train the DQN intelligent agent network model; Step 3 and 4: When the number of training times reaches the maximum number of training times K, the model training ends; the action selection is performed through the ε-greedy algorithm, the action is randomly selected with a probability of ε, and the action with the largest Q value is selected with a probability of 1-ε; Step 35: store the four-tuple formed after model training in the memory bank. When the memory bank is full, perform a loop overwrite and randomly select the stored samples in the memory bank to calculate the estimated Q value and the target Q value. Step 36: Using the difference between the estimated Q value and the target Q value, the deep learning optimization algorithm RMSProp is used to optimize the network parameters through the mean square error function; the parameter structure of the estimated Q value table is assigned to the target Q value table; Step 37: When the mean square error function converges, the process ends.

2. According to claim 1, the method for selecting a communication protocol for a power optical communication digital twin system based on a deep Q network is characterized in that: The power optical communication cable data described in step 1 includes physical data of the optical fiber, performance parameters of the optical fiber and environmental parameters of the optical fiber.

3. According to claim 2, the method for selecting a communication protocol for a power optical communication digital twin system based on a deep Q network is characterized in that: The physical data of the optical fiber includes the optical fiber material data; the performance parameters of the optical fiber include the optical fiber attenuation coefficient, the optical fiber refractive index and the optical fiber cut-off wavelength; the environmental parameters of the optical fiber include the ambient temperature and humidity of the optical fiber and the stress to which the optical fiber is subjected.

4. According to claim 1, the method for selecting a communication protocol for a power optical communication digital twin system based on a deep Q network is characterized in that: In step 2, according to the transmission process at different stages, the delay caused by different communication protocols is selected as the criterion for judging the reward feedback value; the total delay formula is: T=t1+t2+t3 Where T is the total delay of the entire transmission process, t1 is the delay caused by the selection of the communication protocol for transmission from the physical entity to the digital twin, t2 is the delay caused by the selection of the communication protocol for transmission between digital twins, and t3 is the delay caused by the selection of the communication protocol for transmission from the digital twin to the database.

5. According to claim 1, the method for selecting a communication protocol for a power optical communication digital twin system based on a deep Q network is characterized in that: In step 3, by establishing a Q value table, an estimated Q value table and a target Q value table are obtained; The Q value is the expected future reward of the action when the state is s and the action is a; and the target Q value is assigned using the estimated Q value.

6. According to claim 1, the method for selecting a communication protocol for a power optical communication digital twin system based on a deep Q network is characterized in that: In step three, the system status includes the form, size and complexity of the transmitted data.

Citation Information

Patent Citations

  • Selection for managing bus communication protocols

    CN115048326A

  • Method for selecting communication protocol for data transmission

    CN115623089B

  • Network technology and protocol test platform based on digital twinning and test method thereof

    CN114520781A

  • Underwater acoustic network MAC protocol based on Double DQN

    CN117061014A