A time delay dynamic compensation method, device, equipment, medium and system
By using a reinforcement learning model for dynamic latency compensation, the problem of the inability to automatically adjust latency asymmetry in existing technologies is solved, enabling real-time detection and automatic compensation, thereby improving network synchronization accuracy and service quality.
Patent Information
- Application Number
- CN202310808322.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-03
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-07-03
AI Technical Summary
In existing technologies, asymmetric delay compensation methods cannot be automatically adjusted, which affects network synchronization accuracy and service quality. Furthermore, the delay asymmetry value cannot be detected in real time, requiring frequent manual intervention.
A time delay dynamic compensation method based on a reinforcement learning model is adopted. Through the signal interaction between the synchronization network probe device and the transmission device, the time delay compensation is performed using a reinforcement learning model, including the measurement, deviation analysis and dynamic compensation of the time synchronization signal, and the action compensation decision is made using a value network and a policy network.
It enables real-time detection and automatic compensation of latency asymmetry values, reduces manual intervention, maintains equipment synchronization accuracy, avoids instantaneous clock precision jumps, and realizes the automation and intelligence of network operation and maintenance.
Smart Images

Figure CN116708240B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication network, and particularly relates to a time delay dynamic compensation method, device, equipment, medium and system. BACKGROUND
[0002] In the prior art, the main way of asymmetric time delay compensation is to measure the fixed two-way delay difference by means of probes, instruments or 5G base station built-in, and then to compensate the time delay, but this way cannot automatically compensate, and if it cannot automatically compensate, the following shortcomings exist:
[0003] 1) When the application environment changes, such as the optical fiber path changes caused by network OLP switching, the two-way delay difference will change, affecting the synchronization accuracy and service quality of the network, at this time, it is necessary to manually obtain the two-way delay difference again and recompensate;
[0004] 2) Each measurement is compensated once according to the measurement value, and when the correction value is large, it may cause the instantaneous jump of clock accuracy and affect the service;
[0005] 3) It cannot detect the time delay asymmetric value in real time. SUMMARY
[0006] The present application solves the above-mentioned problems of the prior art, and provides a time delay dynamic compensation method, device, equipment, medium and system based on a reinforcement learning model, which can automatically compensate the time delay.
[0007] In a first aspect, the present application provides a time delay dynamic compensation method, which is applied to a synchronous network transmission device, and the method comprises the following steps:
[0008] Step K1: sending a first time synchronization signal to a synchronous network probe device;
[0009] Step K2: receiving a first time synchronization signal deviation returned by the synchronous network probe device;
[0010] Wherein:
[0011] The first time synchronization signal deviation is obtained by measuring and deviation analyzing the first time synchronization signal by the synchronous network probe device, the first time synchronization signal deviation is the difference between the first theoretical time and the first actual time, the first theoretical time is the theoretical time when the first time synchronization signal reaches the synchronous network probe device, and the first actual time is the actual time when the first time synchronization signal reaches the synchronous network probe device;
[0012] Step K3: compensating the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model to obtain a second time synchronization signal;
[0013] Step K4: sending the second time synchronization signal to the synchronization network probe device, so as to complete the dynamic compensation of the time delay.
[0014] Further, in the step K3, the first time synchronization signal is compensated according to the first time synchronization signal deviation and the reinforcement learning model to obtain the second time synchronization signal, and the step specifically comprises steps of:
[0015] Step K31: comparing the first time synchronization signal deviation with a preset threshold value:
[0016] When the first time synchronization signal deviation is greater than the preset threshold value, step K32 is entered; when the first time synchronization signal deviation is less than or equal to the preset threshold value, the first time synchronization signal is not compensated for time delay, and the first time synchronization signal is modified as the second time synchronization signal.
[0017] Step K32: obtaining the environment state S t corresponding to the first time synchronization signal t as the input of the reinforcement learning model; the environment state comprises a time error TE, a maximum time interval error MTIE, a time deviation TDEV and a frequency deviation;
[0018] Step K33: outputting an action compensation in the reinforcement learning model according to the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation to obtain the second time synchronization signal; the outputting of the action compensation in the reinforcement learning model comprises outputting a compensation direction and an output compensation value step length in the reinforcement learning model, the compensation direction comprises a positive compensation, a negative compensation and a zero compensation, and the compensation value step length is a preset value;
[0019] Step K34: measuring and analyzing the second time synchronization signal to obtain a second time synchronization signal deviation, and comparing the second time synchronization signal deviation with a preset threshold value:
[0020] When the second time synchronization signal deviation is greater than the preset threshold value, steps K32 to K33 are repeated until the second time synchronization signal deviation is less than or equal to the preset threshold value, and the first time synchronization signal is modified as the second time synchronization signal.
[0021] Further, the reinforcement learning model is a value network and a policy network based on the first time synchronization signal deviation,
[0022] In the step K33, the action compensation is outputted in the reinforcement learning model according to the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation to obtain the second time synchronization signal, and the step specifically comprises:
[0023] The environment state S corresponding to the first time synchronization signal t An input value network, which obtains time delay compensated values of each possible action according to the environment state, each possible action value including a positive compensation action, a negative compensation action and a zero compensation action, the sum of the probability of the positive compensation action, the probability of the negative compensation action and the probability of the zero compensation action being 1; the environment state S t Including time error TE, maximum time interval error MTIE, time deviation TDEV and frequency deviation;
[0024] The policy network selects time delay compensation according to the maximum probability value of each possible action value;
[0025] After the synchronous network transmission equipment receives the action, the environment state changes to the environment state S t+1 , and returns the reward value r to the probe equipment t ;
[0026] Wherein:
[0027] When the absolute value of the maximum value of the up and down peak values of the TE clock performance at the environment state S t+1 is smaller than or equal to the absolute value of the maximum value of the up and down peak values of the TE clock performance at the environment state S t ,
[0028] Or,
[0029] The MTIE of the environment state S t+1 is smaller than the MTIE of the environment state S t ;
[0030] Or,
[0031] The TDEV of the environment state S t+1 is smaller than the TDEV of the environment state S t ;
[0032] Or,
[0033] The MTIE of the environment state S t+1 is smaller than the MTIE specified in the ITU-T G.811 / G.812 standard,
[0034] Or,
[0035] The TDEV of the environment state S t+1 is smaller than the TDEV specified in the ITU-T G.811 / G.812 standard,
[0036] Or,
[0037] The frequency deviation of the environment state S t+1 is smaller than the frequency deviation of the environment state S tThe frequency offset is closer to 0,
[0038] The reward value r t is a plus reward value.
[0039] The reward value r t is a minus reward value.
[0040] Further, in the step K32, the environment state S t is the environment state corresponding to the first time synchronization signal obtained after the time range threshold X;
[0041] In the step K33, the action compensation is output in the reinforcement learning model according to the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation, and a second time synchronization signal is obtained, including the following steps:
[0042] A time range threshold X and a performance threshold Y are determined, wherein the time range threshold X represents X consecutive t time, and the performance threshold Y is a threshold range of the time error TE;
[0043] For each t time range data, the average value μ and the standard deviation σ of TE are calculated;
[0044] It is judged whether the value of TE is distributed in the range of (μ-3σ, μ+3σ):
[0045] If the average value of TE in the X consecutive t time range is less than Y, and the value of TE is distributed in the range of (μ-3σ, μ+3σ), the action compensation output in the reinforcement learning model is zero compensation, it is determined that the second time synchronization signal is the same as the first time synchronization signal, and the first time synchronization signal is replaced by the second time synchronization signal; if the average value of TE of any t time range data is greater than Y or the value of TE is not distributed in the range of (μ-3σ, μ+3σ), it is determined that the second time synchronization signal is the first time synchronization signal for positive compensation or negative compensation.
[0046] In a second aspect, the present application provides a time delay dynamic compensation method, which is applied to a synchronization network probe device, and includes the following steps:
[0047] Step S1: receiving a first time synchronization signal transmitted by a synchronization network transmission device;
[0048] Step S2: measuring and analyzing the deviation of the first time synchronization signal to obtain a first time synchronization signal deviation, and sending the first time synchronization signal deviation to the synchronization network transmission device;
[0049] Wherein:
[0050] The first time synchronization signal deviation is a difference between a first theoretical time and a first actual time, the first theoretical time is a theoretical time when the first time synchronization signal reaches the synchronization network probe device, and the first actual time is an actual time when the first time synchronization signal reaches the synchronization network probe device.
[0051] Step S3: receiving a second time synchronization signal transmitted by the synchronization network transmission device, so as to complete dynamic delay compensation; the second time synchronization signal is obtained by compensating the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model.
[0052] In a third aspect, the present application provides a dynamic delay compensation device, which is applied to a synchronization network probe device, and comprises:
[0053] A first receiving unit is configured to receive a first time synchronization signal transmitted by a synchronization network transmission device.
[0054] A measurement and analysis unit is connected to the first receiving unit and is configured to measure and analyze the first time synchronization signal to obtain a first time synchronization signal deviation, and send the first time synchronization signal deviation to the synchronization network transmission device; the first time synchronization signal deviation is a difference between a first theoretical time and a first actual time, the first theoretical time is a theoretical time when the first time synchronization signal reaches the synchronization network probe device, and the first actual time is an actual time when the first time synchronization signal reaches the synchronization network probe device.
[0055] The first receiving unit is further configured to receive a second time synchronization signal transmitted by the synchronization network transmission device, so as to complete dynamic delay compensation.
[0056] The second time synchronization signal is obtained by compensating the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model.
[0057] In a fourth aspect, the present application provides a dynamic delay compensation device, which is applied to a synchronization network transmission device, and comprises:
[0058] A first sending unit is configured to send a first time synchronization signal to a synchronization network probe device.
[0059] A second receiving unit is connected to the first sending unit and is configured to receive a first time synchronization signal deviation returned by the synchronization network probe device; the first time synchronization signal deviation is a difference between a first theoretical time and a first actual time, the first theoretical time is a theoretical time when the first time synchronization signal reaches the synchronization network probe device, and the first actual time is an actual time when the first time synchronization signal reaches the synchronization network probe device.
[0060] The compensation unit is connected with the second receiving unit and is configured to compensate the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model to obtain a second time synchronization signal.
[0061] The second sending unit is connected with the compensation unit and is configured to send the second time synchronization signal to the synchronization network probe device, so as to complete the dynamic compensation of the time delay.
[0062] In a fifth aspect, the present application provides an electronic device, which comprises a memory and a processor, and the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the time delay dynamic compensation method according to the above.
[0063] In a sixth aspect, the present application provides a computer readable storage medium, which stores a computer program, and when the processor executes the computer program, the processor executes the time delay dynamic compensation method according to the above.
[0064] In a seventh aspect, the present application provides a time delay dynamic compensation system, which comprises a synchronization network probe device and a synchronization network transmission device,
[0065] The synchronization network transmission device transmits a first time synchronization signal to the synchronization network probe device.
[0066] The synchronization network probe device sends a first time synchronization signal deviation to the synchronization network transmission device.
[0067] The synchronization network probe device receives a second time synchronization signal transmitted by the synchronization network transmission device, so as to complete the dynamic compensation of the time delay; the second time synchronization signal is obtained by compensating the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model.
[0068] Wherein:
[0069] The first time synchronization signal deviation is obtained by measuring and deviation analyzing the first time synchronization signal by the synchronization network probe device, the first time synchronization signal deviation is the difference between a first theoretical time and a first actual time, the first theoretical time is the theoretical time when the first time synchronization signal reaches the synchronization network probe device, and the first actual time is the actual time when the first time synchronization signal reaches the synchronization network probe device.
[0070] The present application has the following beneficial effects:
[0071] 1. The present application can detect the time delay asymmetry value in real time and automatically compensate based on the reinforcement learning model for dynamic compensation of time delay.
[0072] 2. The application can detect the delay asymmetry value in real time and automatically compensate when the network environment changes, greatly simplifying the workload of delay measurement and asymmetry compensation, and without manual intervention.
[0073] 3. When the device is in a hold state after losing lock, the application can correct the device synchronization accuracy deviation back to the normal range during the hold state through continuous dynamic compensation.
[0074] 4. The application performs delay dynamic compensation based on a reinforcement learning model, so that delay compensation can be implemented in a gradual and progressive manner, avoiding sudden large value compensation that can cause instantaneous accuracy jumps and affect services, and achieving automation and intelligentization of network operation.
[0075] 5. The application can effectively ensure the stability of delay compensation by setting the time range and performance threshold, thereby avoiding the impact of sudden fluctuations on the communication system. BRIEF DESCRIPTION OF DRAWINGS
[0076] Figure 1 A schematic diagram of a synchronous network transmission device in an embodiment of the application using reinforcement learning to implement dynamic delay compensation;
[0077] Figure 2 A 1588v2 topology diagram in an embodiment of the application;
[0078] Figure 3 A policy network in a reinforcement learning model in an embodiment of the application;
[0079] Figure 4 A schematic diagram of a 5G base station device in an embodiment of the application using reinforcement learning to implement dynamic delay compensation;
[0080] Figure 5 A schematic diagram of a delay dynamic compensation process in an embodiment of the application;
[0081] Figure 6 A schematic diagram of a delay dynamic compensation device in an embodiment of the application;
[0082] Figure 7 Another schematic diagram of a delay dynamic compensation device in an embodiment of the application.
[0083] Wherein, the reference signs: 10, first receiving unit, 20, measurement and analysis unit, 30, first sending unit, 40, second receiving unit, 50, compensation unit, 60, second sending unit. DETAILED DESCRIPTION
[0084] To enable those skilled in the art to better understand the technical solutions of the present application, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0085] It can be understood that the specific embodiments and drawings described herein are merely intended to explain the present application, but not to limit the present application.
[0086] It can be understood that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0087] It can be understood that, for the convenience of description, only parts related to the present application are shown in the drawings of the present application, and parts irrelevant to the present application are not shown in the drawings.
[0088] It can be understood that each unit and module involved in the embodiments of the present application can correspond to only one entity structure, or can be composed of multiple entity structures, or multiple units and modules can be integrated into one entity structure.
[0089] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present application can occur in an order different from that marked in the drawings.
[0090] It can be understood that, in the flowcharts and block diagrams of the present application, the architecture, functions and operations of the possible implementations of the system, device, equipment and method according to the embodiments of the present application are shown. Each block in the flowchart or block diagram can represent a unit, module, program segment, code, which contains executable instructions for realizing the specified functions. Moreover, each block or combination of blocks in the block diagram and flowchart can be realized by a hardware-based system for realizing the specified functions, or by a combination of hardware and computer instructions.
[0091] It can be understood that the units and modules involved in the embodiments of the present application can be realized in the form of software or hardware, for example, the units and modules can be located in a processor.
[0092] Embodiment 1:
[0093] As shown in Figure 1 , Figure 4 and Figure 5 , the present embodiment provides a time delay dynamic compensation method, which is applied to a synchronous network transmission device,
[0094] The method comprises the following steps K1 to K4:
[0095] Step K1: sending a first time synchronization signal to a synchronous network probe device.
[0096] Dynamic Delay Compensation is a technology for measuring and correcting the time delay in a synchronous network. In a synchronous network, due to the time delay of network transmission, data packets will inevitably experience a certain delay during transmission. In order to ensure the synchronization of clocks between nodes in the network, it is necessary to measure and compensate for these time delays. The basic principle of dynamic delay compensation is to estimate the time delay by measuring the time difference of data packets arriving in the network, and to correct the clock according to the estimated time delay value. The specific implementation can be to record the sending and receiving time stamps of data packets on the nodes in the network, then calculate the time delay of data packets during transmission, and according to the measurement results, adjust the frequency or phase of the local clock to achieve compensation for the time delay.
[0097] In the process of dynamic delay compensation, the first time synchronization signal needs to be sent to the synchronization network probe device. The first time synchronization signal is an accurate time mark or a periodic synchronization signal. By sending the synchronization signal to the synchronization network probe device, the probe device can obtain the most accurate synchronization reference time. The synchronization network probe device usually communicates with the synchronization network transmission device and determines the time delay in the network by measuring the delay. The synchronization network probe device can detect various time delays introduced during data transmission, such as transmission delay, processing delay, and propagation delay. By collecting and analyzing these time delay information, the time delay difference between different nodes in the network can be calculated and dynamically compensated, and the time delay compensation strategy can be continuously monitored and adjusted in real time, so as to ensure that the entire synchronization network always maintains high precision and stable synchronization state.
[0098] Step K2: receiving the first time synchronization signal deviation returned by the synchronization network probe device.
[0099] Specifically, the first time synchronization signal deviation is obtained by measuring and analyzing the first time synchronization signal by the synchronization network probe device. The first time synchronization signal deviation is the difference between the first theoretical time and the first actual time. The first theoretical time is the theoretical time when the first time synchronization signal arrives at the synchronization network probe device, and the first actual time is the actual time when the first time synchronization signal arrives at the synchronization network probe device.
[0100] Step K3: compensating the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model to obtain the second time synchronization signal; the reinforcement learning model is a value network and a policy network based on the first time synchronization signal deviation.
[0101] Specifically, in step K3, the first time synchronization signal is compensated according to the first time synchronization signal deviation and the reinforcement learning model to obtain the second time synchronization signal, which specifically includes the following steps:
[0102] Step K31: comparing the first time synchronization signal deviation with a preset threshold value:
[0103] When the first time synchronization signal deviation is greater than the preset threshold value, step K32 is entered; when the first time synchronization signal deviation is less than or equal to the preset threshold value, no time delay compensation is performed on the first time synchronization signal, and the first time synchronization signal is replaced by the second time synchronization signal;
[0104] Step K32: obtaining the environment state S t corresponding to the first time synchronization signal t as the input of the reinforcement learning model; the environment state includes the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation;
[0105] Step K33: outputting the action compensation in the reinforcement learning model according to the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation, to obtain the second time synchronization signal; the outputting the action compensation in the reinforcement learning model includes outputting the compensation direction and the compensation value step length in the reinforcement learning model, the compensation direction includes positive compensation, negative compensation and zero compensation, and the compensation value step length is a preset value;
[0106] In an embodiment, the reinforcement learning model is a value network and a policy network based on the first time synchronization signal deviation.
[0107] In step K33, the action compensation is outputted in the reinforcement learning model according to the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation, to obtain the second time synchronization signal, specifically including:
[0108] the environment state S t is inputted into the value network, and the value network obtains the values of each possible action of time delay compensation according to the environment state, the values of each possible action including positive compensation action, negative compensation action and zero compensation action, the sum of the probability of the positive compensation action, the probability of the negative compensation action and the probability of the zero compensation action being 1; the environment state S t includes the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation;
[0109] The policy network selects the time delay compensation according to the maximum probability value of each possible action value;
[0110] After the synchronization network transmission device receives the action, the environment state changes to the environment state S t+1 , and returns the reward value r t to the probe device;
[0111] in:
[0112] When the environment state S t+1 The absolute maximum value of the peak and trough of the TE clock performance at the current time, compared to the environmental state S at the previous time. t When the absolute values of the peak and trough values of the TE clock performance are smaller or equal,
[0113] or,
[0114] Environmental State S t+1 MTIE compared to environmental state S t The MTIE is smaller;
[0115] or,
[0116] Environmental State S t+1 TDEV ratio compared to environmental state S t The TDEV is smaller;
[0117] or,
[0118] Environmental State S t+1 When the MTIE is smaller than the value specified in the ITU-T G.811 / G.812 standard,
[0119] or,
[0120] Environmental State S t+1 The TDEV is smaller than the value specified in the ITU-T G.811 / G.812 standard.
[0121] or,
[0122] Environmental State S t+1 Frequency offset ratio of environmental state S t When the frequency offset is closer to 0,
[0123] Then the reward value r t Bonus points;
[0124] Otherwise, the reward value r t This is a deduction from the reward value.
[0125] In one implementation, in step K32, the environmental state S corresponding to the first time synchronization signal is acquired. t The environmental state corresponding to the first time synchronization signal obtained after the time range threshold X;
[0126] In step K33, based on the time error TE, maximum time interval error MTIE, time deviation TDEV, and frequency deviation, action compensation is output in the reinforcement learning model to obtain the second time synchronization signal, including the following steps:
[0127] determining a time range threshold X and a performance threshold Y; wherein: the time range threshold X represents X consecutive t time ranges, and the performance threshold Y is a threshold range of the time error TE;
[0128] calculating the average value μ and the standard deviation σ of the TE for each t time range of data;
[0129] judging whether the value of the TE is distributed within the range of (μ-3σ, μ+3σ):
[0130] if the average value of the TE in the X consecutive t time ranges is less than Y, and the value of the TE is distributed within the range of (μ-3σ, μ+3σ), then outputting the action compensation as zero compensation in the reinforcement learning model, determining that the second time synchronization signal is the same as the first time synchronization signal, and replacing the first time synchronization signal with the second time synchronization signal; if the average value of the TE in any t time range of data is greater than Y or the value of the TE is not distributed within the range of (μ-3σ, μ+3σ), then determining that the second time synchronization signal is the first time synchronization signal with positive or negative compensation.
[0131] Step K34: measuring and analyzing the deviation of the second time synchronization signal to obtain a second time synchronization signal deviation, and comparing the second time synchronization signal deviation with a preset threshold:
[0132] when the second time synchronization signal deviation is greater than the preset threshold, repeating steps K32 to K33 until the second time synchronization signal deviation is less than or equal to the preset threshold, and replacing the first time synchronization signal with the second time synchronization signal.
[0133] Step K4: sending the second time synchronization signal to the synchronization network probe device, thereby completing the delay dynamic compensation.
[0134] wherein:
[0135] The first time synchronization signal deviation is obtained by measuring and analyzing the first time synchronization signal by the synchronization network probe device, and the first time synchronization signal deviation is the difference between the first theoretical time and the first actual time. The first theoretical time is the theoretical time at which the first time synchronization signal arrives at the synchronization network probe device, and the first actual time is the actual time at which the first time synchronization signal arrives at the synchronization network probe device. Similarly, the Nth time synchronization signal deviation is obtained by measuring and analyzing the Nth time synchronization signal by the synchronization network probe device, and the Nth time synchronization signal deviation is the difference between the Nth theoretical time and the Nth actual time. The Nth theoretical time is the theoretical time at which the Nth time synchronization signal arrives at the synchronization network probe device, and the Nth actual time is the actual time at which the Nth time synchronization signal arrives at the synchronization network probe device. N is a natural number greater than 1.
[0136] Embodiment 2:
[0137] As Figure 1 , Figure 3 and Figure 4 illustrated, the embodiment provides a time delay dynamic compensation method, which is applied to a synchronous network probe device, and the method comprises the following steps:
[0138] Step S1: receiving a first time synchronization signal transmitted by a synchronous network transmission device;
[0139] Step S2: measuring and analyzing the first time synchronization signal to obtain a first time synchronization signal deviation, and sending the first time synchronization signal deviation to the synchronous network transmission device;
[0140] Specifically, the first time synchronization signal deviation is the difference between a first theoretical time and a first actual time, the first theoretical time is the theoretical time for the first time synchronization signal to arrive at the synchronous network probe device, and the first actual time is the actual time for the first time synchronization signal to arrive at the synchronous network probe device;
[0141] Step S3: receiving a second time synchronization signal transmitted by the synchronous network transmission device, thereby completing time delay dynamic compensation; the second time synchronization signal is obtained by compensating the first time synchronization signal according to the first time synchronization signal deviation and a reinforcement learning model.
[0142] Embodiment 3:
[0143] As Figure 6 illustrated, the embodiment provides a time delay dynamic compensation device, which corresponds to the method of embodiment 1, and the device is applied to a synchronous network probe device, and the device comprises:
[0144] A first receiving unit 10 is configured to receive a first time synchronization signal transmitted by a synchronous network transmission device;
[0145] A measurement and analysis unit 20 is connected to the first receiving unit 10 and is configured to measure and analyze the first time synchronization signal to obtain a first time synchronization signal deviation, and send the first time synchronization signal deviation to the synchronous network transmission device; the first time synchronization signal deviation is the difference between a first theoretical time and a first actual time, the first theoretical time is the theoretical time for the first time synchronization signal to arrive at the synchronous network probe device, and the first actual time is the actual time for the first time synchronization signal to arrive at the synchronous network probe device;
[0146] The first receiving unit 10 is further configured to receive a second time synchronization signal transmitted by the synchronous network transmission device, thereby completing time delay dynamic compensation;
[0147] The second time synchronization signal is obtained by compensating the first time synchronization signal according to the first time synchronization signal deviation and a reinforcement learning model.
[0148] Wherein, the time delay compensation is the environment state S t corresponding to the first time synchronization signal t as the input of the reinforcement learning model, and outputs the action compensation in the reinforcement learning model; the environment state includes the time error TE, the maximum time interval error MTIE, the time deviation TDEV and the frequency deviation.
[0149] Embodiment 4:
[0150] As shown in Figure 7 , the embodiment provides a time delay dynamic compensation device, applied to a synchronous network transmission device, which comprises:
[0151] The first sending unit 30 is configured to send the first time synchronization signal to the synchronous network probe device;
[0152] The second receiving unit 40 is connected with the first sending unit 30, and is configured to receive the first time synchronization signal deviation returned by the synchronous network probe device; the first time synchronization signal deviation is the difference between the first theoretical time and the first actual time, the first theoretical time is the theoretical time when the first time synchronization signal reaches the synchronous network probe device, and the first actual time is the actual time when the first time synchronization signal reaches the synchronous network probe device;
[0153] The compensation unit 50 is connected with the second receiving unit 40, and is configured to compensate the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model, to obtain the second time synchronization signal;
[0154] The second sending unit 60 is connected with the compensation unit 50, and is configured to send the second time synchronization signal to the synchronous network probe device, so as to complete the time delay dynamic compensation.
[0155] Embodiment 5:
[0156] As shown in Figure 1 , the scheme adopts the reinforcement learning technology in machine learning, adds a related decision function module in the synchronous network probe device, the synchronous network transmission device or the 5G base station device, and realizes a dynamic time delay compensation mechanism. The specific process of the time delay compensation of the synchronous network transmission device is as follows:
[0157] For the synchronous network transmission device, since it has no reference signal, it needs to be externally connected with the synchronous network probe device with a satellite antenna to measure the synchronous network clock time performance.
[0158] When the initial state is started:
[0159] 1) The clock performance state of the synchronous network transmission equipment as the environment, output the clock and time signal to the synchronous network probe equipment.
[0160] 2) The synchronous network probe equipment corresponds to the agent in the reinforcement learning, based on its own satellite signal as the reference, the clock and time signal of the synchronous network transmission equipment are measured for the performance, and the environment state s is obtained. s contains TIE, TE and other characteristics.
[0161] 3) A state analysis function is added to the probe equipment, for the environment state s, the outliers are removed, and according to the 3σ principle, the average value of the data in (μ-3σ, μ+3σ) and the absolute maximum value of the upper and lower peak values of the TE value distribution are calculated. The MTIE, TDEV, and frequency offset are calculated.
[0162] 4) On each timestamp t, the probe equipment outputs the action a t to the synchronous network transmission equipment. The compensation direction includes positive compensation, negative compensation and no compensation. The compensation value step can be set each time, for example, 1 ns.
[0163] 5) The synchronous network transmission equipment receives the action at time t, and performs corresponding delay compensation on the input port of the synchronous network transmission equipment.
[0164] 6) The state of the synchronous network transmission equipment changes to s t+1 after receiving the action, and the reward value r t is returned.
[0165] 7) The reward value r t returned by the synchronous network transmission equipment to the probe equipment. When the absolute maximum value of the upper and lower peak values of the TE clock performance of s t+1 is smaller or equal to the absolute maximum value of the upper and lower peak values of the TE of s t , or the MTIE and TDEV of s t+1 are smaller than the MTIE and TDEV of s t , or the frequency offset of s t+1 is closer to 0 than the frequency offset of s t , the reward value r t is the plus reward value, otherwise the reward value r t is the minus reward value.
[0166] The most critical link is the strategy of action decision. Here, a neural network is created, which takes the environment state s as input, containing four characteristics: TE, MTIE, TDEV, and frequency offset, i.e. a vector of length 4. The action a is taken as output. The probability of each action is estimated, and a action is randomly selected according to the estimated probability. The probability distribution of the action is π θ(a|s), where θ is the parameter of the policy function π, which can be parameterized using a neural network. θ Functions. For example... Figure 3 As shown. The sum of the probabilities of all actions is 1, i.e., ∑ a∈A πθ(a|s) = 1, where A is the set of all actions. θ A network representing the agent's policies is called a policy network, such as... Figure 3 As shown.
[0167] The neural network consists of an input layer, multiple fully connected hidden layers, and an output layer. The output has three nodes, representing the probability distribution of the three actions. Gradient descent is used to optimize the network, with the goal of increasing the probability of a positive reward and decreasing the probability of a negative reward. During the interaction, the action 'a' with the highest probability is selected. t =argmax a πθ(a|st) is used as the decision result in the time delay compensation environment.
[0168] During the stable period after compensation, the clock performance normally changes only slightly, and there is no need to perform dynamic compensation every second. For this purpose, a time range threshold X and a performance threshold Y can be set. When the average value of the TE value distributed in (μ-3σ,μ+3σ) is less than Y within X consecutive time ranges t, dynamic compensation is not performed; otherwise, dynamic compensation is restarted.
[0169] Specifically, the 1588v2 topology diagram is as follows: Figure 2 As shown, the end-of-line transmission equipment of the synchronization network outputs clock and time signals to the 5G base station equipment to meet the high-precision synchronization requirements of 5G services. Simultaneously, it can also output clock and time signals to the synchronization network probe equipment for signal accuracy monitoring and measurement. Output clock signals include, but are not limited to, Synchronization Ethernet (SyncE), 2 Mbit / s, 2 MHz, and 10 MHz signals, while output time signals include, but are not limited to, 1 PPS, 1 PPS+ToD, and PTP signals.
[0170] Synchronous network probe equipment and 5G base station equipment that support measurement functions can measure the output signals of terminal equipment based on BeiDou / GPS satellite reference signals, including measuring time interval error (TIE) and time error (TE), and calculating maximum time interval error (MTIE), time deviation (TDEV), frequency offset, etc. based on TIE.
[0171] like Figure 4As shown, in this embodiment, the time delay compensation of the 5G base station device is as follows:
[0172] For the 5G device, since the 5G device itself has the condition of installing a satellite antenna, that is, it has the ability to use satellite signals as a reference to measure the performance of the input signal, the 5G base station device itself can integrate the environment and the agent. As shown Figure 4 The running steps are the same as described above.
[0173] For example, when the TE deviation range of the 5G base station clock / time input signal itself is 100-160ns. When X is set to 10 and the performance threshold Y is set to 10ns. Since the average value 130ns exceeds the threshold, the dynamic compensation mechanism is run. The neural network is trained through reinforcement learning, and the first action is output. If the action is to compensate 1ns, the compensated state becomes 101-161ns, and the absolute value is larger than the last maximum value of 160 to 161, so the penalty value is reduced. If the action is to compensate -1ns, the compensated state becomes 99-159ns, and the penalty value is increased. In this way, the final dynamic compensation is -130ns, and the signal deviation range is -30ns-30ns. At this time, the steady state is reached, and only monitoring is performed without compensation. When the network structure suddenly changes, such as interface switching, failure, etc., the signal deviation exceeds the threshold, and dynamic compensation is performed again, thereby solving the problem of dynamic time delay compensation.
[0174] Embodiment 5 is corresponding to embodiments 1-4.
[0175] Embodiment 6:
[0176] Based on the same technical concept, the embodiment provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the time delay dynamic compensation method according to embodiment 1.
[0177] Embodiment 7:
[0178] Based on the same technical concept, the embodiment provides a computer readable storage medium, which stores a computer program, and when the processor executes the computer program, the processor executes the time delay dynamic compensation method according to embodiment 1.
[0179] Embodiment 8:
[0180] The embodiment provides a time delay dynamic compensation system, which comprises a synchronization network probe device and a synchronization network transmission device,
[0181] The synchronization network transmission device transmits a first time synchronization signal to the synchronization network probe device;
[0182] The synchronization network probe device sends a first time synchronization signal deviation to the synchronization network transmission device;
[0183] The synchronization network probe device receives a second time synchronization signal transmitted by the synchronization network transmission device; the second time synchronization signal is obtained by compensating the first time synchronization signal according to the first time synchronization signal deviation and the reinforcement learning model.
[0184] Wherein:
[0185] The first time synchronization signal deviation is obtained by measuring and deviation analyzing the first time synchronization signal by the synchronization network probe device, the first time synchronization signal deviation is the difference between the first theoretical time and the first actual time, the first theoretical time is the theoretical time when the first time synchronization signal reaches the synchronization network probe device, and the first actual time is the actual time when the first time synchronization signal reaches the synchronization network probe device.
[0186] It can be understood that the above embodiments are only exemplary embodiments adopted for illustrating the principles of the present application, and the present application is not limited thereto. Various modifications and improvements can be made by those skilled in the art without departing from the spirit and essence of the present application, and these modifications and improvements are also considered as the protection scope of the present application.
Claims
1. A method for compensating time delay, applied to a synchronous network transmission device, characterized in that, The method comprises the following steps: Step K1: sending a current time synchronization signal to a synchronization network probe device to make the synchronization network probe device calculate a current compensation action; wherein the compensation action is calculated by the synchronization network probe device based on a measured synchronization signal deviation and in combination with a reinforcement learning model; the synchronization signal deviation is the difference between the theoretical time and the actual time of the current time synchronization signal arriving at the synchronization network probe device; the reinforcement learning model receives a current environment state as input; the current environment state comprises a time error TE, a maximum time interval error MTIE, a time deviation TDEV, and a frequency deviation; the compensation action comprises a compensation direction and a compensation value step, the compensation direction comprises positive compensation, negative compensation, and zero compensation, and the compensation value step is a preset value; Step K2: receiving a compensation action from the synchronization network probe device; Step K3: performing time delay compensation on the current time synchronization signal according to the compensation action to generate a compensated time synchronization signal; Step K4: sending the compensated current time synchronization signal to the synchronization network probe device to make the synchronization network probe device take the compensated current time synchronization signal as a new current time synchronization signal to calculate a next compensation action; Repeat steps K2 to K4 until the synchronization network probe device confirms that the synchronization signal deviation is less than or equal to a preset threshold.
2. A method for time delay dynamic compensation applied to a network probe device, characterized in that, The method comprises the following steps: Step S1: receiving a current time synchronization signal sent by a synchronization network transmission device; Step S2: measuring a synchronization signal deviation of the current time synchronization signal; the synchronization signal deviation is the difference between the theoretical time and the actual time of the current time synchronization signal arriving at the synchronization network probe device; Step S3: determining whether the synchronization signal deviation exceeds a preset threshold; If the synchronization signal deviation exceeds the preset threshold, a current environment state is obtained, and the current environment state is input into a pre-trained reinforcement learning model to output a compensation action; Wherein, the current environment state comprises a time error TE, a maximum time interval error MTIE, a time deviation TDEV, and a frequency deviation; the compensation action comprises a compensation direction and a compensation value step, the compensation direction comprises positive compensation, negative compensation, and zero compensation, and the compensation value step is a preset value; Step S4: sending the compensation action to the synchronization network transmission device to make the synchronization network transmission device perform time delay compensation on the current time synchronization signal according to the compensation action to generate a compensated time synchronization signal; Step S5: receiving the compensated time synchronization signal sent by the synchronization network transmission device and taking the compensated time synchronization signal as a new current time synchronization signal; Repeat steps S2 to S5 until the measured synchronization signal deviation is less than or equal to the preset threshold.
3. The time delay dynamic compensation method according to claim 2, wherein: The reinforcement learning model is a value network and a policy network based on the time synchronization signal deviation; The compensation action specifically comprises: corresponding to the first time synchronization signal t a value network is input, the value network obtains time delay compensation for each possible action value according to the environment state, each possible action value includes a positive compensation action, a negative compensation action and a zero compensation action, the sum of the probability of the positive compensation action, the probability of the negative compensation action and the probability of the zero compensation action is 1; the environment state S t includes time error TE, maximum time interval error MTIE, time deviation TDEV and frequency deviation; The policy network selects the time delay compensation according to the maximum probability value of each possible action value; The synchronous network transmission device receives the action, and the environment state S t is changed to the environment state S t+1 , and returns the reward value r t to the probe device. Wherein: When the absolute value of the up-and-down peak of the TE clock performance at the environmental state S t+1 is greater than the absolute value of the up-and-down peak of the TE clock performance at the environmental state S t is less than or equal to the absolute value of the up-and-down peak of the TE clock performance at the environmental state S Or, environmental state S t+1 MTIE of the environmental state S t is smaller than the MTIE of the environmental state S Or, environmental state S t+1 TDEV of the environmental state S t TDEV of the environmental state S Or, environmental state S t+1 MTIE is less than the value specified in the ITU-T G. 811 / G. 812 standard, Or, environmental state S t+1 TDEV is less than the value specified in the ITU-T G. 811 / G. 812 standard, Or, environmental state S t+1 the frequency offset of the environmental state S t is closer to 0, r = r + 1 t r = r + 1 Otherwise the reward value r t is a penalty value.
4. The time delay dynamic compensation method according to claim 2, characterized in that, The judgment of whether the synchronization signal deviation exceeds the preset threshold value specifically comprises: Determine a time range threshold X and determine a performance threshold Y; wherein: the time range threshold X represents continuous X t time, and the performance threshold Y is the threshold range of the time error TE; For each t time range data, calculate the average value μ and the standard deviation σ of TE; Judge whether the value of TE is distributed in the range of (μ-3σ, μ+3σ): If the average value of TE in the continuous X t time range is less than Y, and the value of TE is distributed in the range of (μ-3σ, μ+3σ), it is determined that the synchronization signal deviation does not exceed the preset threshold value; if the average value of TE of any t time range data is greater than Y or the value of TE is not distributed in the range of (μ-3σ, μ+3σ), it is determined that the synchronization signal deviation exceeds the preset threshold value.
5. An electronic device, comprising: The computer program is executed by the processor, and the processor executes the time delay dynamic compensation method according to any one of claims 1 or 2 to 4.
6. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor, and the processor executes the time delay dynamic compensation method according to any one of claims 1 or 2 to 4.
7. A time delay dynamic compensation system characterized by, The synchronization network probe device and the synchronization network transmission device are included, The synchronization network transmission device transmits the current time synchronization signal to the synchronization network probe device; The synchronization network probe device sends a compensation action to the synchronization network transmission device; wherein the compensation action is calculated and generated by the synchronization network probe device based on the measured synchronization signal deviation and via a reinforcement learning model, and the synchronization signal deviation is the difference between the theoretical time and the actual time of the current time synchronization signal reaching the synchronization network probe device; The synchronization network transmission device compensates the current time synchronization signal based on the received compensation action to obtain a compensated time synchronization signal; The compensated time synchronization signal is sent to the synchronization network probe device, thereby completing the time delay dynamic compensation; Wherein, the reinforcement learning model receives the current environment state as input; the current environment state includes time error TE, maximum time interval error MTIE, time deviation TDEV and frequency deviation; the compensation action includes compensation direction and compensation value step, the compensation direction includes positive compensation, negative compensation and zero compensation, and the compensation value step is a preset value.
Citation Information
Patent Citations
Satellite communication ground synchronization simulation system based on IEEE 1588v2 and application method thereof
CN107947848A
Later 5G fronthaul network time synchronization method and device based on deep reinforcement learning
CN110896556A