Value-based optical cable state sensing and energy consumption optimization method

CN120343678APending Publication Date: 2025-07-18FOSHAN GUYUXUAN BRAND MANAGEMENT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510260347.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing optical cable status monitoring methods lack effective technical support, resulting in delayed detection of optical cable impregnation and icing faults, and maintenance is time-consuming and labor-intensive. Traditional methods cannot meet the needs of optical cable status monitoring.

Method used

Build a digital twin architecture based on value-based multi-agent body reinforcement learning, and realize optical cable state perception and energy consumption optimization through the collaborative work of the main agent and the perceptual agent, and use the meta-reinforcement learning algorithm to optimize the accuracy and energy consumption control of the digital twin architecture.

Benefits of technology

Real-time monitoring and accurate prediction of optical cable status are realized, the system's adaptability is improved, manual intervention is reduced, and the efficiency and energy consumption optimization effect of optical cable status monitoring are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343678A_ABST
    Figure CN120343678A_ABST
Patent Text Reader

Abstract

The invention discloses a value-based optical cable state sensing and energy consumption optimization method. According to the method, a digital twinning architecture composed of a main agent and a sensing agent is constructed, the main agent obtains the state of an optical cable through interaction with a physical environment, the sensing agent is responsible for observing the environment and communicating with a wireless access point, the precision of the digital twinning architecture is continuously optimized through a meta reinforcement learning algorithm, and the accuracy of the digital twinning architecture is improved. And scheduling and energy consumption control are carried out on the sensing intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electric power communication, and in particular relates to a value-based optical cable state perception and energy consumption optimization method. Background Art

[0002] With the rapid development of my country's smart grid with UHV power grid as the backbone, the information, automation and interactive business needs of smart grids have promoted the rapid development of power system communications with fiber-optic communication as the core. As the physical layer of fiber-optic communication, optical cables cover a very wide range. At present, the optical cables of the entire power system have exceeded 1 million kilometers. Therefore, the operational risks of the communication system that may be caused by optical cable failures have greatly increased. Among them, the immersion and freezing of optical cables is the pain point, blocking point and difficulty in the maintenance of the power grid optical fiber communication line system, and has become a bottleneck for the efficient and reliable operation of the optical fiber communication network and the improvement of the accessibility rate. The problem needs to be solved urgently, and the communication needs are realistic and urgent.

[0003] However, the current maintenance and inspection of optical cable lines is still mainly done manually, lacking effective technical support. Usually, the spare optical fiber is tested regularly or irregularly, and the loss changes at each fusion point of the spare optical fiber in the relay section are compared to determine whether the optical cable is soaked in water or even frozen. However, since this method is a passive protection, it can only be found after the optical cable is severely soaked in water or even iced to damage the fiber core. There may even be a problem that the service fiber core is blocked but the spare fiber core is good and cannot be found. On the other hand, some optical cable large attenuation points found through the optical fiber reflection test curve are often time-consuming and labor-intensive to find and judge the fault points on site in actual maintenance and repair work, due to the huge differences in the mileage of the completion data and the actual geographical location.

[0004] Therefore, considering the difficulty of inspecting optical cable lines for water immersion and icing, traditional optical cable status monitoring methods obviously cannot meet the needs well. Summary of the invention

[0005] In response to the above problems, the inventors found that the multi-agent system uses multiple interacting agents, each of which can independently perceive the environment and make decisions. For the problem of optical cable state perception and energy consumption optimization, multiple agents can be responsible for monitoring the state of optical cables in different areas, or can work together and share tasks in the process of energy consumption optimization; value-based meta-reinforcement learning is reinforcement learning, which guides the behavior of agents by estimating the "value" of a certain state or action, and in the context of optical cable state perception and energy consumption optimization, meta-reinforcement learning can help agents quickly adapt to new network environments or optical cable failure modes, thereby improving the system's adaptive ability.

[0006] Based on this, the present invention proposes a value-based optical cable status perception and energy consumption optimization method. Specifically, it can be said to be a value-based multi-agent meta-reinforcement learning optical cable status perception and energy consumption optimization method. This method uses artificial intelligence algorithms to improve the efficiency of optical cable monitoring while optimizing energy consumption. Specifically, this method constructs a digital twin architecture composed of a main agent and perception agents. The main agent obtains the status of the optical cable through interaction with the physical environment, and the perception agents are responsible for observing the environment and communicating with wireless access points. Among them, the meta-reinforcement learning algorithm is used to continuously optimize the accuracy of the digital twin architecture and perform scheduling and energy consumption control on the perception agents.

[0007] The embodiments of the present invention provide the following technical solutions: Step A, construct a digital twin DT architecture for optical cable status perception. This architecture consists of a single main agent PA and a group of perception agents SA working together. The SA observes the environmental status and transmits data to the access point and PA to support the update of the DT model. In each cycle, the PA updates the multi-dimensional system status based on the dynamic state equation , while the SA is scheduled by the PA through the observation equation. This step establishes the mathematical basis for the DT model to dynamically track the physical environment through the above definitions.

[0008] Step B, establish an optimization process for the main agent based on the value function. With the dual goals of maximizing the state value function of the PA and minimizing the transmission energy consumption of the SA, calculate the estimated value of the initial state of the PA of the estimated discounted value function under the policy , and determine the optimized reward function based on the above definitions. This step integrates the state estimation accuracy and energy consumption optimization through the meta-reinforcement learning framework, providing a quantitative evaluation and iterative optimization path for dynamic decision-making.

[0009] Step C, the optimization process of dynamically scheduling the perception agents. First, initialize the set of available SAs and the set to be scheduled, and introduce an error level variable to constrain the accuracy requirements of the SAs. Subsequently, construct a transmission energy consumption allocation vector and design a joint optimization function, with the goal of balancing error control and energy consumption minimization under the premise of meeting the channel capacity limit. This step realizes high-efficiency scheduling under the cooperation of multiple SAs through a dynamic resource allocation and constraint satisfaction mechanism.

[0010] Step D, channel modeling and power optimization strategy for SA communication. Derive the instantaneous signal-to-noise ratio expression and its probability density function, which characterize the combined effects of line-of-sight transmission and small-scale fading. Combine Shannon's theorem to establish a channel capacity model and analyze the optimal power allocation solution that meets the transmission requirements. This step provides a theoretical basis for power configuration for the reliable data transmission of the SA through strict mathematical modeling and communication theory analysis.

[0011] Among them, step A specifically includes: A1. The digital twin DT model architecture consists of a single primary agent PA and a group of sensing agents SA. These sensing agents are responsible for observing the environment and communicating with the wireless access point through the wireless channel to facilitate the construction of the DT model of PA.

[0012] A2. At the beginning of each cycle, the DT model needs to update the state of SA, then estimate the state of the entire system, update the policy, calculate the optimal control signal, schedule SA, and apply the fusion algorithm on the wireless access point to execute the control command on SA.

[0013] A3. Define the time sequence of each cycle as , so the state obtained by PA interacting with the environment at the -th cycle is a -dimensional state vector. When the state includes the conditions of optical cable water immersion, icing, temperature, humidity, stress and breakage, at this time , its evolution process is described as:

[0014] Among them represents the state update function, describing how the state at the previous moment changes to the state at the current moment. The matrix represents the influence of the control signal on SA, is the control signal, the process noise , which follows a distribution with a mean of 0 and a variance of .

[0015] A4. When SA receives the state of PA at the moment, the observed value of SA is:

[0016] Among them is the observation matrix, is the observation noise, with a mean of 0 and a variance of . In addition, the variance of is defined as , and the maximum allowable standard deviation of the SA observed value .

[0017] The purpose of the DT model is to maintain an accurate estimate of the state of PA. Therefore, the estimated value of the state is expressed as .

[0018] Among them, step B specifically includes: B1. The value function of PA in the initial state , under the parameter and the control strategy is .

[0019] B2. Define the transmission energy consumption when SA transmits the observation value to the wireless access point at the -th loop as

[0020] B3. Therefore, the optimization objective is expressed as maximizing while minimizing .

[0021] B4. Abstract the problem in B3 as a Markov decision process, represented by the tuple , where is a finite set of possible states, is a set of control actions, is a set of possible observations, is the probability that the agent takes the action and its state transfers from to , is the probability of observing from the system, is the reward received when the agent takes the action and its state transfers from to , is the discount factor.

[0022] B5. The estimated value of the initial state of PA under the policy has the following estimated discounted value function:

[0023] where is the discount factor and is the mean calculation.

[0024] B6. The estimated state of PA at the -th loop is , and the conditional probability distribution function of is expressed as , where is the parameterization of the precision vector, and each element represents the accuracy of the estimated state at the -th process. Therefore, as increases, the estimated state The accuracy will also increase, which helps the agent make correct decisions.

[0025] B7, so in the th cycle, the reward function is defined as , where is the reward for the th time, where is the weight coefficient, represents the non-increasing cost function.

[0026] B8, the parameter update strategy is: , where is 's previous state, where is the meta-learning rate, is the gradient function.

[0027] Among them, step C specifically includes: C1, define the available SA set as , the scheduled SA set is , at the beginning of the th cycle, let represent that the total number of available SAs is ones, represents that the scheduled SA set is an empty set.

[0028] C2, introduce any variable to represent the error level of the DT expectation at the th cycle of the system. Therefore, the error level that SA should satisfy at the th cycle is .

[0029] C3, define the SA transmission energy consumption allocation vector as , and .

[0030] C4, so the optimization function of SA is expressed as: , , is the maximum number of SAs that can be scheduled when reaching the channel capacity limit.

[0031] C5, use the state update formula of A3 and the SA observation value formula of A4 to iterate cyclically to meet the conditions of C4.

[0032] Among them, step D specifically includes: D1, model the channel between the th SA and the AP as a channel with a strong line-of-sight transmission link and small-scale fading, where the instantaneous signal-to-noise ratio is modeled as: , , where is a constant, depending on system parameters (operating frequency and antenna gain), represents SA the distance between and AP, is the path loss factor, is the transmission noise, is the bandwidth, is the fading power.

[0033] D2, defined as follows the Rice distribution, so its probability density function is expressed as: , where is the Rice factor, is the constant coefficient, represents the modified Bessel function of the first kind of order zero, is the mean calculation.

[0034] D3, according to Shannon's theorem, the channel capacity can be obtained , and further the optimal allocation power for the th SA to meet the C4 condition is represented by , , where , is the length of the transmission data packet, is the inverse function, is the block error probability.

[0035] Compared with the prior art, the above technical solution has the following advantages: The present invention uses a digital twin architecture to achieve a high degree of mapping and collaborative work between the physical environment and the virtual model, can take into account multiple state dimensions (including but not limited to optical cable water immersion, icing, temperature, humidity, stress, breakage conditions), and perform real-time perception and monitoring of the optical cable state, so as to improve the accuracy of the system and achieve precise monitoring, prediction and optimization; by constructing a multi-agent system composed of a main agent and a sensing agent, the division of labor and cooperation of tasks are realized. The main agent is responsible for interacting with the physical environment and making decisions, while the sensing agent provides real-time optical cable state data by monitoring the surrounding environment and communicating with the wireless access point. The distributed role of the sensing agent enables the system to work in parallel in different regions, thus achieving the high efficiency and robustness of the system; the meta-reinforcement learning algorithm optimizes the accuracy of the digital twin architecture, and at the same time schedules the sensing agent and performs energy consumption control. The introduction of meta-reinforcement learning enables the system to have the ability to learn how to quickly adapt to new environments and new tasks, and can continuously improve the policy optimization efficiency by accumulating experience and reduce manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 It is a schematic flow chart of a value-based optical cable status perception and energy consumption optimization method provided by an embodiment of the present invention. Specific embodiments

[0038] As described in the background art section, how to meet the overall performance of communication perception and calculation in the process of comprehensively improving the optical cable status monitoring to ensure the normal operation of the power system is an urgent problem for those skilled in the art.

[0039] The core idea of the present invention is that in the monitoring of the optical cable status, the digital twin architecture is used to monitor the optical cable status in real time, and at the same time, the value-based meta-reinforcement learning algorithm is used to optimize the system perception accuracy and energy consumption control.

[0040] See Figure 1 , an embodiment of the present invention provides a value-based optical cable status perception and energy consumption optimization method, and the method includes: Step A, construct a digital twin (Digital Twin, abbreviated as DT) architecture for optical cable status perception. This architecture consists of a single primary agent (Primary Agent, abbreviated as PA) and a group of sensing agents (Sensing Agent, abbreviated as SA) working together. The SA observes the environmental state and transmits data to the access point and the PA to support the update of the DT model; in each cycle, the PA updates the multi-dimensional system state based on the dynamic state equation , and at the same time, the SA is scheduled by the PA through the observation equation. Through the above definition, this step establishes the mathematical basis for the DT model to dynamically track the physical environment.

[0041] Step B, establish an optimization process for the primary agent based on the value function. With the dual goals of maximizing the PA state value function and minimizing the SA transmission energy consumption, calculate the estimated value of the initial state of the PA of the estimated discounted value function under the policy , and determine the optimized reward function based on the above definition. This step integrates the state estimation accuracy and energy consumption optimization through the meta-reinforcement learning framework, providing a quantitative evaluation and iterative optimization path for dynamic decision-making.

[0042] Step C, the optimization process of the dynamic scheduling-aware agent. First, initialize the available SA set and the set to be scheduled, and introduce an error level variable to constrain the accuracy requirements of the SA; subsequently, construct a transmission energy consumption allocation vector and design a joint optimization function, with the goal of balancing error control and energy consumption minimization while meeting the channel capacity limit. This step realizes high-efficiency scheduling under the cooperation of multiple SAs through a dynamic resource allocation and constraint satisfaction mechanism.

[0043] Step D, the channel modeling and power optimization strategy for SA communication. The instantaneous signal-to-noise ratio expression and its probability density function are derived, which characterize the combined influence of line-of-sight transmission and small-scale fading; a channel capacity model is established in combination with Shannon's theorem, and the optimal power allocation solution to meet the transmission requirements is analyzed. This step provides a theoretical basis for power configuration for the reliable data transmission of the SA through strict mathematical modeling and communication theory analysis.

[0044] Among them, step A specifically includes: A1, the digital twin DT model architecture consists of a single primary agent PA and a group of sensing agents SA. These sensing agents are responsible for observing the environment, communicating with the wireless access point through the wireless channel, and further transmitting the data to the primary agent PA for the construction and maintenance of the DT model by the PA. The PA predicts the state of the optical cable (such as fault prediction, production optimization) through the data received from the SA and schedules the SA subsequently. This design not only realizes the dynamic perception and synchronization of the environmental state but also provides a reliable data basis for the PA to construct a high-precision and real-time updated digital twin model, thus supporting intelligent decision-making and optimization in complex scenarios.

[0045] A2, the sensing agents SA perform sensing in a collaborative manner to improve the sensing accuracy, expand the sensing range, and enhance the robustness. Therefore, at the beginning of each cycle, the DT model needs to update the state of the SA, then estimate the overall system state, update the strategy, calculate the optimal control signal, schedule the SA, and apply a fusion algorithm at the wireless access point to execute the control command on the SA.

[0046] A3, define the time series that occurs in each cycle as , so at the th cycle, the PA interacts with the environment to obtain the state of the system is a -dimensional state. When the state includes the water immersion, icing, temperature, humidity, stress, and breakage conditions of the optical cable, at this time , its evolution process is described as:

[0047] where represents the state update function, which describes the previous moment How to transform into the state at the current moment , the matrix represents the influence of the control signal on SA, is the control signal, and the process noise , which follows a distribution with a mean of 0 and a variance of .

[0048] A4. When SA receives the state of PA at time (SA receives the state of PA aiming to accept the scheduling of PA), the observed value of SA at this time is:

[0049] where is the observation matrix, is the observation noise, with a mean of 0 and a variance of . Additionally, the variance of is defined as .

[0050] A5. The purpose of the DT model is to maintain an accurate estimate of the PA state. Therefore, the estimated value of the state is represented as .

[0051] Among them, step B specifically includes: B1. To achieve the purpose of the DT model, it is necessary to determine the optimization goal of the system. Therefore, the value function of PA in the initial state , under the parameter and the control strategy is defined as .

[0052] B2. Define the transmission energy consumption when SA transmits the observed value to the wireless access point at the th cycle as .

[0053] B3. Therefore, the optimization goal is expressed as maximizing while minimizing . The purpose is to ensure the maximization of the PA state estimation accuracy and minimize the transmission energy consumption when scheduling SA, that is, to improve the system accuracy while reducing the energy consumption.

[0054] B4. Abstract the problem in B3 into a Markov decision process, which is represented by the tuple , where is a finite set of possible states, is a set of control actions, is a set of possible observed values, is the agent taking the action The probability that its state transfers from to is the probability of observing from the system, and the probability that the agent takes action and its state transfers from to is the received reward, where is the discount factor.

[0055] B5. The estimated value of the initial state of PA under the policy is the estimated discounted value function: where

[0056] and is the discount factor, and

[0057] B6. At the -th iteration, the estimated state of PA is , and The conditional probability distribution function of is denoted as where is the parameterization of the precision vector, and each element represents the accuracy of the -th process of the estimated state . Therefore, as increases, the accuracy of the estimated state

[0058] B7. Thus, at the -th iteration, the reward function is defined as where is the reward at the -th time, is the weight coefficient, represents a non-increasing cost function.

[0059] B8. The parameter updates the policy as: where is the previous state, is the meta learning rate, is the gradient function.

[0060] Among them, step C specifically includes: C1. Define the available SA set as , the scheduled SA set as At the At the start of the next cycle, let represent that the total number of available SAs is , indicating that the set of scheduled SAs is an empty set.

[0061] C2. Introduce any variable representing the expected error level of DT at the -th cycle of the system. Therefore, the error level that SA should meet at the -th cycle is .

[0062] C3. Define the SA transmission energy consumption allocation vector as , and .

[0063] C4. Therefore, the optimization function of SA is expressed as: , , is the maximum number of SAs that can be scheduled to reach the channel capacity limit. The purpose of the optimization function is to minimize the weighted sum of the error level of SA and the transmission energy consumption of SA, that is, to reduce the energy consumption while meeting the sensing accuracy of SA.

[0064] C5. Use the state update formula of A3 and the SA observation value formula of A4 to perform iterative loops to meet the conditions of C4.

[0065] Among them, step D specifically includes: D1. Since the data transmission process between SA and PA will be affected by potential delays, which is affected by channel quality and scheduled transmission resources, channel modeling is required. Model the channel between the -th SA and the AP as a channel with a strong line-of-sight transmission link and small-scale fading, where the instantaneous signal-to-noise ratio is modeled as: , , where is a constant, depending on system parameters (operating frequency and antenna gain), represents the distance between SA and the AP, is the path loss factor, is the transmission noise, is the bandwidth, is the fading power.

[0066] D2. Define obeys the Rice distribution, so its probability density function is expressed as: , where is the Rice factor, is a constant coefficient, denotes the first kind of zero-order modified Bessel function, is for mean calculation.

[0067] D3, the channel capacity can be obtained according to Shannon's theorem , at this time according to the requirement of the system outage probability, i.e., the channel capacity shall not be less than the minimum rate of the system ( ), so further obtain the optimal allocated power for the th SA to meet the C4 condition is denoted by , , where , is the length of the transmission data packet, is the inverse function, is the block error probability. At this time, the optimal allocated power of the SA is obtained, which optimizes the scheduling and transmission energy consumption of the SA while ensuring the accuracy of the system optical cable state estimation, and realizes efficient and intelligent decision-making and optimization.

[0068] Compared with the prior art, the above technical solution has the following advantages: The present invention uses a digital twin architecture to achieve a high degree of mapping and collaborative work between the physical environment and the virtual model, and can take into account multiple key dimensions - including but not limited to optical cable immersion, icing, temperature, humidity, stress, and breakage conditions, to perform real-time perception and monitoring of the optical cable state, thereby improving the accuracy of the system to achieve precise monitoring, prediction, and optimization; by constructing a multi-agent system composed of a main agent and a sensing agent, the division of labor and cooperation of tasks are realized. The main agent is responsible for interacting with the physical environment and making decisions, while the sensing agent provides real-time optical cable state data by monitoring the surrounding environment and communicating with wireless access points. The distributed role of the sensing agent enables the system to work in parallel in different regions, thus achieving the efficiency and robustness of the system; the meta-reinforcement learning algorithm optimizes the accuracy of the digital twin architecture, while scheduling the sensing agent and performing energy consumption control. The introduction of meta-reinforcement learning enables the system to have the ability to learn how to quickly adapt to new environments and new tasks, and can continuously improve the strategy optimization efficiency by accumulating experience and reduce manual intervention.

[0069] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A value-based optical cable status perception and energy consumption optimization method, characterized in that It includes the following steps: Step A: Build a digital twin (DT) architecture for optical cable status perception. This architecture consists of a single primary agent (PA) and a group of sensing agents (SA) working in collaboration. The SA observes the environmental status and transmits data to the access point and the PA to support the update of the DT model. In each cycle, the PA updates the multi-dimensional system status based on the dynamic state equation. , and at the same time, the SA is scheduled by the PA through the observation equation to establish the mathematical basis for the DT model to dynamically track the physical environment. Step B: Establish the optimization process of the main agent based on the value function, with the dual objectives of maximizing the PA state value function and minimizing the SA transmission energy consumption, and calculate the estimated value of the initial state of PA of the estimated discounted value function under the policy , determine the optimized reward function, fuse the state estimation accuracy and energy consumption optimization through the meta-reinforcement learning framework, and provide a quantitative evaluation and iterative optimization path for dynamic decision-making; Step C, the optimization process of the dynamic scheduling-aware agent. First, initialize the available SA set and the set to be scheduled, and introduce an error level variable to constrain the accuracy requirements of the SA. Subsequently, construct a transmission energy consumption allocation vector and design a joint optimization function. The goal is to balance error control and energy consumption minimization while satisfying the channel capacity limit, and achieve high-performance scheduling under the cooperation of multiple SAs through a dynamic resource allocation and constraint satisfaction mechanism; Step D, the channel modeling and power optimization strategy for SA communication. Derive the instantaneous signal-to-noise ratio expression and its probability density function, and characterize the combined effects of line-of-sight transmission and small-scale fading. Establish a channel capacity model in combination with Shannon's theorem, and analyze the optimal power allocation solution that meets the transmission requirements. Through strict mathematical modeling and communication theory analysis, provide a theoretical basis for power configuration for the reliable data transmission of the SA.

2. The value-based optical cable status perception and energy consumption optimization method according to claim 1, wherein Step A specifically includes: A1, the digital twin DT model architecture consists of a single primary agent PA and a group of sensing agents SA. These sensing agents are responsible for observing the environment and communicating with the wireless access point through the wireless channel to facilitate the construction of the DT model of the PA; A2, at the beginning of each cycle, the DT model needs to update the state of the SA, then estimate the overall system state, update the strategy, calculate the optimal control signal, schedule the SA, and apply a fusion algorithm at the wireless access point to execute control commands on the SA; A3, define the time series when each cycle occurs as , so at the -th cycle, PA interacts with the environment to obtain the state of the system which is a -dimensional state vector. When the state includes the cases of optical cable water ingress, icing, temperature, humidity, stress and damage, at this time , its evolution process is described as: ; where represents the state update function, which describes how the state at the previous moment transforms into the state at the current moment , and the matrix represents the influence of the control signal on SA,[[]] is the control signal, and the process noise obeys a distribution with a mean of 0 and a variance of ; A4, when SA receives the state of PA at the moment, the observed value of SA is: ; wherein is the observation matrix, is the observation noise, with a mean of 0 and a variance of ; additionally, the variance of is defined as ; the maximum allowable standard deviation of the SA observation value . For A5, the purpose of the DT model is to maintain an accurate estimate of the PA state, so the estimated value of the state is represented as .

3. The value-based optical cable status perception and energy consumption optimization method according to claim 1, characterized in that Step B specifically includes: B1, defined in the control strategy , parameter Under the condition that PA is in the initial state , the value function is ; B2, define the transmission energy consumption when SA transmits the observed value to the wireless access point in the th cycle as B3, so the optimization objective is expressed as maximizing while minimizing ; B4 abstracts the problem of B3 into a Markov decision process, represented by the tuple where is a finite set of possible states, is a set of control actions, is a set of possible observations, is the probability that the agent takes action and its state transfers from to ; is the probability of observing from the system; is the reward received when the agent takes action and its state transfers from to ; is the discount factor; B5, PA initial state Estimated value Under the policy The estimated discounted value function is as follows: ; wherein is a discount factor, is for mean calculation; B6, at the th iteration, the estimated state of PA is , and 's conditional probability distribution function is expressed as , where is the parameterization of the precision vector, and each element represents the estimated state of the th process. Therefore, as increases, the accuracy of the estimated state also increases, which helps the agent make correct decisions; B7, so in the th cycle, the reward function is defined as , where is the reward for the th time, is the weight coefficient, represents a non-increasing cost function; B8, parameter The update strategy is as follows: , where is the previous state of is the meta learning rate, is the gradient function.

4. The value-based optical cable status perception and energy consumption optimization method according to claim 1, wherein Step C specifically includes: C1, define the available SA set as , and the scheduled SA set is , at the start of the -th loop, let denote that the total number of available SAs is ; denote that the scheduled SA set is an empty set; C2, introduce any variable Indicates the expected error level of DT at the th cycle. Therefore, the error level that SA should satisfy at the th cycle is ; Define the SA transmission energy consumption allocation vector as and ; C4, so the optimization function of SA is expressed as: , , is the maximum number of SAs that can be scheduled when the channel capacity limit is reached; C5, use the state update formula of A3 and the SA observation value formula of A4 to iterate cyclically to meet the conditions of C4.

5. The value-based optical cable status perception and energy consumption optimization method according to claim 1, wherein Step D specifically includes: D1 models the channel between the th SA and the AP as a channel with a strong line-of-sight transmission link and small-scale fading, where the instantaneous signal-to-noise ratio is modeled as: , , where is a constant that depends on system parameters, and the system parameters include the operating frequency and antenna gain. represents the distance between the th SA and the AP. is the path loss factor. is the transmission noise. is the bandwidth. is the fading power. D2, Definition Subject to the Rice distribution, its probability density function is expressed as: , where is the Rice factor, is the constant coefficient, represents the first-kind zero-order modified Bessel function, is the mean calculation; The channel capacity can be obtained according to Shannon's theorem for D3. , and further, the optimal allocated power for the -th SA to satisfy the C4 condition is denoted by . , where is the inverse function at the maximum tolerance value of , that is, to ensure the minimum requirement of , is the transmission data packet length, is the inverse function, is the block error probability.

Citation Information

Cited By

  • Network communication construction wiring dynamic adjustment system based on reinforcement learning

    CN122247880A

  • A network communication construction wiring dynamic adjustment system based on reinforcement learning

    CN122247880B