A complex terrain railway slope multi-source heterogeneous agent energy perception scheduling method

CN122547503APending Publication Date: 2026-08-11EAST CHINA JIAOTONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

若为了节省电量而强制设定极低的固定采样频率,一旦遇到突发暴雨导致边坡微小形变加剧,系统将由于处于休眠或低频轮询状态,无法捕获这一瞬态前兆特征,导致错失最佳预警时间窗口

Benefits of technology

[0018]与现有技术相比,本发明的有益效果在于:利用深度强化学习模型根据边坡动态风险态势自适应调节采样频率。在边坡稳定期使高耗能节点进入深度休眠,杜绝能量无效损耗;在气象突变或形变加剧的极端工况下,毫秒级唤醒关键节点高频采样,从根本上解决了高耗能与灾害漏报之间的绝对矛盾。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547503A_ABST
    Figure CN122547503A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-source heterogeneous intelligent agent energy sensing and scheduling method for railway slopes in complex terrain, resolving the contradiction between high-precision sensing and limited power supply at nodes in slope monitoring. The invention constructs a multi-source heterogeneous intelligent agent monitoring platform comprising remote sensing intelligent agents, physical quantity sensing intelligent agents, and meteorological environment monitoring intelligent agents. An edge computing gateway collects multimodal data and the remaining power of each node, constructing a state space input including deformation, hydrology, meteorology, and energy. A scheduling model with a hybrid action space and a composite reward function for multi-objective collaboration are constructed. A deep reinforcement learning model is trained offline iteratively, and the feature extraction network and the policy network of the scheduling model are jointly deployed to the edge computing gateway, outputting differentiated sampling frequency commands to the heterogeneous intelligent agents. This invention breaks through the energy consumption bottleneck of fixed-frequency scheduling, achieving adaptive wake-up and millisecond-level collaborative sensing under disaster environments, and extending the lifecycle of the slope monitoring network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring of geological disasters and collaborative sensing technology of the Internet of Things, and in particular to a method for energy sensing and scheduling of multi-source heterogeneous intelligent agents on railway slopes with complex terrain. Background Technology

[0002] With the rapid development of railway transportation networks, a large number of railway lines pass through mountainous areas with extremely complex geological conditions. To prevent geological disasters such as slope instability, monitoring networks with multi-source heterogeneous intelligent agents, including visual perception, deep tilt angle, and surface displacement sensors, are often deployed in dangerous areas in the field. However, these edge nodes mainly rely on limited self-contained batteries or micro-energy harvesting for power, creating a significant contradiction between high-precision real-time perception and early warning and the limited power of the nodes.

[0003] Currently, existing technologies primarily employ static, uniform sampling and transmission frequencies to address the aforementioned contradictions. This fixed scheduling mechanism has a fatal flaw in practical applications, specifically manifested in the following ways: On the one hand, there is significant energy wastage during the slope stabilization period. Complex railway slopes are in a state of slow deformation or stability for most of the time. Maintaining a high fixed sampling frequency, especially for high-energy-consuming visual sensing nodes, will generate massive amounts of redundant normal data. Continuously collecting and transmitting this worthless data over long distances will cause the battery energy of heterogeneous edge nodes to be rapidly depleted, and may even cause the network to prematurely collapse due to energy depletion, significantly shortening the lifespan of the monitoring system.

[0004] On the other hand, extreme operating conditions can easily lead to missed disaster reports or delayed early warnings. The railway slope environment is greatly affected by sudden factors such as heavy rainfall and train dynamic loads, and instability often occurs in an instant. If an extremely low fixed sampling frequency is forcibly set in order to save power, once a sudden rainstorm causes the small deformation of the slope to intensify, the system will be in a dormant or low-frequency polling state and will be unable to capture this transient precursor feature, resulting in missing the best early warning window.

[0005] In summary, due to the lack of a scheduling mechanism that can adaptively balance the real-time risk status of the slope with the remaining power of each node, the existing fixed-frequency monitoring technology has always been struggling to compromise between node failure caused by high energy consumption and disaster underreporting caused by low-frequency sampling, and cannot fundamentally break the energy bottleneck to achieve reliable all-weather sensing. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a multi-source heterogeneous intelligent agent energy sensing and scheduling method for railway slopes in complex terrain. It utilizes a deep reinforcement learning model to jointly sense the dynamic risk situation of the slope and the current remaining power of each heterogeneous intelligent agent node. During the slope's stable period, the high-energy-consuming remote sensing monitoring agent enters deep dormancy to prevent ineffective energy loss. When high-risk characteristics of slope instability appear, the remote sensing monitoring agent is awakened to perform high-frequency emergency sampling, ensuring timely early warning. Furthermore, by leveraging a cloud-edge collaborative architecture, the lightweight policy network is deployed to an edge computing gateway to achieve real-time closed-loop adaptive control in harsh field environments.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a multi-source heterogeneous intelligent agent energy sensing and scheduling method for railway slopes in complex terrain, comprising the following steps: Step S1: Build a multi-source heterogeneous intelligent agent monitoring platform covering the air, underground and ground, and collect the current remaining power of each heterogeneous intelligent agent node in the network in real time, as well as external meteorological data of the slope, remote sensing image information of the slope and physical quantity signals of the slope. Step S2: Extract the remote sensing image information of the slope into image feature vectors through a feature extraction network, and perform time stamp alignment and normalization processing on the image feature vectors with the external meteorological data of the slope, the physical quantity signals of the slope and the remaining power of each node to construct the state space input of the deep reinforcement learning model; Step S3: Construct a scheduling model with a hybrid action space, in which discrete action branches are used for remote sensing monitoring agents, and continuous action branches are used for physical quantity sensing agents and meteorological environment monitoring agents. Differentiated physical boundary mappings are designed for different agents. At the same time, a multi-objective collaborative composite reward function is constructed to evaluate the long-term cumulative reward of scheduling actions.

[0008] Step S4: Based on the cloud-edge collaborative architecture, a virtual simulation environment for slopes is built on the cloud server and the deep reinforcement learning model is trained offline iteratively. The feature extraction network and the policy network of the scheduling model are jointly deployed to the edge computing gateway, and the online output scheduling instructions control the sampling frequency of heterogeneous intelligent agents to achieve adaptive closed-loop adjustment.

[0009] Further preferably, the multi-source heterogeneous intelligent agent monitoring platform includes: A remote sensing monitoring intelligent agent is used to periodically acquire remote sensing image information of the slope; A physical quantity sensing agent is used to continuously collect the physical quantity signals of the slope, including groundwater level, pressure and angular velocity signals. A meteorological and environmental monitoring intelligent agent is used to continuously monitor external meteorological data of slopes.

[0010] Further preferably, the state space input for constructing the deep reinforcement learning model specifically includes: Multimodal data and the remaining power of each node are spliced ​​together to construct a multidimensional dynamic feature vector covering deformation, hydrology, meteorology and energy status. When there are sudden changes in groundwater level or abnormal changes in pressure in the slope physical quantity signal, the corresponding features are processed by a nonlinear amplification function to increase their numerical representation ratio in the multidimensional dynamic feature vector, thereby characterizing the current slope potential instability risk level.

[0011] Further preferred, the construction of the scheduling model for the hybrid action space, and the design of differentiated physical boundary mappings for different agents, specifically includes: For the physical quantity sensing agent and the meteorological environment monitoring agent, a basic sampling frequency range is set. The original output of the scheduling model is mapped to the basic sampling frequency range through linear mapping or a Sigmoid variant function to maintain the basic sensing of the system. For the remote sensing monitoring agent, a discrete action space including deep sleep, low-frequency inspection and high-frequency emergency response is set. The discrete action branch output of the scheduling model corresponds to the action probability distribution. In the online real-time inference stage, the scheduling model directly selects the action with the largest value in the probability distribution as the current output command based on the input of the multi-dimensional dynamic feature vector, thereby realizing adaptive wake-up control of the remote sensing monitoring agent.

[0012] Further preferably, the composite reward function comprehensively evaluates the long-term cumulative reward of scheduling actions, in At any given moment, the agent executes the currently scheduled action. And transition to the new state Subsequently, the overall immediate reward for environmental feedback. Defined as: ; in, , , These are the perceived benefit functions driven by catastrophe priors. Differential energy penalty function for heterogeneous nodes Ultimate penalty function for network failure based on battery threshold Normalized weighting coefficients; Let this be the current state vector. For all intelligent agents in The set of remaining battery statuses at any given time.

[0013] More preferably, step S4 includes: A virtual slope simulation environment is built on a cloud server. The agent interacts in the virtual slope simulation environment in multiple rounds, updates the network parameters through policy gradient descent, and uses the value evaluation network to calculate the long-term cumulative reward of the action until the composite reward function reaches maximum convergence. After training, the value evaluation network is stripped away, leaving only the policy network and the feature extraction network for image processing. The policy network and the feature extraction network are then subjected to quantization, compression, and lightweighting processes, and deployed to the edge computing gateway on-site.

[0014] Further preferably, the state space input for constructing the deep reinforcement learning model also includes: when the remote sensing monitoring agent is in a deep dormant state and no new images are transmitted back, the corresponding image feature vector is maintained at zero order using the cached features from the previous wake-up time.

[0015] Further preferred, the scheduling model of the hybrid action space adopts a policy network-value evaluation network architecture, wherein the policy network contains two fully connected hidden layers connected in sequence, and the activation function is ReLU; at the output end, the discrete action branch outputs the probability distribution through the Softmax function, and the continuous action branch outputs the mapped sampling frequency through the Sigmoid function.

[0016] The present invention also provides an electronic device, comprising: At least one processor; and A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are then executed by the at least one processor to enable the at least one processor to perform the above-described multi-source heterogeneous intelligent agent energy sensing and scheduling method for railway slopes in complex terrain.

[0017] The present invention also provides a computer-readable storage medium storing computer instructions for causing a computer to execute the above-described method for energy sensing and scheduling of multi-source heterogeneous intelligent agents for railway slopes in complex terrain.

[0018] Compared with existing technologies, the advantages of this invention are as follows: It utilizes a deep reinforcement learning model to adaptively adjust the sampling frequency based on the dynamic risk situation of the slope. During the stable period of the slope, high-energy-consuming nodes enter deep dormancy, eliminating ineffective energy loss; under extreme conditions such as sudden weather changes or intensified deformation, key nodes are awakened for high-frequency sampling within milliseconds, fundamentally resolving the absolute contradiction between high energy consumption and unreported disasters.

[0019] To address the significant differences in energy consumption among remote sensing, physical quantity, and meteorological agents, a multi-branch network output and differentiated physical boundary mapping mechanism was designed. By implementing master-slave collaboration between high-energy-consuming remote sensing agents and low-energy-consuming physical quantity agents, the long-term survival rate of the network under limited energy conditions in the field is significantly improved. Furthermore, subtle changes in physical quantities such as groundwater level fluctuations and pressure anomalies, as well as meteorological information, are assigned as prior features to the high-weight input state vector, enabling the model to possess forward-looking risk perception capabilities.

[0020] Employing a cloud-edge collaborative distributed architecture, the cloud server utilizes massive amounts of data and a virtual slope simulation environment to complete complex strategy iterations, then distributes the simplified and compressed model to the edge computing gateway. The edge terminal only needs to perform a single forward propagation with extremely low computing power consumption to achieve millisecond-level adaptive closed-loop control, perfectly adapting to the hardware deployment constraints of harsh outdoor slope environments. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the architecture of a multi-source heterogeneous intelligent agent collaborative monitoring system for railway slopes in complex terrain, provided by an embodiment of the present invention. Figure 2 This is a main flowchart of a multi-source heterogeneous intelligent agent energy sensing and scheduling method provided in an embodiment of the present invention; Figure 3 This is a diagram illustrating the network structure and physical boundary mapping principle of the deep reinforcement learning multi-branch scheduling model in this embodiment of the invention. Figure 4 This is a time-series interaction diagram of offline training and online real-time inference of a model based on a cloud-edge collaborative architecture in an embodiment of the present invention. Detailed Implementation

[0022] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solution of the present invention, providing detailed implementation methods and specific operating procedures; however, the scope of protection of the present invention is not limited to the following embodiments.

[0023] Please see Figure 1 and Figure 2 This invention provides a method for energy sensing and scheduling of multi-source heterogeneous intelligent agents on railway slopes in complex terrain, specifically including the following steps: Step S1: Build a multi-source heterogeneous intelligent agent monitoring platform covering the air, underground and ground, and collect the current remaining power of each heterogeneous intelligent agent node in the network in real time (the remaining power is expressed as a percentage in this embodiment, but it is actually a percentage of the remaining power. For the sake of simplification, it is uniformly referred to as the remaining power), as well as external meteorological data of the slope, remote sensing image information of the slope and physical quantity signals of the slope. Step S2: Extract the remote sensing image information of the slope into image feature vectors through a feature extraction network, and perform time stamp alignment and normalization processing on the image feature vectors with the external meteorological data of the slope, the physical quantity signals of the slope and the remaining power of each node to construct the state space input of the deep reinforcement learning model; Step S3: Construct a scheduling model with a hybrid action space, in which discrete action branches are used for remote sensing monitoring agents, and continuous action branches are used for physical quantity sensing agents and meteorological environment monitoring agents. Differentiated physical boundary mappings are designed for different agents. At the same time, a multi-objective collaborative composite reward function is constructed to evaluate the long-term cumulative reward of scheduling actions. Step S4: Based on the cloud-edge collaborative architecture, a virtual simulation environment for slopes is built on the cloud server and the deep reinforcement learning model is trained offline iteratively. The feature extraction network and the policy network of the scheduling model are jointly deployed to the edge computing gateway, and the online output scheduling instructions control the sampling frequency of heterogeneous intelligent agents to achieve adaptive closed-loop adjustment.

[0024] Among them, such as Figure 1 As shown, a multi-level sensor network is deployed on the railway slope, specifically including: remote sensing monitoring agents, physical quantity sensing agents, and meteorological environment monitoring agents. The remote sensing monitoring agents include remote sensing sensors such as visible light cameras, InSAR (Intrusive Synthetic Aperture Radar), and LiDAR (LiDAR) carried by UAVs, used to periodically acquire remote sensing image information of the slope, including visible light, InSAR, and LiDAR images. The physical quantity sensing agents include physical quantity sensors such as water level, pressure, and fiber optic gyroscopes, used to continuously collect physical quantity signals of the slope, including groundwater level, pressure, and angular velocity signals. The meteorological environment monitoring agents include meteorological sensors such as temperature, humidity, and rainfall, used to continuously monitor external meteorological data of the slope, including rainfall, temperature, and humidity. The edge computing gateway reads the current remaining power of each sensing node and the multi-source sensing data in real time through underlying communication protocols (such as Modbus communication protocol and Message Queuing Telemetry Transport Protocol MQTT).

[0025] Specifically, the edge computing gateway uploads all sensing data and status information. For slope remote sensing image information (visible light, InSAR, and LiDAR images) acquired by the remote sensing monitoring agent, a pre-trained lightweight convolutional neural network (e.g., MobileNetV3) is first used to extract fixed-dimensional feature vectors. The parameters of this feature extraction network are frozen during training on the cloud server or fine-tuned jointly with the policy network. Assume the system includes one UAV (remote sensing monitoring agent), three stress gauges and one fiber optic gyroscope (physical quantity sensing agent), and one rain gauge (meteorological environment monitoring agent). For the slope remote sensing image acquired by the UAV at the current moment, the edge computing gateway inputs it into the pre-trained MobileNetV3 network to extract a 512-dimensional image feature vector. Simultaneously obtain the current remaining power of each node: if The drone currently has 15% battery remaining, the three stress gauges have 40%, 45%, and 80% battery remaining respectively, and the rain gauge has 100% battery remaining. (External meteorological data for the slope) Includes current rainfall (0.2 mm / h); slope physical quantity signals. This includes groundwater level (2.3 m, safety threshold 2.5 m), angular velocity (0.01 rad / s, safety threshold 0.05 rad / s), and maximum pressure (1.2 MPa, safety threshold 1.0 MPa). The image feature vector is timestamped and normalized with slope physical quantity signals, external meteorological data, and remaining electrical energy at each node. All data are then concatenated to construct a multi-dimensional dynamic feature vector encompassing deformation, hydrology, meteorology, and energy state, serving as the current state vector. And input it into the deep reinforcement learning model.

[0026] .

[0027] When a sudden change in groundwater level or anomaly in pressure appears in the physical quantity signal of the slope, the corresponding characteristics are... The pressure anomaly feature is assigned a high weight to characterize the level of instability risk. In this embodiment, the pressure anomaly feature is... The proportion of numerical representation in the state vector is increased by using a nonlinear amplification function (e.g., multiplying by an amplification factor of 2) to enhance the model's sensitivity to risk features. If the remote sensing agent is in a deep dormant state and no new slope remote sensing images are transmitted back at the current moment, then the corresponding image feature vector... Zero-order retention is performed using the cache characteristics from the previous wake-up time.

[0028] Please see Figure 3 The deep reinforcement learning model in this embodiment adopts an Actor-Critic architecture and a hybrid action space design. Current state vector. After input, the data passes through two fully connected hidden layers with 256 and 128 neurons respectively, using ReLU activation to extract high-dimensional nonlinear features. Subsequently, the network differentiates at the output, employing discrete and continuous action branches for high-energy-consuming remote sensing agents and low-energy-consuming physical quantity sensing and meteorological environment monitoring agents, resulting in three independent control variables in the output: Among them, the continuous action branch: for the physical quantity sensing agent and the meteorological environment monitoring agent, the original continuous action values ​​of the physical quantity sensing agent are output respectively. and raw continuous action values ​​of meteorological and environmental monitoring intelligent agents The system sets a basic sampling frequency range for these two types of intelligent agents. , The lowest sampling frequency, At the highest sampling frequency, the raw continuous action values ​​of the physical quantity sensing agent are smoothly mapped to the basic sampling frequency range through a linear mapping or a variant of the Sigmoid function. The mapping method is as follows: ; in, σ( represents the actual sampling frequency of the physical quantity sensing agent after mapping.) () represents the Sigmoid function. Similarly, the actual sampling frequency of the meteorological and environmental monitoring agent after mapping can be obtained. Even when the slope is absolutely safe, it is guaranteed to operate at the lowest possible sampling frequency. It runs and maintains the system's basic awareness.

[0029] Among them, the discrete action branch: for high-energy-consuming remote sensing monitoring agents (UAVs), the network outputs three logits (unnormalized log probabilities), which are then processed by the Softmax function to obtain the probability distribution: deep sleep action probability. Low-frequency inspection action probability High-frequency emergency response probability This outputs discrete actions. Deep sleep: sampling frequency is 0 (no image acquisition); low-frequency inspection: sampling frequency is a preset low-frequency value; high-frequency emergency: sampling frequency is a preset high-frequency value.

[0030] During training, the Gumbel-Softmax technique (temperature parameter τ is between 0.5 and 1.0) is used to sample discrete actions so that gradients can be backpropagated; during inference, the action with the highest probability is directly taken as the output command.

[0031] Ultimately, the model outputs the scheduling action. ,in For discrete action indexing, and These are continuous frequency values.

[0032] The specific process of constructing a multi-objective collaborative composite reward function in step S3 to evaluate the long-term cumulative reward of scheduling actions is as follows: exist At any given moment, the intelligent agent executes scheduling actions. And transition to the new state Subsequently, the overall immediate reward for environmental feedback. Defined as: ; Among them, among them, , , These are the perceived benefit functions driven by catastrophe priors. Differential energy penalty function for heterogeneous nodes Ultimate penalty function for network failure based on battery threshold Normalized weighting coefficients; Let this be the current state vector. For all intelligent agents in The set of remaining battery statuses at any given time.

[0033] The specific sub-functions are designed as follows: Catastrophe-prior-driven perceived benefit function: ; in, ; in, Basic risk penalty factor; This represents the rate constant of risk surge. Represents the current state vector The Middle Current observed values ​​of a physical quantity or meteorological characteristic. The set of safety thresholds corresponding to this feature This is a set of risk features. When the current state vector... When the physical quantity sensing agent detects a "sudden change in groundwater level" or the meteorological environment agent detects "heavy rainfall", It will exhibit exponential nonlinear amplification. The sampling frequency for remote sensing monitoring agents; for The average sampling frequency of all physical quantities perceived by the intelligent agent at any given time; These are the adjustment constants for the remote sensing monitoring agent and the physical quantity sensing agent, respectively. In this embodiment, the normalized weighting coefficient is set as... , , Basic risk penalty factor Risk surge rate constant Adjustment constant .

[0034] Differentiated energy penalty function for heterogeneous nodes: ; in, , , These represent the unit sampling energy consumption cost of remote sensing monitoring agents, physical quantity sensing agents, and meteorological environment monitoring agents, respectively. The sampling frequency for meteorological and environmental monitoring intelligent agents.

[0035] Among them, since the energy consumption of remote sensing monitoring agents is much greater than that of physical quantity and meteorological agents, the following settings are made: and Basic monitoring should be maintained primarily by relying on low-energy physical agents, with high-energy remote sensing agents only activated when absolutely necessary.

[0036] Ultimate network failure penalty function based on battery threshold: ; in This indicates that after the scheduling action is executed, the first... The agent predicts the remaining battery power. For indicator functions, As a safety threshold, A constant penalty value is set to a very large value, and N is the number of agents. A safety threshold is set. The remaining power of a node is 5%. Indicator function activation, ultimate penalty Setting it to -1000 will result in a drastic penalty from the system.

[0037] In step S4, the specific process of constructing a virtual slope simulation environment on a cloud server and conducting offline iterative training of a deep reinforcement learning model based on a cloud-edge collaborative architecture is as follows: Please see Figure 1 and Figure 4 A virtual slope simulation environment was built on a cloud server equipped with GPU computing power, based on the deep learning framework PyTorch. This virtual slope simulation environment can simulate sudden changes in rainfall, dramatic changes in groundwater level, and the energy consumption decay process of multi-source heterogeneous intelligent agents. The multi-source heterogeneous intelligent agents engage in multi-round interactions within the virtual slope simulation environment, recording experience tuples in real time, including the current state, actions, rewards, and the next state. The data is then stored in the experience replay memory module. The algorithm randomly extracts experience data from the memory pool and updates the policy network (Actor) parameters using policy gradient descent. For discrete action branches, the Gumbel-Softmax reparameterization technique (temperature parameter τ is set to 0.5~1.0) is used to make the sampling process differentiable, thus ensuring normal gradient backpropagation. Simultaneously, the value evaluation network (Critic) is used to evaluate the long-term cumulative reward of actions, calculate the advantage function, and update the value evaluation network (Critic). Through repeated iterations, a composite reward function is obtained. To achieve maximum convergence.

[0038] In step S4, the lightweight feature extraction network and the scheduling model are jointly deployed to the edge computing gateway, and online scheduling instructions are output to control the sampling frequency of heterogeneous intelligent agents to achieve adaptive closed-loop adjustment. The specific process is as follows: Reference Figure 1 and Figure 4 After training on the cloud server, the value evaluation network (Critic) is stripped away, leaving only the converged policy network (Actor) and the lightweight MobileNetV3 feature extraction network for image processing. Its parameters are then quantized and compressed from FP32 to INT8 format, perfectly adapting to the computing power limitations of low-power edge computing gateways in the field. During the online real-time inference phase, the lightweight Actor network is deployed to the edge computing gateway at the railway slope site. The edge computing gateway collects external meteorological data, remote sensing images, and physical quantity signals of the slope, as well as the remaining power of each node, through multi-source heterogeneous intelligent agents. The remote sensing image information is extracted into image feature vectors by the feature extraction network. Using a built-in multi-dimensional state feature stitching module, the image feature vectors are instantly stitched together with the external meteorological data, physical quantity signals, and remaining power of each node to form the current real-world state vector. .Will The input is fed into a lightweight Actor network, which instantly outputs scheduling instructions for multi-source heterogeneous intelligent agents. Subsequently, the gateway's control command protocol conversion module immediately converts the scheduling commands into underlying communication control frames, and outputs differentiated sampling frequency commands to the multi-source heterogeneous intelligent agents through the local area network segment, thereby realizing adaptive closed-loop control in complex dynamic scenarios.

[0039] To verify the effectiveness of the energy-aware scheduling method based on the Actor-Critic architecture proposed in this invention, this embodiment also provides a set of comparative test case data of a slope virtual simulation environment built based on the deep learning framework (PyTorch).

[0040] During the offline iterative training phase, the analysis of the training loss curve shows that after several rounds of interaction with the virtual slope simulation environment, the loss values ​​of the Actor network and Critic network decrease steadily, and the composite reward function eventually reaches maximum convergence, proving the model's optimization ability in a multidimensional state space.

[0041] During the verification and testing phase, a control group was set up as a traditional fixed-frequency (e.g., continuous sampling at 1Hz) polling scheduling mechanism, and the indicators were compared and verified from two dimensions: "energy consumption and life cycle" and "disaster response accuracy". (1) Energy consumption and life cycle: In a slope stabilization period simulation test lasting up to 30 days, the control group's UAV nodes ran out of power within 72 hours due to high-frequency invalid sampling; while the method of this invention keeps the UAV in a deep hibernation state for a long time, and the physical quantity sensing agent only maintains the basic frequency operation. Due to the significant reduction in the invalid loss of high-energy-consuming equipment, the test shows that the overall effective life cycle of the monitoring network is extended by about 4.5 times compared with the traditional method.

[0042] (2) Disaster response and perception accuracy: A high-risk test set containing sudden heavy rainfall and drastic changes in groundwater level is injected into the simulation environment. After sensing the prior characteristics of groundwater level and meteorological anomalies, the scheduling model of this invention adaptively outputs the high-frequency emergency action command with the highest probability in a very short time (less than 500 milliseconds), effectively avoiding the underreporting of disaster precursors in critical time windows, and achieving a high-risk detection rate of 96.55% in the test set.

[0043] Experiments show that the energy-sensing scheduling method of the present invention can perfectly balance low-power operation of edge nodes and high-precision early warning under extreme conditions in practical engineering applications.

[0044] In this embodiment, an electronic device is also provided, which can serve as part of the aforementioned edge computing gateway or cloud server. The electronic device includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed, enable the at least one processor to perform the operations described in the method embodiments. Specifically, by executing the instructions, the processor performs the following functions: acquiring remote sensing image information of the slope, meteorological data, and slope physical quantity signals, as well as the remaining power of each node, and performing timestamp alignment and normalization processing; running a deep reinforcement learning model for inference, and outputting sampling frequency instructions for different heterogeneous intelligent agents; and converting the instructions into control frames and sending them to each intelligent agent. This electronic device can be an industrial computer, an embedded ARM motherboard, a GPU workstation, or other hardware platform with computing capabilities.

[0045] In this embodiment, a computer-readable storage medium is also provided, which stores computer instructions. The computer instructions are used to cause a computer to perform the operations described in the method embodiments above. The computer-readable storage medium may be a read-only memory (ROM), random access memory (RAM), flash memory, hard disk, solid-state drive (SSD), optical disk (CD-ROM, DVD), or any other non-transitory, tangible storage medium. When the computer instructions are loaded and executed by a computer or processor, the following processes can be achieved: receiving data collected by multi-source heterogeneous intelligent agents and constructing the state space input of a deep reinforcement learning model; invoking a trained policy network for forward inference to generate a sampling frequency scheduling instruction; and sending the scheduling instruction to the corresponding intelligent agent node. This storage medium can be sold independently or pre-installed as firmware in an edge computing gateway device.

[0046] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A complex terrain railway slope multi-source heterogeneous agent energy perception scheduling method, characterized in that, Includes the following steps: Step S1: Build a multi-source heterogeneous intelligent agent monitoring platform covering the air, underground and ground, and collect the current remaining power of each heterogeneous intelligent agent node in the network in real time, as well as external meteorological data of the slope, remote sensing image information of the slope and physical quantity signals of the slope. Step S2: Extract the remote sensing image information of the slope into image feature vectors through a feature extraction network, and perform time stamp alignment and normalization processing on the image feature vectors with the external meteorological data of the slope, the physical quantity signals of the slope and the remaining power of each node to construct the state space input of the deep reinforcement learning model; Step S3: Construct a scheduling model with a hybrid action space, in which discrete action branches are used for remote sensing monitoring agents, and continuous action branches are used for physical quantity sensing agents and meteorological environment monitoring agents. Differentiated physical boundary mappings are designed for different agents. At the same time, a multi-objective collaborative composite reward function is constructed to evaluate the long-term cumulative reward of scheduling actions. Step S4: Based on the cloud-edge collaborative architecture, a virtual simulation environment for slopes is built on the cloud server and the deep reinforcement learning model is trained offline iteratively. The feature extraction network and the policy network of the scheduling model are jointly deployed to the edge computing gateway, and the online output scheduling instructions control the sampling frequency of heterogeneous intelligent agents to achieve adaptive closed-loop adjustment.

2. The complex terrain railway slope multi-source heterogeneous agent energy-aware scheduling method according to claim 1, characterized in that, The multi-source heterogeneous intelligent agent monitoring platform includes: A remote sensing monitoring intelligent agent is used to periodically acquire remote sensing image information of the slope; A physical quantity sensing agent is used to continuously collect the physical quantity signals of the slope, including groundwater level, pressure and angular velocity signals. A meteorological and environmental monitoring intelligent agent is used to continuously monitor external meteorological data of slopes.

3. The complex terrain railway slope multi-source heterogeneous agent energy-aware scheduling method according to claim 1, characterized in that, The state space input for constructing the deep reinforcement learning model specifically includes: Multimodal data and the remaining power of each node are spliced ​​together to construct a multidimensional dynamic feature vector covering deformation, hydrology, meteorology and energy status. When there are sudden changes in groundwater level or abnormal changes in pressure in the slope physical quantity signal, the corresponding features are processed by a nonlinear amplification function to increase their numerical representation ratio in the multidimensional dynamic feature vector, thereby characterizing the current slope potential instability risk level.

4. The multi-source heterogeneous intelligent agent energy sensing and scheduling method for railway slopes in complex terrain according to claim 1, characterized in that, The construction of a scheduling model for a hybrid action space, and the design of differentiated physical boundary mappings for different agents, specifically includes: For the physical quantity sensing agent and the meteorological environment monitoring agent, a basic sampling frequency range is set. The original output of the scheduling model is mapped to the basic sampling frequency range through linear mapping or a Sigmoid variant function to maintain the basic sensing of the system. For the remote sensing monitoring agent, a discrete action space including deep sleep, low-frequency inspection and high-frequency emergency response is set. The discrete action branch output of the scheduling model corresponds to the action probability distribution. In the online real-time inference stage, the scheduling model directly selects the action with the largest value in the probability distribution as the current output command based on the input of the multi-dimensional dynamic feature vector, thereby realizing adaptive wake-up control of the remote sensing monitoring agent.

5. The complex terrain railway slope multi-source heterogeneous agent energy-aware scheduling method according to claim 1, characterized in that, The composite reward function comprehensively evaluates the long-term cumulative reward of scheduling actions. At any given moment, the agent executes the currently scheduled action. And transition to the new state Subsequently, the overall immediate reward for environmental feedback. Defined as: ; in, , , These are the perceived benefit functions driven by catastrophe priors. Differential energy penalty function for heterogeneous nodes Ultimate penalty function for network failure based on battery threshold Normalized weighting coefficients; Let this be the current state vector. For all intelligent agents in The set of remaining battery statuses at any given time.

6. The complex terrain railway slope multi-source heterogeneous agent energy-aware scheduling method according to claim 1, characterized in that, Step S4 includes: A virtual slope simulation environment is built on a cloud server. The agent interacts in the virtual slope simulation environment in multiple rounds, updates the network parameters through policy gradient descent, and uses the value evaluation network to calculate the long-term cumulative reward of the action until the composite reward function reaches maximum convergence. After training, the value evaluation network is stripped away, leaving only the policy network and the feature extraction network for image processing. The policy network and the feature extraction network are then subjected to quantization, compression, and lightweighting processes, and deployed to the edge computing gateway on-site.

7. The complex terrain railway slope multi-source heterogeneous agent energy-aware scheduling method according to claim 1, characterized in that, The state space input for constructing the deep reinforcement learning model also includes: when the remote sensing agent is in a deep dormant state and no new images are transmitted, the corresponding image feature vector is maintained at zero order using the cached features from the previous wake-up time.

8. The multi-source heterogeneous intelligent agent energy sensing and scheduling method for railway slopes in complex terrain according to claim 1, characterized in that, The scheduling model of the hybrid action space adopts a policy network-value evaluation network architecture, wherein the policy network contains two fully connected hidden layers connected in sequence, and the activation function is ReLU; at the output end, the discrete action branch outputs the probability distribution through the Softmax function, and the continuous action branch outputs the mapped sampling frequency through the Sigmoid function.

9. An electronic device, comprising: include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to execute the multi-source heterogeneous intelligent agent energy sensing and scheduling method for complex terrain railway slopes as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the energy sensing and scheduling method for multi-source heterogeneous intelligent agents on complex terrain railway slopes as described in any one of claims 1 to 8.