Disaster environment distributed collaborative perception method based on multi-agent reinforcement learning
By employing multi-agent reinforcement learning and utilizing data preprocessing and dynamic compensation techniques, the data quality and collaborative control issues of distributed sensing systems in disaster environments were addressed. This enabled adaptive collaborative deployment of agents, improving sensing efficiency and robustness in disaster environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN ZHONGDI YUNSHEN TECH CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-29
Smart Images

Figure CN121765288B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computational processing technology. More specifically, this invention relates to a distributed collaborative perception method for disaster environments based on multi-agent reinforcement learning. Background Technology
[0002] In extreme environments such as earthquake secondary disasters and mine collapses, due to fragmented terrain, smoke and dust obstruction, and signal transmission blockage, single detection equipment often struggles to achieve efficient coverage of large and complex areas. Utilizing distributed systems composed of multiple agents to collaboratively perform sensing tasks has become a common technical approach in this field. However, the mechanical vibrations, electromagnetic pulses, and non-stationary fluctuations present in disaster environments can introduce a large amount of interference noise into the raw data sequences acquired by various sensors. If this low-quality observation data is directly used in subsequent calculations without targeted purification and standardization, it can easily cause the sensing system to make biased judgments about the environmental state, severely weakening the ability of distributed nodes to extract the essential characteristics of the disaster environment.
[0003] Without global scheduling, distributed sensing nodes can typically make decisions based solely on local observation information. This leads to multiple nodes blindly clustering in high-value areas to maximize individual coverage during search tasks, resulting in severe physical overlap and information redundancy in the detection fields of different nodes. This not only fails to effectively improve the overall spatial breadth of sensing but also causes an imbalance in the allocation of limited detection resources. Furthermore, the dynamic evolution of disaster environments is highly unpredictable, and drastic fluctuations in environmental field characteristics often trigger frequent fluctuations in multi-agent strategies. Currently, when mapping abstract decision-making logic to underlying physical control commands, there is a lack of coupling consideration between task completion and environmental evolution rate. This causes drive motors to continue ineffective reciprocating motion or generate mechanical flutter in areas where sensing tasks are approaching saturation, limiting the collaborative efficiency of cluster sensing and the robustness of physical operations. Summary of the Invention
[0004] To address the technical problems of poor sensing data quality, severe information redundancy between nodes, and inaccurate physical cooperative control in distributed sensing systems under disaster environments, this invention provides a distributed cooperative sensing method for disaster environments based on multi-agent reinforcement learning. The method includes: acquiring and preprocessing the initial sequence of the disaster environment perceived by the agents to obtain standardized state observations; extracting the spatial entropy of the agent at the current moment and calculating the perception difference between the agent and its neighboring agents, obtaining cooperative sensing interaction gain based on the spatial entropy and perception difference; acquiring the environmental field change rate, dynamically compensating the original action evaluation value output by the basic reinforcement learning network based on the environmental field change rate to obtain a corrected action evaluation value; acquiring the agent's perception task saturation at the current moment, mapping the physical control vector based on the corrected action evaluation value and the perception task saturation, and adjusting the agent's perception task intensity through the physical control vector.
[0005] This invention acquires and preprocesses the initial sequence of the disaster environment, uses a collaborative perception interaction gain model and the rate of change of the environmental field to dynamically compensate the action evaluation value, and adjusts the perception operation intensity of the agent based on the perception task saturation. This enables the distributed perception system to balance the relationship between individual coverage and cluster cooperation in complex disaster environments, thereby improving the overall task execution quality of distributed collaborative perception.
[0006] Preferably, the preprocessing includes: applying standard discrete wavelet transform to perform frequency domain denoising on the initial sequence of the disaster environment, extracting signal components, and performing linear normalization.
[0007] This invention applies standard discrete wavelet transform to perform frequency domain denoising and linear normalization on the initial sequence of disaster environment, reducing the interference of impulse noise and non-stationary fluctuations in sensor data from earthquake secondary disasters or mine collapse sites on subsequent calculation logic, and providing a high-quality and scale-uniform data foundation for logic calculation in disaster environments.
[0008] Preferably, the collaborative sensing interaction gain satisfies the expression: ;in, For the first An intelligent agent in Moment-to-moment collaborative perception interaction gain, For indexing intelligent agents, For a moment, For the first An intelligent agent in Spatial entropy at time, For the first An intelligent agent in The degree of perceptual difference between itself and its neighboring agents at any given time. This is the preset gain sensitivity coefficient. This is for natural exponent calculations.
[0009] This invention extracts the spatial entropy of an agent and calculates the perceptual difference between it and neighboring agents. It evaluates the actual contribution of individual agents to the cluster through a collaborative perception interaction gain model, thereby reducing the overlap of observation fields and information redundancy when different agents perform tasks and improving the breadth of disaster environment perception.
[0010] Preferably, the corrected action evaluation value satisfies the expression: ;in, For the first An intelligent agent in The motion evaluation value after real-time correction For the first An intelligent agent in The raw action evaluation values output by the basic reinforcement learning network at each moment. For the first An intelligent agent in Moment-to-moment collaborative perception interaction gain, As a robust regulator, For the first An intelligent agent in The rate of change of the environmental field sensed at all times. For natural logarithm operations, This is a modulo operation.
[0011] This invention utilizes the rate of change of the environmental field combined with a robust adjustment factor and natural logarithm calculation to dynamically compensate the original action evaluation value, enabling the distributed sensing system to adaptively correct its sensing strategy and proactively reduce the evaluation of aggressive actions when the environment undergoes unpredictable changes, thereby improving the robustness of the distributed sensing system when operating in hazardous areas.
[0012] Preferably, the physical control vector satisfies the expression: ;in, For the first An intelligent agent in The physical control vector executed at each moment. As the reference control vector constant, This is for natural index calculations. The kurtosis coefficient of the mapping function. For the first An intelligent agent in The motion evaluation value after real-time correction For the first An intelligent agent in Perceive task saturation at any given moment.
[0013] Preferably, the adjustment of the sensing operation intensity of the intelligent agent includes: using the vehicle-mounted embedded control system to adjust the movement speed of the intelligent agent by changing the pulse width modulation duty cycle of the drive motor through physical control vector, and adjusting the sensing sampling frequency of the intelligent agent by changing the point cloud emission trigger period of the laser scanner through physical control vector.
[0014] This invention utilizes an in-vehicle embedded control system to change the pulse width modulation duty cycle of the drive motor and the point cloud emission trigger cycle of the laser scanner through physical control vectors. This maps the abstract action evaluation logic into specific instructions that adjust the movement rate and sensing sampling frequency of the intelligent agent, thereby realizing the dynamic adaptive collaborative deployment of the intelligent agent in physical space.
[0015] Preferably, the rate of change of the environmental field is obtained by calculating the first-order difference modulus of the normalized sensing feature vector at adjacent sampling times.
[0016] Preferably, the perception task saturation is obtained by calculating the ratio of the cumulative entropy of the covered features before the current moment to the expected total entropy.
[0017] Preferably, adjusting the motion rate of the intelligent agent includes: adjusting the output intensity of the physical control vector based on the difference between the corrected action evaluation value and the perception task saturation.
[0018] Preferably, the gain sensitivity coefficient is set to .
[0019] The beneficial effects of this invention are as follows:
[0020] This invention obtains collaborative sensing interaction gain based on spatial entropy and perception difference, and guides the agent to avoid areas with highly overlapping information when performing tasks through numerical feedback, thereby reducing the blind aggregation phenomenon in the distributed sensing process of disaster environment and improving the collaborative operation efficiency of distributed sensing system in complex disaster environment.
[0021] This invention obtains the rate of change of the environmental field by calculating the first-order difference modulus of the normalized sensing feature vector at adjacent sampling times. This is used to correct the action evaluation value and reduce the impact of false evaluation value fluctuations caused by environmental oscillations on decision-making, thereby improving the robustness of the distributed sensing system in the face of rapidly evolving disaster fields.
[0022] This invention combines the perception of task saturation with the acquisition of physical control vectors through a mapping function and changes the working state of the drive motor and laser scanner. This enables the intelligent agent to automatically and smoothly decelerate when the task approaches saturation, reducing the invalid reciprocating motion of the intelligent agent in the covered area and improving the precision of search perception. Attached Figure Description
[0023] Figure 1This is a flowchart illustrating the distributed collaborative perception method for disaster environments based on multi-agent reinforcement learning in this invention.
[0024] Figure 2 This is a schematic diagram illustrating how the collaborative perception interaction gain changes with the degree of perception difference;
[0025] Figure 3 This is a schematic diagram illustrating the dynamic compensation effect of motion evaluation values;
[0026] Figure 4 This is a schematic diagram illustrating the trend of sensing sampling frequency adjustment with sensing task saturation. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0029] This invention discloses a distributed collaborative perception method for disaster environments based on multi-agent reinforcement learning, referring to... Figure 1 This includes steps S1 to S4:
[0030] S1. Obtain the initial sequence of the disaster environment perceived by the agent and preprocess it to obtain standardized state observation values.
[0031] It should be noted that sensor data from disaster environments such as earthquake secondary disasters or mine collapses often contain a large amount of impulse noise and non-stationary fluctuations. If these data containing interference are directly input into the calculation process, it will cause serious deviations in the system's judgment of the environment. Therefore, the distributed sensing system of this invention purifies the input data through standardized transformation, providing a high-quality and scale-uniform data foundation for subsequent logical calculations.
[0032] Specifically, the distributed sensing system acquires the initial sequence of the disaster environment through sensors mounted on the intelligent agent. The distributed sensing system then uses standard discrete wavelet transform to perform frequency domain denoising on the initial sequence of the disaster environment, extracting signal components that reflect the essential characteristics of the disaster environment. The distributed sensing system performs linear normalization on the denoised signal, mapping the data to a unified feature space between 0 and 1, and obtaining standardized state observations.
[0033] S2. Extract the spatial entropy of the agent at the current moment and calculate the perceptual difference between the agent and its neighboring agents. Based on the spatial entropy and the perceptual difference, obtain the collaborative perceptual interaction gain.
[0034] It should be noted that during the process of agents performing distributed perception tasks, if only the local perception information of individuals is relied upon, agents will often gather in easily observable areas in order to pursue their own coverage, resulting in serious overlap of the observation fields of different agents. By acquiring the cooperative perception interaction gain, the distributed perception system can evaluate the actual contribution of each agent to the entire distributed perception system, thereby suppressing information redundancy in the distributed perception system at the logical level and improving the overall breadth of perception.
[0035] Specifically, the distributed perception system extracts the spatial entropy of the agent at the current moment and calculates the perceptual difference between the agent and its neighboring agents. Based on the spatial entropy and the perceptual difference, the distributed perception system obtains the cooperative perception interaction gain through a cooperative perception interaction gain model.
[0036] The collaborative sensing interaction gain satisfies the expression:
[0037]
[0038] in, For the first An intelligent agent in Moment-to-moment collaborative perception interaction gain, For indexing intelligent agents, For a moment, For the first An intelligent agent in Spatial entropy at time, For the first An intelligent agent in The degree of perceptual difference between itself and its neighboring agents at any given time. This is the preset gain sensitivity coefficient. This is for natural exponent calculations.
[0039] In the formula, the perceived difference degree A smaller value means that the observations of the agents are more similar, which leads to a larger exponential term in the denominator, thus reducing the cooperative perception interaction gain. Rapidly decreasing. Spatial entropy. The larger the value, the more complex the distribution of information points observed by the agent, and the higher the corresponding cooperative sensing interaction gain. The higher the base value, the better. This collaborative perception interaction gain, through numerical feedback, can guide the agent to avoid areas where information overlaps significantly.
[0040] It should be added that the gain sensitivity coefficient in this invention The value range is set between 0.5 and 2. If the gain sensitivity coefficient... If the value is set too low, the system's ability to filter out redundant information will be weakened; if the value is set too high, the system will overreact to normal perceptual fluctuations. This invention uses experimental data to adjust the gain sensitivity coefficient... Setting it to 1.2 ensures that the distributed sensing system can maintain stable collaborative performance even in complex environments.
[0041] For example, Figure 2 This is a schematic diagram illustrating the change in cooperative sensing interaction gain as sensing difference increases, showing the evolution trend of cooperative sensing interaction gain with increasing sensing difference under the influence of random environmental noise. The curve exhibits significant oscillations at low differences, reflecting the sensitivity of cooperative sensing interaction gain when the agent is in a highly overlapping region. As sensing difference increases, the gain value steadily increases amidst fluctuations, demonstrating the incentive effect of this invention on high-contribution sensing behaviors.
[0042] S3. Obtain the rate of change of the environmental field, and dynamically compensate the original action evaluation value output by the basic reinforcement learning network based on the rate of change of the environmental field to obtain the corrected action evaluation value.
[0043] It should be noted that the evolution of disaster environments is highly unpredictable, and sudden changes in the environment can cause erroneous fluctuations in the strategies of distributed sensing systems. By introducing the rate of change of the environmental field as a penalty factor, the distributed sensing system can adaptively correct the original action evaluation values of the agent. This allows the distributed sensing system to proactively reduce the evaluation of aggressive actions when the environment becomes unstable, thus making the entire distributed sensing system more robust in dangerous areas.
[0044] Specifically, the distributed perception system obtains the rate of change of the environmental field, reflecting the speed of environmental evolution, by calculating the feature changes between adjacent sampling times. The distributed perception system then uses this rate of change of the environmental field to correct the original action evaluation value output by the basic reinforcement learning network, obtaining the corrected action evaluation value.
[0045] The corrected action evaluation value satisfies the expression:
[0046]
[0047] in, For the first An intelligent agent in The motion evaluation value after real-time correction For the first An intelligent agent in The raw action evaluation values output by the basic reinforcement learning network at each moment. For the first An intelligent agent in Moment-to-moment collaborative perception interaction gain, As a robust regulator, For the first An intelligent agent in The rate of change of the environmental field sensed at all times. For natural logarithm operations, This is a modulo operation.
[0048] In the formula, the rate of change of the environmental field The larger the value, the more intense the environmental fluctuations, and the higher the value of the denominator will be. When the denominator increases, the corrected action evaluation value... This will decrease accordingly. This adjustment logic forces the distributed sensing system to adopt a more conservative sensing strategy in unstable environments. In the formula, by introducing the cooperative sensing interaction gain as a product term, nonlinear suppression of redundant sensing actions is achieved. When the agent is in the sensing overlap region, The reduction of the direct drive correction of the motion evaluation value The decline forces the perceptual agent to avoid redundant areas through policy optimization.
[0049] It should be added that the robustness regulating factor in this invention The value range is set between 0.5 and 1.2. When the robustness adjustment factor... When the value is set too low, the system's ability to suppress drastic environmental changes will be insufficient; when the value is too high, the system will become too sluggish in responding to normal environmental conditions. In this invention, the value is set to 0.8, which can effectively filter out false evaluation value fluctuations caused by environmental interference.
[0050] For example, Figure 3 This is a schematic diagram illustrating the dynamic compensation effect of action evaluation values. The diagram compares the changes in the original and corrected action evaluation values under conditions of drastic environmental changes. In the middle region of the time frame, when the rate of change of the environmental field increases significantly, the corrected action evaluation values show a suppressed downward trend, effectively filtering out the radical bias in the original evaluation. This demonstrates that the system improves decision-making robustness in unstable disaster environments through a dynamic compensation mechanism.
[0051] S4. Obtain the perception task saturation of the agent at the current moment, complete the physical control vector mapping based on the corrected action evaluation value and perception task saturation, and adjust the perception task intensity of the agent through the physical control vector.
[0052] It should be noted that, in order for the distributed sensing system to perform precise search operations in real disaster environments, abstract numerical evaluation values must be converted into control commands that the agent can directly execute. By calculating the saturation of the sensing task, the distributed sensing system can determine whether the current area has been sufficiently observed. Through a nonlinear mapping structure, the distributed sensing system can ensure that the agent maintains efficient operation when the task is not completed, and automatically and smoothly decelerates when the task approaches saturation, avoiding ineffective repetitive movements.
[0053] Specifically, the distributed perception system calculates the perception task saturation of the agent at the current moment. The distributed perception system constructs a mapping function that includes a correction for the perception task saturation, converting the corrected action evaluation value into a physical control vector for the agent to execute.
[0054] The physical control vector satisfies the expression:
[0055]
[0056] in, For the first An intelligent agent in The physical control vector executed at each moment. As the reference control vector constant, This is for natural index calculations. The kurtosis coefficient of the mapping function. For the first An intelligent agent in The motion evaluation value after real-time correction For the first An intelligent agent in Perceive task saturation at any given moment.
[0057] In the formula, when the corrected action evaluation value Significantly higher than the perception task saturation When this occurs, it indicates that the location has extremely high perceptual value and is not yet covered; the mapping term will tend towards 1, thus allowing the physical control vector to... Reach maximum output. When the perceived task saturation... As the intensity of perception increases gradually, the agent will automatically reduce the intensity of its perception.
[0058] It should be added that the kurtosis coefficient of the mapping function in this invention The value range is between 2 and 8. If the kurtosis coefficient of the mapping function... Choosing a value that is too low will result in overly sluggish action feedback from the agent; choosing a value that is too high will lead to frequent mechanical vibrations. This invention sets this value to 5 based on mechanical smoothness requirements, which ensures the agent's adaptive and cooperative deployment in physical space.
[0059] Furthermore, the present invention obtains the physical control vector through calculation. When each agent independently identifies the characteristics of the local disaster field, the vehicle-mounted embedded control system changes the pulse width modulation (PWM) duty cycle of the drive motor and the point cloud emission trigger cycle of the laser scanner in real time, thereby realizing the dynamic adaptive collaborative deployment of each sensing node involved in the distributed collaborative perception method of disaster environment based on multi-agent reinforcement learning in physical space.
[0060] For example, Figure 4 This diagram illustrates the trend of the sensing sampling frequency adjustment with the saturation of the sensing task. It shows the adaptive adjustment process of the agent to the intensity of the sensing task under the physical control vector mapping. Due to the influence of hardware system jitter, the sensing sampling frequency curve contains slight sampling noise, but its overall trend shows a significant non-linear decline as the saturation of the sensing task increases. This indicates that the system can automatically reduce the sampling frequency when the task approaches saturation, reducing the resource consumption caused by ineffective reciprocating motion.
Claims
1. A distributed collaborative perception method for disaster environments based on multi-agent reinforcement learning, characterized in that, include: Acquire the initial sequence of the disaster environment perceived by the intelligent agent and preprocess it to obtain standardized state observations; Extract the spatial entropy of the agent at the current moment and calculate the perceptual difference between the agent and its neighboring agents. Based on the spatial entropy and perceptual difference, obtain the cooperative perceptual interaction gain; the cooperative perceptual interaction gain satisfies the expression: ; For the first An intelligent agent in Moment-to-moment collaborative perception interaction gain, For indexing intelligent agents, For a moment, For the first An intelligent agent in Spatial entropy at time, For the first An intelligent agent in The degree of perceptual difference between itself and its neighboring agents at any given time. This is the preset gain sensitivity coefficient. For natural index calculations; The rate of change of the environmental field is obtained, and the original action evaluation value output by the basic reinforcement learning network is dynamically compensated based on the rate of change of the environmental field and the cooperative perception interaction gain to obtain the corrected action evaluation value. The rate of change of the environmental field is obtained by calculating the first-order difference modulus of the normalized sensing feature vectors at adjacent sampling times; Obtain the agent's perception task saturation at the current moment, complete the physical control vector mapping based on the corrected action evaluation value and perception task saturation, and adjust the agent's perception task intensity through the physical control vector.
2. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 1, characterized in that, The preprocessing includes: The standard discrete wavelet transform is applied to perform frequency domain denoising on the initial sequence of the disaster environment, extracting signal components and performing linear normalization.
3. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 1, characterized in that, The corrected action evaluation value satisfies the expression: ; in, For the first An intelligent agent in The motion evaluation value after real-time correction For the first An intelligent agent in The raw action evaluation values output by the basic reinforcement learning network at each moment. For the first An intelligent agent in Moment-to-moment collaborative perception interaction gain, As a robust regulator, For the first An intelligent agent in The rate of change of the environmental field sensed at all times. For natural logarithm operations, This is a modulo operation.
4. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 1, characterized in that, The physical control vector satisfies the expression: ; in, For the first An intelligent agent in The physical control vector executed at each moment. As the reference control vector constant, This is for natural index calculations. The kurtosis coefficient of the mapping function. For the first An intelligent agent in The motion evaluation value after real-time correction For the first An intelligent agent in Perceive task saturation at any given moment.
5. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 1, characterized in that, The adjustment of the perception task intensity of the intelligent agent includes: The vehicle-mounted embedded control system uses physical control vectors to change the pulse width modulation duty cycle of the drive motor to adjust the movement speed of the intelligent agent, and uses physical control vectors to change the point cloud emission trigger period of the laser scanner to adjust the sensing sampling frequency of the intelligent agent.
6. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 4, characterized in that, The perception task saturation is obtained by calculating the ratio of the cumulative entropy of the covered features before the current moment to the expected total entropy.
7. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 5, characterized in that, The regulation of the agent's movement rate includes: The output intensity of the physical control vector is adjusted based on the difference between the corrected action evaluation value and the perceived task saturation.
8. The distributed collaborative perception method for disaster environment based on multi-agent reinforcement learning according to claim 1, characterized in that, The gain sensitivity coefficient is set to .