Hybrid model control method and system applied to negative pressure adsorption type wall-climbing robot

Through the hybrid model control method, combined with deep learning and reinforcement learning, the negative pressure adsorption force is dynamically adjusted, which solves the problems of wasteful energy consumption and poor adaptability of traditional wall-climbing robots in complex environments, and realizes efficient autonomous navigation and stable adsorption of robots on complex walls.

CN120503902APending Publication Date: 2025-08-19WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510702528.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional wall-climbing robots rely on fixed negative pressure control systems in high altitude operations and complex environments, resulting in waste of energy consumption or adsorption failure, and have poor adaptability in complex terrain, low motion efficiency and risk of slippage.

Method used

The hybrid model control method is adopted to collect environmental and state information in real time through multi-source sensors, use deep learning and reinforcement learning models to predict future adsorption state changes, dynamically adjust the negative pressure adsorption force, and combine multimodal path planning to optimize action selection to realize adaptive adsorption of robots on complex walls.

Benefits of technology

It significantly improves the adsorption stability and environmental adaptability of the robot in complex wall environments, achieves safe and efficient autonomous navigation, and optimizes the dynamic coupling relationship between energy consumption and adsorption force.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120503902A_ABST
    Figure CN120503902A_ABST
Patent Text Reader

Abstract

The invention provides a hybrid model control method and system applied to a negative pressure adsorption type wall-climbing robot, and relates to the technical field of robotics.The method comprises the steps that wall surface environment information, robot state information and energy consumption information are collected in real time through a multi-source sensing module; performing feature extraction on the acquired wall surface environment information by using a deep learning model to obtain a multi-dimensional feature vector representing wall surface features; predicting a future adsorption state change trend through a neural network model based on the multi-dimensional feature vector, the robot state information and the energy consumption information; constructing a reinforcement learning model by using a current state space and an action space, optimizing action selection by comprehensively considering a reward function of state information and energy consumption information of the robot, and dynamically adjusting negative pressure to maintain stable adsorption between the robot and a wall surface; an environment state space is constructed based on multi-mode sensor data, real-time feedback of dynamic negative pressure adjustment is combined, an optimal path is generated through a deep reinforcement learning model, and a control instruction is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to a hybrid model control method and system applied to a negative pressure adsorption wall-climbing robot. Background Art

[0002] Currently, operations in high-altitude and extreme environments are primarily manual labor. This manual approach not only presents challenges such as low efficiency and high costs, but also poses significant safety risks for operators, resulting in numerous casualties. With the increasing number of high-rise buildings and large ships, robots are facing challenges in cleaning, maintenance, welding, and inspection, among other tasks. Robots must maintain safe operation while navigating high-rise walls.

[0003] However, existing wall-climbing robot technology still has some shortcomings in applications such as high-altitude operations and complex environments. In terms of energy efficiency and endurance, traditional wall-climbing robots rely on a fixed negative pressure control system and cannot dynamically adjust the suction force based on wall roughness, tilt angle, and other factors, which can easily lead to energy waste or suction failure. Furthermore, traditional steering wheels have poor adaptability in complex terrain, and path planning algorithms fail to consider the dynamic coupling between energy consumption and suction force, resulting in low motion efficiency and the risk of slippage. Summary of the Invention

[0004] The purpose of the present invention is to provide a hybrid model control method and system for a negative pressure adsorption wall-climbing robot, so as to solve the problems mentioned in the above background technology that traditional wall-climbing robots rely on a fixed negative pressure control system, which easily leads to energy waste or adsorption failure, poor adaptability in complex terrain, resulting in low movement efficiency and the risk of slippage.

[0005] To achieve the above-mentioned objectives, the present invention provides the following technical solutions: a hybrid model control method for a negative pressure adsorption wall-climbing robot, wherein the negative pressure adsorption wall-climbing robot includes a fan, which is located at the bottom of the robot. When the robot moves on the wall, the rotation of the fan forms a negative pressure adsorption force between the robot and the wall. The method steps include: real-time acquisition of wall environment information, robot state information and energy consumption information through a multi-source sensing module; feature extraction of the acquired wall environment information using a deep learning model to obtain a multi-dimensional feature vector characterizing the wall characteristics; based on the multi-dimensional feature vector, robot state information and energy consumption information, predicting the future adsorption state change trend through a neural network model; constructing a reinforcement learning model with the current state space and action space, optimizing action selection through a reward function that comprehensively considers the robot state information and energy consumption information, and dynamically adjusting the negative pressure to maintain stable adsorption between the robot and the wall; constructing an environmental state space based on multimodal sensor data, combining real-time feedback of dynamic negative pressure adjustment, generating an optimal path through a deep reinforcement learning model, and outputting control instructions.

[0006] Optionally, the wall environment information includes wall texture feature images, surface roughness data, environmental parameters and hazardous gas concentrations; the robot status information includes the pressure difference between the inside and outside of the negative pressure chamber, the robot's tilt angle and vibration amplitude, the four-steering wheel contact force distribution and the actual movement speed; the energy consumption information includes the real-time power consumption of the fan, the remaining power and the predicted endurance time based on historical energy consumption data.

[0007] Optionally, the step of using a deep learning model to extract features from the collected wall environment information to obtain a multi-dimensional feature vector characterizing the wall characteristics specifically includes: inputting the wall texture feature image into a pre-trained convolutional neural network model, processing it through a 5-layer convolution layer with a 3×3 convolution kernel and a ReLU activation function, and finally outputting a 128-dimensional feature vector to characterize texture roughness and crack density.

[0008] Optionally, the step of predicting the future adsorption state change trend through a neural network model based on the multidimensional feature vector, robot state information and energy consumption information specifically includes: inputting the 128-dimensional feature vector, historical pressure difference sequence and robot movement speed into a bidirectional long-short-term memory network to predict the pressure difference change rate and wall adhesion risk level within a set time period in the future.

[0009] Optionally, the bidirectional long short-term memory network adopts a two-layer structure, each layer contains 256 hidden units, and an anti-overfitting mechanism is set; the reward function includes a pressure difference deviation index, an energy consumption index and a vibration amplitude index, and the weight of each index can be dynamically adjusted.

[0010] Optionally, the reinforcement learning model is constructed with the current state space and action space, the action selection is optimized by comprehensively considering the reward function of the robot state information and energy consumption information, and the steps of dynamically adjusting the negative pressure to maintain stable adsorption between the robot and the wall specifically include: the state space includes the pressure difference between the inside and outside of the negative pressure chamber, the inclination angle of the robot, the roughness of the wall contact surface, and the wall adhesion risk level output by the model; the action space includes the fan speed adjustment and the opening adjustment of the negative pressure chamber pressure relief valve; the Q-learning model is constructed with the current state space and action space, the expected cumulative reward value is updated according to the reward function considering the robot state information and energy consumption information, and then the control instructions are output to dynamically adjust the negative pressure.

[0011] Optionally, the steps of constructing an environmental state space based on multimodal sensor data, combining real-time feedback of dynamic negative pressure regulation, and generating an optimal path through a deep reinforcement learning model specifically include: constructing a multimodal state space, wherein the multimodal state space includes robot state information, wall environment information, wall adhesion risk level, and energy consumption information, which is used to characterize the robot's posture, terrain characteristics, adhesion requirements and energy status in a complex environment in real time; defining a composite action space, wherein the composite action space includes steering wheel control parameters and negative pressure regulation parameters, which are used to realize the robot's motion control and ground adaptability adjustment; designing a multi-objective reward function, wherein the reward function integrates navigation efficiency, energy consumption optimization and safe obstacle avoidance goals, and dynamically adjusts the weight of each goal according to the remaining power and task urgency; using the TD3 reinforcement learning algorithm to optimize the path planning strategy, and generating a motion control strategy that adapts to complex environments through interactive training.

[0012] Optionally, the multimodal state space specifically includes four-steering wheel pressure sensor array data, surface roughness and friction coefficient estimation values, robot tilt angle, RGB-D environmental perception data, hazardous gas concentration vector, wall adhesion risk level, current energy consumption and remaining battery life; wherein, the RGB-D environmental perception data includes color image information and corresponding three-dimensional depth information, which is used to detect obstacles, passable areas and terrain features in the environment in real time.

[0013] Optionally, the interactive training steps specifically include: randomly extracting batches of quadruple data from the cache, each quadruple containing the current state, executed action, immediate reward and next state; generating the action of the next state through the target policy network, and superimposing Gaussian noise to enhance exploratory power, while limiting the action space to a preset range; using a dual Critic network to calculate the Q value of the next state and the target action respectively, taking the smaller value and combining it with the discount factor and the immediate reward to generate the target Q value; calculating the time difference error between the predicted Q value and the target Q value of the current Critic network, and reversely updating the Critic network parameters based on the batch average error; updating the policy network every preset round, and maximizing the Q value output by the Critic network through gradient ascent; using a soft update mechanism for the target policy network and the target Critic network, and mixing the parameters of the current network and the target network in a fixed ratio.

[0014] On the other hand, the present invention also provides a hybrid model control system for a negative pressure adsorption wall-climbing robot, wherein the negative pressure adsorption wall-climbing robot includes a fan, which is located at the bottom of the robot. When the robot moves on the wall, the rotation of the fan forms a negative pressure adsorption force between the robot and the wall. The control system also includes: an acquisition module for real-time acquisition of wall environment information, robot state information and energy consumption information through a multi-source sensing module; a feature extraction module for using a deep learning model to perform feature extraction on the collected wall environment information to obtain a multi-dimensional feature vector characterizing the wall characteristics; an adsorption state prediction module for predicting future adsorption state change trends through a hybrid model based on the multi-dimensional feature vector, robot state information and energy consumption information; an adsorption state adjustment module for constructing a reinforcement learning model with the current state space and action space, optimizing action selection through a reward function that comprehensively considers the robot state information and energy consumption information, and dynamically adjusting the negative pressure to maintain stable adsorption between the robot and the wall; a path optimization module for constructing an environmental state space based on multimodal sensor data, combining real-time feedback of dynamic negative pressure adjustment, generating an optimal path through a deep reinforcement learning algorithm, and outputting control instructions.

[0015] Compared with the prior art, the present invention has the following beneficial effects:

[0016] This application dynamically adjusts negative pressure using a hybrid model. It uses visual sensors to input environmental features, and a deep learning model, a convolutional neural network (CNN), to extract environmental and surface features. A bidirectional long short-term memory (Bi-LSTM) neural network model is used for time series prediction, predicting environmental state changes within the next three seconds. Based on environmental changes, such as changes in surface roughness, the robot's next action is optimized using a reinforcement learning strategy layer (Q-Learning) to calculate the optimal negative pressure for optimal negative pressure regulation. Through multi-source sensor fusion and hybrid model control, the robot achieves adaptive negative pressure adsorption on complex surface environments, such as smooth, rough, and inclined surfaces, significantly improving adsorption stability and environmental adaptability. This application upgrades traditional passive negative pressure control to active intelligent regulation using a "perception-prediction-decision" approach. By combining dynamic negative pressure regulation with multimodal path planning and considering the dynamic coupling relationship between energy consumption and adsorption force, the robot achieves safe and efficient autonomous navigation in complex terrain, achieving coordinated optimization of dynamic negative pressure regulation and path planning, and thus improving the robot's adaptability on different surface types. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the process steps of the present invention.

[0018] Figure 2 This is a schematic structural diagram of the negative pressure adsorption wall-climbing robot of the present invention.

[0019] Figure 3 It is a schematic diagram of the structure of the negative pressure adsorption fan of the present invention.

[0020] Figure 4 This is a flow chart of the dynamic negative pressure regulation method based on the hybrid model of the present invention.

[0021] Figure 5 This is a framework diagram of the multimodal deep reinforcement learning path planning algorithm of the present invention.

[0022] Figure 6 Schematic diagram of the system structure of the present invention.

[0023] In the figure: 1-robot body, 2-fan, 21-drive motor, 22-upper outer cover, 23-coupling, 24-fan blade, 25-lower outer cover, 3-camera, 4-steering wheel, 5-photovoltaic panel, 6-micro wind turbine, 10-acquisition module, 20-feature extraction module, 30-adsorption state prediction module, 40-adsorption state adjustment module, 50-path optimization module. DETAILED DESCRIPTION

[0024] The following will provide a clear and complete description of the solutions of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0027] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0028] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.

[0029] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0030] Please refer to Figure 1-Figure 5 The present invention provides a hybrid model control method for a negative pressure adsorption wall-climbing robot, wherein the negative pressure adsorption wall-climbing robot includes a fan, which is located at the bottom of the robot. When the robot moves on the wall, the rotation of the fan forms a negative pressure adsorption force between the robot and the wall. The method steps include: collecting wall environment information, robot state information and energy consumption information in real time through a multi-source sensing module; using a deep learning model to perform feature extraction on the collected wall environment information to obtain a multi-dimensional feature vector that characterizes the wall characteristics; based on the multi-dimensional feature vector, robot state information and energy consumption information, predicting the future adsorption state change trend through a neural network model; constructing a reinforcement learning model with the current state space and action space, optimizing action selection through a reward function that comprehensively considers the robot state information and energy consumption information, and dynamically adjusting the negative pressure to maintain stable adsorption between the robot and the wall; constructing an environmental state space based on multimodal sensor data, combining real-time feedback of dynamic negative pressure adjustment, generating an optimal path through a deep reinforcement learning model, and outputting control instructions.

[0031] Specifically, this application dynamically adjusts negative pressure using a hybrid model. It uses visual sensors to input environmental features, and a deep learning model, convolutional neural network (CNN), to extract environmental and surface features. A bidirectional long short-term memory (Bi-LSTM) neural network model is used for time series prediction, predicting environmental state changes within the next three seconds. Based on environmental changes, such as changes in surface roughness, the robot's next action is calculated using a reinforcement learning strategy layer (Q-Learning) to optimize the Q-Learning strategy. This optimizes negative pressure regulation. Through multi-source sensor fusion and hybrid model control, the robot achieves adaptive negative pressure adsorption on complex surface environments, such as smooth, rough, and inclined surfaces, significantly improving adsorption stability and environmental adaptability. This upgrades traditional passive negative pressure control to active intelligent regulation based on a "perception-prediction-decision" approach. By combining dynamic negative pressure regulation with multimodal path planning and considering the dynamic coupling relationship between energy consumption and adsorption force, the robot achieves safe and efficient autonomous navigation in complex terrain, achieving coordinated optimization of dynamic negative pressure regulation and path planning, and thus improving its adaptability on different surface types.

[0032] Optionally, the wall environment information includes wall texture feature images, surface roughness data, environmental parameters and hazardous gas concentrations; the robot status information includes the pressure difference between the inside and outside of the negative pressure chamber, the robot's tilt angle and vibration amplitude, the four-steering wheel contact force distribution and the actual movement speed; the energy consumption information includes the real-time power consumption of the fan, the remaining power and the predicted endurance time based on historical energy consumption data.

[0033] Specifically, the wall environment information includes: wall texture feature images acquired by a high-resolution RGB-D camera; surface roughness data measured by a laser triangulation rangefinder; environmental parameters acquired by temperature and humidity sensors; and hazardous gas concentrations detected by multi-gas sensors. The robot status information includes: the pressure difference between the inside and outside of the negative pressure chamber acquired by a distributed pressure sensor array; the robot's tilt angle and vibration amplitude measured by a 9-axis IMU (accelerometer, gyroscope, magnetometer); the contact force distribution acquired by the four steering wheel embedded pressure sensors; and the actual motion speed calculated by the encoder. The energy consumption information includes: the real-time power consumption of the fan acquired by the power monitoring module; the remaining power monitored by the battery management system; and the battery life predicted based on historical energy consumption data. Through the collaborative work of multiple types of sensors, including vision, mechanics, and environment, a comprehensive environmental perception and status monitoring network is constructed, providing a high-precision data foundation for intelligent control. This allows for real-time acquisition of multi-dimensional information, including wall physical properties, robot motion status, and energy consumption, to support precise decision-making by subsequent algorithms.

[0034] Optionally, the step of using a deep learning model to extract features from the collected wall environment information to obtain a multi-dimensional feature vector characterizing the wall characteristics specifically includes: inputting the wall texture feature image into a pre-trained convolutional neural network model, processing it through a 5-layer convolution layer with a 3×3 convolution kernel and a ReLU activation function, and finally outputting a 128-dimensional feature vector to characterize texture roughness and crack density.

[0035] Specifically, through 5 layers of 3×3 convolutional kernels and ReLU activation function processing, a 128-dimensional feature vector is finally output, which can effectively extract the deep features of the wall texture, such as roughness or crack distribution. The lightweight network structure runs efficiently on embedded devices and meets real-time requirements.

[0036] Optionally, the step of predicting the future adsorption state change trend through a neural network model based on the multidimensional feature vector, robot state information and energy consumption information specifically includes: inputting the 128-dimensional feature vector, historical pressure difference sequence and robot movement speed into a bidirectional long-short-term memory network to predict the pressure difference change rate and wall adhesion risk level within a set time period in the future.

[0037] Specifically, the 128-dimensional feature vector obtained from the feature extraction layer at the current moment, the historical 10-second pressure difference sequence, and the robot motion speed obtained by the 9-axis IMU are input into the time modeling layer. Through a two-layer bidirectional LSTM network with 256 hidden units in each layer, the 128-dimensional wall feature vector, the historical 10-second pressure difference sequence and the IMU real-time motion speed are multimodally fused to achieve accurate prediction of the environmental state within the next 3 seconds; the wall texture features, dynamic pressure difference sequence and motion state are jointly input to overcome the limitations of traditional single-signal prediction, and the correlation between historical and future trends is captured through the Bi-LSTM forward and reverse bidirectional method. The Dropout layer with a 30% neuron discarding probability is combined to suppress noise interference, improve prediction robustness, and synchronously output the pressure difference change rate and attachment risk probability to provide a high-precision time series basis for reinforcement learning decision-making.

[0038] Optionally, the bidirectional long short-term memory network adopts a two-layer structure, each layer contains 256 hidden units, and an anti-overfitting mechanism is set; the reward function includes a pressure difference deviation index, an energy consumption index and a vibration amplitude index, and the weight of each index can be dynamically adjusted.

[0039] Specifically, the bidirectional network structure simultaneously considers historical and future trend correlation information to improve prediction accuracy, predicts future adsorption state change trends through time series modeling, and adjusts control parameters in advance to avoid stability problems caused by sudden changes in adsorption force.

[0040] Optionally, the reinforcement learning model is constructed with the current state space and action space, the action selection is optimized by comprehensively considering the reward function of the robot state information and energy consumption information, and the steps of dynamically adjusting the negative pressure to maintain stable adsorption between the robot and the wall specifically include: the state space includes the pressure difference between the inside and outside of the negative pressure chamber, the inclination angle of the robot, the roughness of the wall contact surface, and the wall adhesion risk level output by the model; the action space includes the fan speed adjustment and the opening adjustment of the negative pressure chamber pressure relief valve; the Q-learning model is constructed with the current state space and action space, the expected cumulative reward value is updated according to the reward function considering the robot state information and energy consumption information, and then the control instructions are output to dynamically adjust the negative pressure.

[0041] Specifically, the state space is Where ΔP is the pressure difference between the inside and outside of the negative pressure chamber, and θ is the robot's tilt angle. The action space includes fan speed adjustment ±20% of the current value and the negative pressure relief valve opening range of 0%-100%. The reward function is calculated as follows: Where R is the comprehensive reward value, ω1 is the negative pressure stability weight, ΔP is the pressure difference between the inside and outside of the negative pressure chamber, and P target is the ideal negative pressure value that the system needs to maintain, ω2 is the energy efficiency weight, e is the Euler number, E poweris the energy consumption per unit time of the fan, ω3 is the vibration suppression weight, and F is the vibration amplitude; the expected cumulative reward value Q-table update rule formula is: Where Q(s,a) is the expected cumulative reward Q value for taking action a in state s, α is the learning rate, r is the reward value calculated by the reward function, γ is the discount factor, α′ and s′ are the next action and the next state respectively. is the maximum Q value of all possible actions for the new state s′. This application achieves global optimization of the control strategy by balancing adsorption performance, energy consumption, and motion stability through a multi-objective reward function. A dynamic weighting mechanism is used to autonomously adjust control priorities based on task requirements, enhancing system flexibility. This upgrades the passive response of traditional negative pressure control to an integrated intelligent control system of "perception-prediction-decision-making," enabling dynamic regulation of negative pressure, thereby enabling the robot to adhere to the wall more stably and utilize energy more efficiently.

[0042] Optionally, the steps of constructing an environmental state space based on multimodal sensor data, combining real-time feedback of dynamic negative pressure regulation, and generating an optimal path through a deep reinforcement learning model specifically include: constructing a multimodal state space, wherein the multimodal state space includes robot state information, wall environment information, wall adhesion risk level, and energy consumption information, which is used to characterize the robot's posture, terrain characteristics, adhesion requirements and energy status in a complex environment in real time; defining a composite action space, wherein the composite action space includes steering wheel control parameters and negative pressure regulation parameters, which are used to realize the robot's motion control and ground adaptability adjustment; designing a multi-objective reward function, wherein the reward function integrates navigation efficiency, energy consumption optimization and safe obstacle avoidance goals, and dynamically adjusts the weight of each goal according to the remaining power and task urgency; using the TD3 reinforcement learning algorithm to optimize the path planning strategy, and generating a motion control strategy that adapts to complex environments through interactive training.

[0043] Specifically, the multimodal state space is constructed as: S t =[P s , R s ,θ t , V d , G d , E p , F], where S t is the multimodal state space vector, P s ∈R 4 It is a four-steering wheel pressure sensor array, including contact force and slip detection, R s ∈R 2 is the estimated value of surface roughness index and friction coefficient, θ t ∈R 3 is the three-axis attitude angle, V d ∈R {H×W×3}G is the RGB-D environment perception data collected by the camera, d ∈R 3 is the dangerous gas concentration vector, E p ∈R 2 is the current energy consumption and the remaining endurance time, and F is the wall adhesion risk level.

[0044] The composite action space is defined as: A t =[Δω,Δα,F α ] T , where A t is the composite action space vector, Δω∈R 4 is the angular velocity increment of the four steering wheels, Δα∈[-π,π] 4 is the steering wheel steering angle correction, F a ∈R 2 is the dynamic negative pressure adjustment parameter [P v ,Q v ] T , that is, pressure value and flow rate.

[0045] The navigation reward function is: Where R nav is the navigation reward, d is the Euclidean distance from the current pose to the target, which can be obtained through the visual SLAM module and GPS sensor, θ err is the deviation angle between the current heading and the target direction, which can be calculated by the IMU attitude sensor. λ1,λ2,k are adjustment coefficients.

[0046] The energy consumption reward function is: Where R energy is the energy consumption reward, Δω i is the angular velocity change of the i-th steering wheel, measured by the steering wheel encoder, ΔP v is the pressure adjustment value of the negative pressure system, obtained through the pressure sensor, I slip is the slip indication function. When the pressure sensor array detects wheel slip, I slip =1, γ1, γ2, γ3 are the corresponding penalty coefficients.

[0047] The safety reward function is: Where R safety For safety rewards, G d is the concentration vector of dangerous gas, obtained by gas sensor, I collision For collision indication, when the lidar or ultrasonic sensor detects an obstacle contact, I collision =1,σ(P s ) is the standard deviation of the pressure distribution on the four steering wheels, calculated by the pressure sensor array, and η1, η2, and η3 are the adjustment coefficients.

[0048] The multi-objective reward function is: R(S,A)=ω1·R nav +ω2·R energy +ω3·R safety , where ω1, ω2, ω3 are corresponding weight values. The weights ω1, ω2, ω3 can be adaptively adjusted according to the remaining power and task urgency. nav is the navigation reward, R energy is the energy consumption reward, R safety Reward for safety.

[0049] First define the power factor Remaining power ratio, emergency factor Among them, E remaining is the current remaining power, E total is the total capacity of the battery, t remaining ,t critical They represent the remaining time of the task and the critical time threshold respectively, e is a constant, and k is the slope parameter of the Sigmoid function, then we can get: ω3=1-ω1-ω2; where ω1 is the navigation reward weight value, ω2 is the energy consumption reward weight value, and ω3 is the safety reward weight value. The default initial weight for navigation rewards, The default initial weight for energy consumption rewards.

[0050] This application integrates the environment, robot status, adhesion requirements and energy consumption information through multimodal state space, deeply integrates path planning and dynamic adjustment of adsorption, generates a global optimal path that takes into account navigation efficiency, safety and energy consumption, and realizes coordinated control of steering wheel movement and negative pressure regulation through composite action space, avoiding system performance bottlenecks caused by single parameter optimization.

[0051] Optionally, the multimodal state space specifically includes four-steering wheel pressure sensor array data, surface roughness and friction coefficient estimation values, robot tilt angle, RGB-D environmental perception data, hazardous gas concentration vector, wall adhesion risk level, current energy consumption and remaining battery life; wherein, the RGB-D environmental perception data includes color image information and corresponding three-dimensional depth information, which is used to detect obstacles, passable areas and terrain features in the environment in real time.

[0052] Optionally, the interactive training steps specifically include: randomly extracting batches of quadruple data from the cache, each quadruple containing the current state, executed action, immediate reward and next state; generating the action of the next state through the target policy network, and superimposing Gaussian noise to enhance exploratory power, while limiting the action space to a preset range; using a dual Critic network to calculate the Q value of the next state and the target action respectively, taking the smaller value and combining it with the discount factor and the immediate reward to generate the target Q value; calculating the time difference error between the predicted Q value and the target Q value of the current Critic network, and reversely updating the Critic network parameters based on the batch average error; updating the policy network every preset round, and maximizing the Q value output by the Critic network through gradient ascent; using a soft update mechanism for the target policy network and the target Critic network, and mixing the parameters of the current network and the target network in a fixed ratio.

[0053] This application significantly improves the convergence speed and robustness of the algorithm through a dual critic network and experience replay mechanism, avoids control incoherence caused by strategy oscillation, effectively solves the problem of response lag of traditional methods in dynamic environments, and balances exploration and utilization efficiency through noise injection and action space restriction, accelerating the adaptive learning of the model in real environments.

[0054] On the other hand, please refer to Figure 6 The present invention also provides a hybrid model control system for a negative pressure adsorption wall climbing robot, wherein the negative pressure adsorption wall climbing robot includes a fan, which is located at the bottom of the robot. When the robot moves on the wall, the rotation of the fan forms a negative pressure adsorption force between the robot and the wall. The control system also includes: an acquisition module 10 for collecting wall environment information, robot state information and energy consumption information in real time through a multi-source sensing module; a feature extraction module 20 for extracting features from the collected wall environment information using a deep learning model to obtain a multi-dimensional feature vector that characterizes the characteristics of the wall; and an adsorption state prediction module. 30, used to predict the future adsorption state change trend through a hybrid model based on the multidimensional feature vector, robot state information and energy consumption information; the adsorption state adjustment module 40, used to construct a reinforcement learning model with the current state space and action space, optimize the action selection through a reward function that comprehensively considers the robot state information and energy consumption information, and dynamically adjust the negative pressure to maintain stable adsorption between the robot and the wall; the path optimization module 50, used to construct an environmental state space based on multimodal sensor data, combine the real-time feedback of dynamic negative pressure adjustment, generate the optimal path through a deep reinforcement learning algorithm, and output control instructions.

[0055] Specifically, the negative pressure adsorption wall-climbing robot includes a robot body 1, a fan 2 is installed inside the robot body 1, and the fan 2 is located at the bottom of the robot body 1. When the robot body 1 moves on the wall, the rotation of the fan 2 forms a negative pressure adsorption force between the robot body 1 and the wall.

[0056] The upper housing 22 of the fan 2 is fixedly connected to the top of the robot body 1, and the lower housing 25 of the fan 2 is fixedly connected to the bottom of the robot body 1. The drive motor 21 is fixed at the center of the upper housing 22, and the fan blades 24 are fixedly connected to the output end of the drive motor via a coupling 23. The negative pressure suction force of the fan is adjusted by adjusting the rotation speed of the motor.

[0057] The negative pressure adsorption wall-climbing robot is also equipped with a sensor 3, a photovoltaic panel 5 and a micro wind generator.

[0058] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0059] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. In the embodiments provided by the present invention, any reference to memory, storage, database, or other media can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0060] The above are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A hybrid model control method for a negative pressure adsorption wall-climbing robot, wherein the negative pressure adsorption wall-climbing robot includes a fan located at the bottom of the robot. When the robot moves on a wall, the rotation of the fan generates a negative pressure adsorption force between the robot and the wall. The method is characterized in that: The method steps include: Through multi-source sensing modules, real-time collection of wall environment information, robot status information and energy consumption information; Use deep learning models to extract features from the collected wall environment information and obtain multi-dimensional feature vectors that represent the wall characteristics; Based on the multidimensional feature vector, robot state information and energy consumption information, predicting the future adsorption state change trend through a neural network model; A reinforcement learning model is constructed based on the current state space and action space. The reward function, which comprehensively considers the robot's state information and energy consumption information, optimizes action selection and dynamically adjusts the negative pressure to maintain stable adhesion between the robot and the wall. The environmental state space is constructed based on multimodal sensor data, combined with real-time feedback of dynamic negative pressure regulation, and the optimal path is generated through a deep reinforcement learning model, and control instructions are output.

2. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 1, characterized in that: The wall surface environment information includes wall surface texture feature images, surface roughness data, environmental parameters and hazardous gas concentrations; The robot status information includes the pressure difference between the inside and outside of the negative pressure chamber, the robot's tilt angle and vibration amplitude, the contact force distribution of the four steering wheels, and the actual movement speed; The energy consumption information includes the real-time power consumption of the fan, the remaining power, and the life time predicted based on historical energy consumption data.

3. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 2, characterized in that: The step of extracting features from the collected wall environment information using a deep learning model to obtain a multi-dimensional feature vector representing the wall characteristics specifically includes: The wall texture feature image is input into a pre-trained convolutional neural network model, processed through 5 layers of 3×3 convolution kernels and ReLU activation function, and finally outputs a 128-dimensional feature vector to characterize texture roughness and crack density.

4. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 3, characterized in that: The step of predicting the future adsorption state change trend through a neural network model based on the multidimensional feature vector, robot state information and energy consumption information specifically includes: The 128-dimensional feature vector, historical pressure difference sequence, and robot motion speed are input into a bidirectional long short-term memory network to predict the pressure difference change rate and wall adhesion risk level within a future set time period.

5. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 4, characterized in that: The bidirectional long short-term memory network adopts a two-layer structure, each layer contains 256 hidden units, and is equipped with an anti-overfitting mechanism; The reward function includes a pressure difference deviation index, an energy consumption index, and a vibration amplitude index, and the weight of each index can be dynamically adjusted.

6. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 5, characterized in that: The steps of constructing a reinforcement learning model based on the current state space and action space, optimizing action selection through a reward function that comprehensively considers the robot's state information and energy consumption information, and dynamically adjusting the negative pressure to maintain stable adsorption between the robot and the wall specifically include: The state space includes the pressure difference between the inside and outside of the negative pressure chamber, the robot's tilt angle, the wall contact surface roughness, and the wall adhesion risk level output by the model; The action space includes fan speed adjustment and negative pressure chamber pressure relief valve opening adjustment; A Q-learning model is constructed based on the current state space and action space. The expected cumulative reward value is updated according to the reward function that considers the robot's state information and energy consumption information, and then control instructions are output to dynamically adjust the negative pressure.

7. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 6, characterized in that: The steps of constructing an environmental state space based on multimodal sensor data, combining real-time feedback of dynamic negative pressure regulation, and generating an optimal path through a deep reinforcement learning model specifically include: Constructing a multimodal state space that includes robot state information, wall environment information, wall adhesion risk level, and energy consumption information to characterize the robot's position, terrain characteristics, adhesion requirements, and energy status in a complex environment in real time; Defining a composite motion space, wherein the composite motion space includes steering wheel control parameters and negative pressure adjustment parameters, for realizing motion control and ground adaptability adjustment of the robot; Design a multi-objective reward function that integrates navigation efficiency, energy optimization, and safe obstacle avoidance, and dynamically adjusts the weights of each objective based on remaining battery power and mission urgency. The TD3 reinforcement learning algorithm is used to optimize the path planning strategy, and a motion control strategy that adapts to complex environments is generated through interactive training.

8. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 7, characterized in that: The multimodal state space specifically includes the four-wheel pressure sensor array data, surface roughness and friction coefficient estimates, robot tilt angle, RGB-D environmental perception data, hazardous gas concentration vector, wall adhesion risk level, and current energy consumption and remaining flight time; The RGB-D environmental perception data includes color image information and corresponding three-dimensional depth information, which is used to detect obstacles, traversable areas and terrain features in the environment in real time.

9. The hybrid model control method for a negative pressure adsorption wall-climbing robot according to claim 8, characterized in that: The steps of the interactive training specifically include: Randomly extract batches of four-tuple data from the cache, each of which contains the current state, executed action, immediate reward, and next state; The target policy network generates the action for the next state, and Gaussian noise is superimposed to enhance exploration, while the action space is limited to a preset range. A dual critic network is used to calculate the Q value of the next state and the target action respectively, and the smaller value is taken and combined with the discount factor and the immediate reward to generate the target Q value; Calculate the time difference error between the predicted Q value of the current critic network and the target Q value, and reversely update the critic network parameters based on the batch average error; Update the policy network every preset rounds and maximize the Q value output by the critic network through gradient ascent; A soft update mechanism is adopted for the target policy network and the target critic network, and the parameters of the current network and the target network are mixed in a fixed ratio.

10. A hybrid model control system for a negative pressure adsorption wall-climbing robot, wherein the negative pressure adsorption wall-climbing robot includes a fan located at the bottom of the robot. When the robot moves on the wall, the rotation of the fan generates a negative pressure adsorption force between the robot and the wall. The invention is characterized in that: Also includes: The acquisition module is used to collect wall environment information, robot status information and energy consumption information in real time through the multi-source sensing module; The feature extraction module is used to extract features from the collected wall environment information using a deep learning model to obtain a multi-dimensional feature vector representing the wall characteristics; an adsorption state prediction module, configured to predict future adsorption state change trends through a hybrid model based on the multidimensional feature vector, robot state information, and energy consumption information; The adsorption state adjustment module is used to build a reinforcement learning model based on the current state space and action space. It optimizes action selection through a reward function that comprehensively considers the robot's state information and energy consumption information, and dynamically adjusts the negative pressure to maintain stable adsorption between the robot and the wall. The path optimization module is used to construct the environmental state space based on multimodal sensor data, combine the real-time feedback of dynamic negative pressure regulation, generate the optimal path through the deep reinforcement learning algorithm, and output control instructions.