Depth visual identification-based unhooking robot control method and system

Through the hook-removing robot control method based on deep visual recognition, dynamic focus, frequency domain fusion algorithm and neuromorphic spatiotemporal motion modeling are adopted to solve the bottleneck problem of the fusion efficiency of multi-source heterogeneous data and the collaborative optimization of dynamic coupling systems in the existing technology, high-precision trajectory prediction and robotic arm operation domain construction are realized, collision risk is reduced, and collaborative motion control of biological intelligence is realized.

CN120095829AInactive Publication Date: 2025-06-06ANHUI HUADIAN SUZHOU POWER GENERATION

Patent Information

Application Number
CN202510503256.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has bottlenecks in the efficiency of multi-source heterogeneous data fusion and coordinated optimization of dynamic coupling systems. Especially in highly dynamic unstructured scenarios, it is difficult to take into account perception accuracy, response speed and system robustness.

Method used

The hook-removing robot control method based on deep visual recognition is adopted to generate the target pose sequence through dynamic focus and frequency domain fusion algorithms, and combined with neuromorphic spatiotemporal motion modeling, the prediction trajectory of the dynamic coupling target body and the spatiotemporal constraint domain of the robot end effector are modeled, and the complete three-dimensional model is obtained through tensor decomposition, and the optimal grasping position and operational health status are finally obtained through impedance model and genetic algorithm optimization.

Benefits of technology

It significantly improves the accuracy of trajectory prediction, builds an operation domain that meets the physical limits of the robotic arm, optimizes the efficiency of obstacle avoidance path planning, reduces the risk of collision during movement, and realizes collaborative motion control of biological intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120095829A_ABST
    Figure CN120095829A_ABST
Patent Text Reader

Abstract

The invention discloses an unhooking robot control method and system based on depth visual recognition, and relates to the field of visual recognition, and the method comprises the steps: obtaining an original optical signal and ambient light intensity, carrying out the dynamic focusing of the original optical signal, generating a target pose sequence according to a frequency domain fusion algorithm and a gradient maximization positioning algorithm, and carrying out the dynamic focusing of the original optical signal; meanwhile, the environment light intensity is substituted into adjacent frame pose difference, and a real-time speed vector is obtained; performing modeling on the target pose sequence and the real-time velocity vector by adopting neuromorphic space-time motion to obtain a predicted trajectory of a dynamic coupling target body and a space-time constraint domain of an end effector of the robot; a bionics principle is adopted to simulate a biological motion coordination mechanism, a target motion track and mechanical arm dynamic constraint are jointly described, and the limitation that track prediction and mechanical arm motion ability are disjointed in traditional isolated modeling is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual recognition, and in particular to a method and system for controlling a hook-removing robot based on deep visual recognition. Background Art

[0002] As industrial automation evolves towards precision and intelligence, robot grasping control technology based on deep vision has become the core research direction in the field of intelligent manufacturing. The current technology system mainly revolves around three dimensions: multimodal perception, motion prediction modeling, and adaptive control. At the visual perception level, metasurface imaging technology gradually replaces traditional optical lens systems with its subwavelength phase control capability, and its dynamically adjustable characteristics provide a new way for target recognition in complex light field environments. In the field of motion modeling, neuromorphic computing models significantly improve the real-time prediction of dynamic target trajectories by simulating the spatiotemporal coding mechanism of biological neurons. In terms of control strategies, the combination of impedance control and fault-tolerant mechanisms has become the mainstream. However, existing technologies still have significant bottlenecks in terms of multi-source heterogeneous data fusion efficiency and dynamic coupling system collaborative optimization. Especially in highly dynamic unstructured scenarios, traditional methods are difficult to balance perception accuracy, response speed, and system robustness.

[0003] The defects of existing technologies can be summarized into three core dimensions. First, passive light field perception and low-order data fusion: Traditional visual systems rely on fixed optical architectures and lack the ability to characterize spatiotemporal joint features. They are difficult to adapt to sudden changes in light intensity and dynamic changes in target reflection characteristics, resulting in insufficient modeling of multimodal data correlation; second, decoupled motion modeling and constraint separation: Existing prediction models mostly use isolated trajectory extrapolation methods, which fail to achieve a joint description of the mechanical arm's dynamic limits and the target motion coupling effect, resulting in inaccurate construction of the safe operating domain; third, rigid control strategies and response lags: The current impedance parameters are fixed and nonlinear disturbance compensation is delayed, and the fault-tolerant mechanism lacks adaptive tuning capabilities based on the system degradation state, which significantly reduces robustness in the event of sudden load disturbances or component performance degradation. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a hook-removing robot control method based on deep visual recognition to solve the problems of insufficient real-time posture perception accuracy of dynamic coupling targets in complex light fields and lack of impedance control parameter solidification and system health status adaptive fault tolerance under nonlinear disturbances.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides a control method for a hook-removing robot based on deep visual recognition, which includes obtaining an original light signal and an ambient light intensity, dynamically focusing the original light signal, and generating a target pose sequence according to a frequency domain fusion algorithm and a gradient maximization positioning algorithm, while substituting the ambient light intensity into the pose difference of adjacent frames to obtain a real-time velocity vector; using neuromorphic spatiotemporal motion to model the target pose sequence and the real-time velocity vector, and obtaining a predicted trajectory of a dynamically coupled target body and a spatiotemporal constraint domain of a robot end effector; performing data registration on the predicted trajectory and the spatiotemporal constraint domain to obtain a multi-light spectral data, and perform tensor decomposition on the multispectral data to obtain a complete three-dimensional model. At the same time, the complete three-dimensional model and the space-time constraint domain are feature extracted and substituted into the space-time dominated grasping point optimization function to obtain the optimal grasping posture; six-dimensional force feedback is obtained through the torque sensor in the end effector of the robotic arm, and combined with the optimal grasping posture, the motor control quantity and the actual posture error are obtained through impedance model calculation and actuator dynamic compensation; the actual posture error is substituted into the genetic algorithm to obtain the optimized phase parameter, and the abnormal marking fault recovery protocol in the motor control quantity is used to obtain the operating health status of the unhooking robot.

[0007] As a preferred solution of the control method of the hook-removing robot based on deep visual recognition described in the present invention, the dynamic weight adjustment refers to automatically adjusting the proportion of different data sources in the final calculation result according to the real-time change of the ambient light intensity. .

[0008] As a preferred solution of the control method of the hook-removing robot based on deep visual recognition described in the present invention, the original light signal is dynamically focused, and a target posture sequence is generated according to the frequency domain fusion algorithm and the gradient maximization positioning algorithm, and the ambient light intensity is substituted into the posture difference of adjacent frames to obtain a real-time velocity vector. The steps are as follows: The light wave phase in the original light signal is dynamically focused by real-time phase adjustment, and normalized at the same time to obtain a multi-focus light field sequence after dynamic focusing; Each focal visual imaging frame in the multi-focal light field sequence is subjected to Fourier transformation and optical transfer function filtering, and combined with weighted multi-focal visual imaging frame fusion based on modulation transfer function to obtain a fused high-resolution visual imaging frame; The Sobel operator is used to calculate the gradient amplitude of the fused high-resolution visual imaging frame, and the position with the maximum gradient amplitude is taken as the target center to obtain the single-frame target pose with the gradient direction as the target orientation. The continuous single-frame target poses are time-aligned to obtain the target pose sequence. The dynamic attenuation factor is obtained by dynamically adjusting the weight of the ambient light intensity, and the real-time velocity vector is obtained based on the position difference between the dynamic attenuation factor and the target posture sequence.

[0009] As a preferred solution of the control method of the hook-removing robot based on deep visual recognition described in the present invention, the steps are as follows: neuromorphic spatiotemporal motion is used to model the target posture sequence and the real-time velocity vector, and the predicted trajectory of the dynamically coupled target body and the spatiotemporal constraint domain of the robot end effector are obtained. The target position sequence and real-time velocity vector are spliced ​​and normalized to obtain the preprocessed spatiotemporal feature data, and a neuromorphic spatiotemporal model is constructed through spiking neural network and spatiotemporal convolution. The preprocessed spatiotemporal feature data are substituted into the neuromorphic spatiotemporal model for iterative prediction and Bayesian uncertainty quantification to obtain the predicted trajectory. The quadratic rule is then used to solve the optimal trajectory that satisfies the constraints, and finally the spatiotemporal constraint domain is obtained.

[0010] As a preferred solution of the control method of the hook-removing robot based on deep visual recognition described in the present invention, the complete three-dimensional model and the spatiotemporal constraint domain are subjected to feature extraction, and the spatiotemporal dominated grasping point optimization function is substituted to obtain the optimal grasping posture. The steps are as follows: By calculating edge curvature, surface thermal gradient and accessibility weight, the geometric features, thermodynamic features and space-time constraint parameters in the complete three-dimensional model and space-time constraint domain are extracted respectively; Based on the extracted geometric features, thermodynamic characteristics and spatiotemporal constraint parameters, the optimal grasping posture is obtained by optimizing geometric stability, thermal safety and spatiotemporal accessibility, and searching in the constraint domain using genetic algorithm and gradient descent.

[0011] As a preferred solution of the control method of the hook-removing robot based on deep visual recognition described in the present invention, the six-dimensional force feedback is obtained through the torque sensor in the end effector of the robot arm, combined with the optimal grasping posture, and the motor control amount and the actual posture error are obtained through impedance model calculation and actuator dynamic compensation. The steps are as follows: A torque sensor is installed at the end of the robotic arm to collect six-dimensional force feedback in real time. Combined with the optimal grasping posture, the calibrated six-dimensional force feedback and target contact force are obtained by converting signal filtering and coordinate system. The motor control quantity of the six-dimensional force feedback after the target contact force calibration is calculated, and the position error and attitude error are evaluated to obtain the motor control quantity and the actual posture error.

[0012] As a preferred solution of the control method of the hook-removing robot based on deep visual recognition described in the present invention, the actual posture error is substituted into the genetic algorithm to obtain the optimized phase parameter, and the abnormal marking fault recovery protocol in the motor control quantity is used to obtain the running health status of the hook-removing robot. The steps are as follows: The actual posture error is optimized by genetic algorithm according to the fitness function, and the motor control quantity is monitored in real time through multiple sensors to obtain the optimized phase parameters and abnormal marks; According to the fault recovery protocol, the health status of the optimized phase parameters and abnormal marks is evaluated to obtain the system health status of the unhooking robot.

[0013] In a second aspect, the present invention provides a dehooking robot control system based on deep visual recognition, including: a super-table sensing module, which obtains original light signals and ambient light intensity, dynamically focuses the original light signals, generates a target pose sequence according to a frequency domain fusion algorithm and a gradient maximization positioning algorithm, and substitutes the ambient light intensity into the pose difference of adjacent frames to obtain a real-time velocity vector; The pre-control domain model module uses neuromorphic spatiotemporal motion to model the target pose sequence and real-time velocity vector to obtain the predicted trajectory of the dynamically coupled target body and the spatiotemporal constraint domain of the robot end effector; The tensor optimal grasping module performs data registration between the predicted trajectory and the spatiotemporal constraint domain to obtain multispectral data, and performs tensor decomposition on the multispectral data to obtain a complete three-dimensional model. At the same time, the complete three-dimensional model and the spatiotemporal constraint domain are extracted and substituted into the spatiotemporal dominated grasping point optimization function to obtain the optimal grasping posture. The impedance control module obtains six-dimensional force feedback through the torque sensor in the end effector of the robot arm, combines the optimal grasping posture, calculates the impedance model and dynamically compensates the actuator to obtain the motor control amount and the actual posture error; The evolutionary phase control module substitutes the actual posture error into the genetic algorithm to obtain the optimized phase parameters, and obtains the operating health status of the unhooking robot through the abnormal marking fault recovery protocol in the motor control quantity.

[0014] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for controlling a hooking robot based on deep visual recognition as described in the first aspect of the present invention is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for controlling a hooking robot based on deep visual recognition as described in the first aspect of the present invention.

[0016] The beneficial effects of the present invention are as follows: by simulating the pulse coding mechanism of biological neural networks, integrating visual event streams and inertial sensor data, a motion prediction model with dynamic coupling of space and time is constructed. This step uses the principle of bionics to simulate the coordination mechanism of biological motion, jointly describing the target motion trajectory and the dynamic constraints of the robotic arm, and breaking through the limitation of the disconnection between trajectory prediction and the motion ability of the robotic arm in traditional isolated modeling. The interactive characteristics of the target and the robotic arm are captured through dynamic coupling equations to generate spatiotemporal motion prediction results that include safety boundaries. Its core value lies in significantly improving the accuracy of trajectory prediction, while constructing an operating domain that meets the physical limits of the robotic arm, optimizing the efficiency of obstacle avoidance path planning, reducing the risk of collision during movement, and realizing collaborative motion control similar to biological intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0018] Figure 1 This is a flow chart of the hook-removing robot control method based on deep visual recognition in Example 1.

[0019] Figure 2 This is a flow chart of the neuromorphic pulse-space-time constraint domain in Example 1.

[0020] Figure 3 This is a flow chart of the evolution-fault-tolerant dual-mode adaptive adjustment in Example 1.

[0021] Figure 4 This is a module flow chart of the unhooking robot control system based on deep visual recognition in Example 1. DETAILED DESCRIPTION

[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0025] Example 1, reference Figure 1~Figure 4 , which is the first embodiment of the present invention, and provides a method for controlling a hook-removing robot based on deep visual recognition, comprising the following steps: S1: Obtain the original light signal and ambient light intensity, dynamically focus the original light signal, and generate the target pose sequence based on the frequency domain fusion algorithm and gradient maximization positioning algorithm. At the same time, substitute the ambient light intensity into the pose difference of adjacent frames to obtain the real-time velocity vector.

[0026] S1.1: A tunable phase delay network is constructed using a nano-optical metasurface array. Voltage is applied to different areas to control the local phase response of the tunable phase delay network, so that the incident light forms a multi-view light field map in space. The multi-view light field map is constructed in real time for multiple focal depth layers to obtain a time-continuous multi-focal light field sequence. At the same time, a linear normalization method is used to process the grayscale moment of the original imaging frame to obtain a multi-focal light field sequence.

[0027] The spectrum of each focal imaging frame of the multi-focal light field sequence is analyzed to extract the frequency domain information of the focal visual imaging frame. At the same time, the spectrum is filtered in combination with the optical transfer function to eliminate the distortion or blurring components caused by the optical system and obtain the filtered spectrum data. The two-dimensional Fourier transform formula is as follows: ; in, Represented as the visual imaging frame in spatial coordinates The pixel value on Represents the visual imaging frame in the frequency domain coordinates The frequency components on Indicates the height of the visual imaging frame, Represents the width of the visual imaging frame, Indicates the specific pixel position in the visual imaging frame, represents the horizontal and vertical frequency indices in the frequency domain, represents a complex sine wave, Represents phase rotation in complex arithmetic.

[0028] The filtered spectral data is inversely transformed to obtain multiple visual imaging frames enhanced in the frequency domain. A weight map is constructed based on the modulation transfer function and the local visual imaging frame gradient, and the multi-focus visual imaging frames are weightedly fused to generate high-resolution visual imaging frames with full-focus effect.

[0029] The edge detection operator is applied to process the high-resolution visual imaging frames with full focus effect, and the gradient of the visual imaging frame is extracted. In the gradient response map, the point with the largest gradient in the area where the texture changes most dramatically in the visual imaging frame is found as the target center position, and the target position sequence is obtained by combining the information of multiple focus layers.

[0030] The background light intensity change of each visual imaging frame is measured by the visual imaging frame sensor, and the sensitivity of subsequent velocity estimation to illumination changes is adjusted through the dynamic attenuation factor. At the same time, the adjacent frames in the target posture sequence are differentially processed and weightedly corrected to obtain the real-time target velocity vector.

[0031] S2: Neuromorphic spatiotemporal motion is used to model the target pose sequence and real-time velocity vector to obtain the predicted trajectory of the dynamically coupled target body and the spatiotemporal constraint domain of the robot end-effector.

[0032] S2.1: Based on the position sequence of the dynamic target and the IMU data (inertial measurement unit data) of the robot body, a dynamic visual sensor is used to capture the target motion event stream in real time, and the six-axis inertial measurement unit is used to synchronously collect the robot body motion state. The original visual events are processed by spatiotemporal filtering, and Gaussian kernel function is used for temporal convolution and spatial smoothing.

[0033] The IMU raw data is nonlinearly compressed using the Sigmoid function to convert the analog signal into a pulse rate encoding in the style of biological neurons. The pulse encoder is then used to achieve time synchronization and unified representation of multimodal data, resulting in a pulse-coded visual event stream and IMU pulse sequence.

[0034] Based on pulse-coded visual event streams and IMU pulse sequences, a multi-layer spiking neural network with input layer, hidden layer and output layer is constructed. The input layer receives pulse-coded visual and inertial data; the hidden layer uses leaky integrate-trigger neurons to simulate the temporal dynamics of biological nerves; the output layer generates target three-dimensional trajectory prediction and robot motion planning instructions, and uses an improved pulse timing-dependent plasticity rule to optimize network weights, jointly learn target dynamics and robot kinematic constraints, and obtain the predicted trajectory of the target body and the expected trajectory of the robot terminal.

[0035] Based on the predicted trajectory of the target body and the expected trajectory of the robot end, a dynamic coupling field is constructed. The Level Set method is used to calculate the safety domain in real time to construct a safe operation boundary. GPU acceleration (visual imaging frame processor parallel computing acceleration technology) is used to build a safe operation boundary. The model predictive control algorithm is used to optimize the motion path within the safety domain, balance the trajectory tracking accuracy and energy efficiency, and finally generate the spatiotemporal constraint domain of the robot end effector and the optimized trajectory of the dynamically coupled target body.

[0036] S3: The predicted trajectory and the spatiotemporal constraint domain are aligned to obtain multispectral data, and the multispectral data is tensor decomposed to obtain a complete three-dimensional model. At the same time, the complete three-dimensional model and the spatiotemporal constraint domain are feature extracted and substituted into the spatiotemporal dominated grasping point optimization function to obtain the optimal grasping posture.

[0037] S3.1: Based on the spatiotemporal constraint domain of the robot end effector and the optimized trajectory of the dynamically coupled target, the data streams from multi-source sensors are aligned through a high-precision clock synchronization protocol, and a depth camera is used to capture the three-dimensional structural information of the target object. The surface pressure distribution and the inertial unit are collected in combination with the tactile array to record the motion trajectory. Subsequently, the spatial calibration algorithm is used to unify the coordinate systems of each sensor to construct a four-dimensional spatiotemporal tensor structure that integrates visual, tactile and inertial data. The tensor dimensions include: visual imaging frame channel, time step, execution pressure, and inertial characteristics. With high-resolution visual imaging frames as the basic channel, the tactile pressure field and acceleration characteristics are superimposed to finally form a spatiotemporally aligned multimodal tensor.

[0038] Based on the spatiotemporal aligned multimodal tensor, Tucker decomposition (high-order singular value decomposition) is used to extract multimodal joint features to achieve data dimensionality reduction, and the spatiotemporal coupling loss function is constructed using time series correlation to model the correlation between different modes in the time dimension to obtain the compressed dynamic coupling feature tensor.

[0039] Combining neural radiation field technology with tactile physics modeling, the compressed dynamic coupling feature tensor is reconstructed into a dynamic implicit surface model of the target object. At the same time, the deformation field is predicted and the material stiffness tensor is estimated to obtain a complete dynamic three-dimensional model and material mechanical parameters.

[0040] Based on the complete dynamic three-dimensional model, the grasping optimization problem is constructed in the three-dimensional special Euclidean group manifold space. The stability index is designed by integrating geometric constraints and material mechanical properties. The optimal grasping posture sequence that satisfies the friction cone condition is solved through an iterative optimization algorithm. The contact force distribution and motion smoothness requirements are balanced in trajectory planning to obtain the optimal grasping posture sequence.

[0041] S4: The six-dimensional force feedback is obtained through the torque sensor in the end effector of the robot arm. Combined with the optimal grasping posture, the motor control amount and the actual posture error are obtained through impedance model calculation and actuator dynamic compensation.

[0042] S4.1: According to the optimal grasping posture of the robot end effector, a virtual mass-spring-damper system model is constructed through the Cartesian space mechanical characteristics, and the target inertia, damping and stiffness parameters are set. The expected end force and posture are converted into the joint space driving torque. The equivalent mass, damping and stiffness parameters are introduced to adjust the control response performance. Then, the kinematic Jacobian matrix is ​​used to map the spatial mechanical characteristics to the joint torque. A dynamic stiffness adjustment mechanism is designed, which can adaptively adjust the spring stiffness according to the end contact state, and finally the joint impedance torque of the robot is obtained.

[0043] According to the complete three-dimensional model of the dynamically coupled target body, the relationship between the idle stroke range and the torque is recorded by driving the joints forward and reverse, and a lookup table is constructed. Then, the viscous-Coulomb friction parameters are fitted through the ramp speed experiment, and the friction model is fitted using the least squares method. The viscous-Coulomb friction parameters include the Coulomb friction parameters and the viscous damping coefficient. The state observer is deployed to estimate the unmodeled dynamic disturbance in real time to obtain the compensation torque of the robot.

[0044] The total driving torque is generated by integrating the robot's compensation torque and joint impedance torque, and the driving current command is inversely calculated based on the electromagnetic characteristics of the motor. Then, the inertia torque feedforward compensation mechanism is introduced to offset the dynamic load caused by the acceleration motion of the robot arm, and the gradient limited algorithm is used to smooth the current command to obtain the motor control quantity of the robot joint drive unit.

[0045] The actual posture data of the end effector is obtained through high-precision measuring equipment, and the trajectory tracking error is calculated in the three-dimensional special Euclidean group space. The error calculation adopts the logarithmic mapping operation to obtain the posture residual, and a deep reinforcement learning framework is constructed to optimize the impedance parameters online. Then, the evaluation network is used to evaluate the stability benefit of the parameter adjustment strategy, and the stiffness and damping corrections are output in real time through the execution network to obtain the actual posture error of the robot end effector.

[0046] S5: Substitute the actual posture error into the genetic algorithm to obtain the optimized phase parameter, and obtain the operating health status of the unhooking robot through the abnormal marking fault recovery protocol in the motor control quantity.

[0047] S5.1: Based on the motor current error and posture error, the dynamic health factor is calculated through dynamic weighting and exponential decay, and three different health levels are divided according to the health factor: when the health factor exceeds 90%, it is judged to be in a healthy state and full power operation is allowed; when it is between 70% and 90%, it enters sub-health mode and automatically limits the load power; below 70%, the emergency shutdown protection mechanism is triggered.

[0048] Based on the dynamic health factor and target imaging resolution, and with the health status as a constraint, a multi-objective optimization model of phase distribution including focus energy concentration, material response uniformity and error convergence speed is constructed. Through a hybrid optimization algorithm, a genetic algorithm is used for global search to generate the initial population, and the quasi-Newton method is used to perform gradient refinement on the high-quality solution, ultimately obtaining the optimized phase parameters and actual spot size.

[0049] The time-domain statistical characteristics of joint torque and vibration spectrum characteristics are extracted, and a five-dimensional health feature vector including mean, variance, kurtosis, flat-band energy distribution and harmonic ratio is constructed. The matrix divergence difference between real-time data and benchmark health status is calculated through sliding window covariance analysis, and the divergence value is mapped into an intuitive health score using the hyperbolic tangent function to accurately quantify the degree to which the system deviates from normal operating conditions. When abnormal harmonic distortion is detected in a specific frequency band, a fault feature list is generated to locate potential damaged components, ultimately obtaining an accurate health score and fault feature list.

[0050] This embodiment also provides a hook-removing robot control system based on deep visual recognition, including: The super-meter sensing module obtains the original light signal and the ambient light intensity, dynamically focuses the original light signal, generates the target pose sequence according to the frequency domain fusion algorithm and the gradient maximization positioning algorithm, and substitutes the ambient light intensity into the pose difference of adjacent frames to obtain the real-time velocity vector; Pre-control domain model module: uses neuromorphic spatiotemporal motion to model the target pose sequence and real-time velocity vector to obtain the predicted trajectory of the dynamically coupled target and the spatiotemporal constraint domain of the robot end effector; Tensor optimal grasping module: the predicted trajectory and the spatiotemporal constraint domain are registered to obtain multispectral data, and the multispectral data is decomposed into tensors to obtain a complete three-dimensional model. At the same time, the complete three-dimensional model and the spatiotemporal constraint domain are extracted and substituted into the spatiotemporal dominated grasping point optimization function to obtain the optimal grasping posture; Impedance control module: obtains six-dimensional force feedback through the torque sensor in the end effector of the robot arm, combines it with the optimal grasping posture, calculates the impedance model and dynamically compensates the actuator to obtain the motor control amount and the actual posture error; Evolutionary phase control module: Substitute the actual posture error into the genetic algorithm to obtain the optimized phase parameters, and obtain the operating health status of the unhooking robot through the abnormal marking fault recovery protocol in the motor control quantity.

[0051] This embodiment also provides a computer device, which is suitable for the case of a hook-removing robot control method based on deep visual recognition, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the hook-removing robot control method based on deep visual recognition as proposed in the above embodiment.

[0052] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0053] The present embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for controlling a hook-removing robot based on deep visual recognition as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage device, a flash memory, a disk or an optical disk.

[0054] In summary, the present invention achieves this by: simulating the pulse coding mechanism of biological neural networks, integrating visual event streams and inertial sensor data, and constructing a motion prediction model with dynamic coupling in space and time. This step uses the principle of bionics to simulate the coordination mechanism of biological motion, jointly describing the target motion trajectory and the dynamic constraints of the robotic arm, and breaking through the limitation of the disconnection between trajectory prediction and the motion ability of the robotic arm in traditional isolated modeling. The interactive characteristics of the target and the robotic arm are captured through dynamic coupling equations to generate spatiotemporal motion prediction results that include safety boundaries. Its core value lies in significantly improving the accuracy of trajectory prediction, while constructing an operating domain that meets the physical limits of the robotic arm, optimizing the efficiency of obstacle avoidance path planning, reducing the risk of collision during motion, and realizing collaborative motion control similar to biological intelligence.

[0055] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A hook-removing robot control method based on deep visual recognition, characterized in that: include, The original light signal and ambient light intensity are obtained. The original light signal is dynamically focused, and the target pose sequence is generated according to the frequency domain fusion algorithm and the gradient maximization positioning algorithm. At the same time, the ambient light intensity is substituted into the pose difference of adjacent frames to obtain the real-time velocity vector. Using neuromorphic spatiotemporal motion, we model the target pose sequence and real-time velocity vector to obtain the predicted trajectory of the dynamically coupled target and the spatiotemporal constraint domain of the robot end-effector; The predicted trajectory and the spatiotemporal constraint domain are registered to obtain multispectral data, and the multispectral data is tensor-decomposed to obtain a complete three-dimensional model. At the same time, the complete three-dimensional model and the spatiotemporal constraint domain are feature extracted and substituted into the spatiotemporal-dominated grasping point optimization function to obtain the optimal grasping posture. The six-dimensional force feedback is obtained through the torque sensor in the end effector of the robot arm. Combined with the optimal grasping posture, the motor control amount and the actual posture error are obtained through impedance model calculation and actuator dynamic compensation. The actual posture error is substituted into the genetic algorithm to obtain the optimized phase parameter, and the operating health status of the unhooking robot is obtained through the abnormal marking fault recovery protocol in the motor control quantity.

2. The method for controlling a hook-removing robot based on deep visual recognition according to claim 1, characterized in that: The original light signal is dynamically focused, and a target pose sequence is generated according to a frequency domain fusion algorithm and a gradient maximization positioning algorithm. At the same time, the ambient light intensity is substituted into the pose difference of adjacent frames to obtain a real-time velocity vector. The steps are as follows: The light wave phase in the original light signal is dynamically focused by real-time phase adjustment, and normalized at the same time to obtain a multi-focus light field sequence after dynamic focusing; Each focal visual imaging frame in the multi-focal light field sequence is subjected to Fourier transformation and optical transfer function filtering, and combined with weighted multi-focal visual imaging frame fusion based on modulation transfer function to obtain a fused high-resolution visual imaging frame; The Sobel operator is used to calculate the gradient amplitude of the fused high-resolution visual imaging frame, and the position with the maximum gradient amplitude is taken as the target center to obtain the single-frame target pose with the gradient direction as the target orientation. The continuous single-frame target poses are time-aligned to obtain the target pose sequence. The dynamic attenuation factor is obtained by dynamically adjusting the weight of the ambient light intensity, and the real-time velocity vector is obtained based on the position difference between the dynamic attenuation factor and the target posture sequence.

3. The method for controlling a hook-removing robot based on deep visual recognition according to claim 2, characterized in that: The dynamic weight adjustment refers to automatically adjusting the proportion of different data sources in the final calculation result according to the real-time changes in the ambient light intensity.

4. The method for controlling a hook-removing robot based on deep visual recognition according to claim 1, characterized in that: The neuromorphic spatiotemporal motion is used to model the target pose sequence and real-time velocity vector to obtain the predicted trajectory of the dynamically coupled target body and the spatiotemporal constraint domain of the robot end effector. The steps are as follows: The target position sequence and real-time velocity vector are spliced ​​and normalized to obtain the preprocessed spatiotemporal feature data, and a neuromorphic spatiotemporal model is constructed through spiking neural network and spatiotemporal convolution. The preprocessed spatiotemporal feature data are substituted into the neuromorphic spatiotemporal model for iterative prediction and Bayesian uncertainty quantification to obtain the predicted trajectory. The quadratic rule is then used to solve the optimal trajectory that satisfies the constraints, and finally the spatiotemporal constraint domain is obtained.

5. The method for controlling a hook-removing robot based on deep visual recognition according to claim 1, characterized in that: The complete three-dimensional model and the spatiotemporal constraint domain are subjected to feature extraction, and the spatiotemporal dominated grasping point optimization function is substituted to obtain the optimal grasping posture. The steps are as follows: By calculating edge curvature, surface thermal gradient and accessibility weight, the geometric features, thermodynamic features and space-time constraint parameters in the complete three-dimensional model and space-time constraint domain are extracted respectively; Based on the extracted geometric features, thermodynamic characteristics and spatiotemporal constraint parameters, the optimal grasping posture is obtained by optimizing geometric stability, thermal safety and spatiotemporal accessibility, and searching in the constraint domain using genetic algorithm and gradient descent.

6. The method for controlling a hook-removing robot based on deep visual recognition according to claim 1, characterized in that: The six-dimensional force feedback is obtained through the torque sensor in the end effector of the manipulator, combined with the optimal grasping posture, and the motor control amount and the actual posture error are obtained through impedance model calculation and actuator dynamic compensation. The steps are as follows: A torque sensor is installed at the end of the robotic arm to collect six-dimensional force feedback in real time. Combined with the optimal grasping posture, the calibrated six-dimensional force feedback and target contact force are obtained by converting signal filtering and coordinate system. The motor control quantity of the six-dimensional force feedback after the target contact force calibration is calculated, and the position error and attitude error are evaluated to obtain the motor control quantity and the actual posture error.

7. The method for controlling a hook-removing robot based on deep visual recognition according to claim 1, characterized in that: The actual posture error is substituted into the genetic algorithm to obtain the optimized phase parameter, and the abnormal marking fault recovery protocol in the motor control quantity is used to obtain the healthy operation status of the hook removal robot. The steps are as follows: The actual posture error is optimized by genetic algorithm according to the fitness function, and the motor control quantity is monitored in real time through multiple sensors to obtain the optimized phase parameters and abnormal marks; According to the fault recovery protocol, the health status of the optimized phase parameters and abnormal marks is evaluated to obtain the system health status of the unhooking robot.

8. A hook-removing robot control system based on deep visual recognition, based on the hook-removing robot control method based on deep visual recognition according to any one of claims 1 to 7, characterized in that: include, The super-meter sensing module obtains the original light signal and the ambient light intensity, dynamically focuses the original light signal, generates the target pose sequence according to the frequency domain fusion algorithm and the gradient maximization positioning algorithm, and substitutes the ambient light intensity into the pose difference of adjacent frames to obtain the real-time velocity vector; The pre-control domain model module uses neuromorphic spatiotemporal motion to model the target pose sequence and real-time velocity vector to obtain the predicted trajectory of the dynamically coupled target body and the spatiotemporal constraint domain of the robot end effector; The tensor optimal grasping module performs data registration between the predicted trajectory and the spatiotemporal constraint domain to obtain multispectral data, and performs tensor decomposition on the multispectral data to obtain a complete three-dimensional model. At the same time, the complete three-dimensional model and the spatiotemporal constraint domain are extracted and substituted into the spatiotemporal dominated grasping point optimization function to obtain the optimal grasping posture. The impedance control module obtains six-dimensional force feedback through the torque sensor in the end effector of the robot arm, combines the optimal grasping posture, calculates the impedance model and dynamically compensates the actuator to obtain the motor control amount and the actual posture error; The evolutionary phase control module substitutes the actual posture error into the genetic algorithm to obtain the optimized phase parameters, and obtains the operating health status of the unhooking robot through the abnormal marking fault recovery protocol in the motor control quantity.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the hook-removing robot control method based on deep visual recognition as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the hook-removing robot control method based on deep visual recognition as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-perception fusion bionic search robot with man-machine bidirectional interaction function

    CN117021112A

  • Brain vision plasticity checking and training system based on multi-focus depth stack

    CN117434726A

  • Robot real-time voice interaction method and system based on spiking neural network

    CN119049460A

  • Feature extraction and encoding of spiking neural networks using convolutional neural network and trainable encoders for deployment in neuromorphic chips

    WO2025062034A1

Cited By

  • Automatic cargo hooking and unhooking method and system based on machine vision technology

    CN120664450A

  • Robot touch and vision fused double-loop coupling self-adaptive grabbing control method

    CN121061892A

  • Robot tactile and visual fusion double ring coupling adaptive grasping control method

    CN121061892B

  • Uncoupling robot control system and method based on multi-source visual fusion

    CN121105033A

  • Linear servo joint reverse driving control method and system

    CN121340307A