Robot alarm processing method and system based on target detection algorithm and cloud platform

By deploying sensor arrays in industrial sites, using object detection algorithms and cloud platforms to process image and depth information, constructing dynamic anomaly correlation diagrams and generating anomaly development prediction sequences, the problems of inefficient exception alarm processing in the existing technology are solved, and high-precision abnormality detection and prediction are achieved, enhancing the autonomy and adaptability of the robot.

CN119418174BActive Publication Date: 2025-05-06北京网藤科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510019961.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The prior art relies on manual intervention in abnormal alarm processing, is inefficient and prone to errors, and is difficult to adapt to complex industrial environments and diverse abnormal situations, and lacks the ability to predict the dynamic evolution of abnormal events and future development trends.

Method used

The robot alarm processing method based on object detection algorithm and cloud platform is adopted. Image and depth information are collected through sensor arrays, adaptive image segmentation and preprocessing are performed, convolutional neural network object detection model is called for abnormal object detection, combined with the feature fusion network guided by depth information to obtain accurate three-dimensional characterization data, build dynamic anomaly correlation diagrams and generate anomaly development prediction sequences, and generate multi-level decision sequences through hierarchical reinforcement learning algorithms and deep deterministic strategy gradient algorithms, optimize task execution strategies and generate optimal motion trajectory.

Benefits of technology

It improves the accuracy and efficiency of abnormal detection, realizes dynamic monitoring and prediction of abnormal events, enhances the autonomy and adaptability of the robot, can avoid obstacles in real time and conducts dynamic skills migration, ensuring the robustness and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418174B_ABST
    Figure CN119418174B_ABST
Patent Text Reader

Abstract

The present invention provides a robot alarm processing method and system based on a target detection algorithm and a cloud platform, which relates to the technical field of robot alarm processing, including: a convolutional neural network target detection model based on a cloud platform, combining a residual structure, a channel attention mechanism and an improved anchor-free frame target detection algorithm to perform abnormal target detection, and perform feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data, use an improved timing graph neural network to construct a dynamic abnormal association graph, and combine an improved hierarchical reinforcement learning algorithm to generate a multi-level decision sequence, and finally perform collaborative task planning through a task decomposition network to obtain a task allocation plan, and the robot performs abnormal processing according to the optimal motion trajectory, obstacle avoidance trajectory data, real-time posture data, key operation sequence data and contact force feedback data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot alarm processing, and in particular to a robot alarm processing method and system based on a target detection algorithm and a cloud platform. Background Art

[0002] Robotics plays an increasingly important role in industrial automation, especially in abnormal alarm processing. Traditional abnormality handling methods usually rely on manual intervention, which is inefficient and error-prone.

[0003] With the development of technology, automated abnormal alarm processing has achieved certain success in specific scenarios, but there are still problems such as the need to manually define rules, difficulty in adapting to complex industrial environments and various abnormal situations, reliance on a large amount of labeled data for training, while abnormal data in industrial sites is usually scarce and difficult to obtain, and lack of ability to predict the dynamic evolution of abnormal events and future development trends, making it difficult to take preventive measures in advance.

[0004] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention

[0005] The embodiments of the present invention provide a robot alarm processing method and system based on a target detection algorithm and a cloud platform, which can at least solve some of the problems existing in the prior art.

[0006] A first aspect of an embodiment of the present invention provides a robot alarm processing method based on a target detection algorithm and a cloud platform, comprising:

[0007] Based on a preset sensor array, image information and depth information corresponding to the industrial site are collected and uploaded to the cloud platform. The image information and depth information are divided by an adaptive image segmentation algorithm to obtain a multidimensional feature image block. The multidimensional feature image block is preprocessed based on a wavelet transform operation and a Laplace edge enhancement to obtain a standard image block. The cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, and determines the multi-scale deep features and key area features corresponding to the standard image block in combination with a residual structure and a channel attention mechanism. Abnormal target detection is performed in combination with an improved anchor-free target detection algorithm to obtain an abnormal detection result and perform feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data;

[0008] Based on the precise three-dimensional characterization data and the anomaly detection results, a dynamic anomaly association graph is constructed through an improved time-series graph neural network. By adding time coding to the graph attention layer to model long-range temporal dependencies, the anomaly evolution law is obtained and an anomaly development prediction sequence is generated. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized through an improved deep deterministic policy gradient algorithm combined with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through a bidirectional attention mechanism and a graph convolution layer to obtain a task allocation plan. The optimal motion trajectory is generated through adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm.

[0009] The optimal motion trajectory is received, an improved dynamic path planning algorithm is executed, and real-time obstacle avoidance navigation is performed in combination with a topological map and a sampling tree to obtain obstacle avoidance trajectory data of the robot. The depth information and the image features collected in real time are integrated through an improved visual servo algorithm for precise positioning to obtain real-time posture data. Key action features are extracted through an improved imitation learning algorithm and a pre-built action value assessment model. Dynamic skill transfer is performed in combination with an online adaptation module to obtain key operation sequence data and add it to an impedance adaptive control algorithm. The impedance parameters are adaptively adjusted through a neural network compensator to obtain contact force feedback data. State monitoring is performed in combination with an improved abnormality prediction algorithm to obtain abnormality handling results.

[0010] In an optional embodiment,

[0011] Based on the preset sensor array, the image information and depth information corresponding to the industrial site are collected and uploaded to the cloud platform. The image information and the depth information are divided by an adaptive image segmentation algorithm to obtain a multidimensional feature image block. The multidimensional feature image block is preprocessed based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block. The cloud platform calls the convolutional neural network target detection model in the pre-trained target detection model library based on the image features of the standard image block, and determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism. The improved anchor-free target detection algorithm is combined to perform abnormal target detection, obtain abnormal detection results, and perform feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data including:

[0012] The image information and depth information corresponding to the industrial site are collected through a pre-set sensor array and uploaded to the cloud platform, wherein the image information is stored in three channels of RGB and the value range of each pixel is 0-255, and the depth information records the distance value corresponding to each pixel;

[0013] For the image information and the depth information, the contrast and spatial position information of each pixel are calculated by an adaptive image segmentation algorithm to construct a saliency map, color quantization is performed based on the image information and the depth information is normalized, the saliency map is used as a probability weight in combination with the normalized depth gradient distribution to dynamically adjust the seed point distribution density, the seed point density is increased in an area where the saliency is higher than a preset saliency threshold, and the seed point density is reduced in an area where the saliency is lower than the saliency threshold, the color similarity, texture features and the depth information are integrated to construct an affinity relationship between pixels, and the cluster center position and boundary are iteratively optimized until convergence to obtain the multidimensional feature image block;

[0014] Performing a wavelet transform operation on the multidimensional feature image block, extracting frequency domain features by multi-scale decomposition, performing contrast enhancement on low-frequency components by piecewise linear mapping, and dynamically adjusting the enhancement coefficient according to the grayscale distribution of the multidimensional feature image block, performing adaptive threshold denoising on high-frequency components, wherein the threshold parameter is determined according to the local variance of the multidimensional feature image block, performing Laplace edge enhancement, and adaptively adjusting the sub-kernel coefficient according to the gradient amplitude of the multidimensional feature image block, uniformly scaling the enhanced image block to a standard size, and performing pixel value normalization processing to obtain a standard image block;

[0015] The cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, introduces a channel attention mechanism in the residual structure of the convolutional neural network target detection model, and acquires statistical information of the channel dimension through global feature pooling to learn the importance weights of different channel features, wherein the main branch of the residual structure performs conventional convolution operations, the attention branch calculates channel weights, and the short-circuit branch maintains feature identity mapping, and the enhanced feature map is obtained by weighted fusion of the outputs of the three branches, and the multi-scale deep features and the key area features are extracted based on the enhanced feature map;

[0016] Combined with the improved anchor-free target detection algorithm, the multi-scale deep features and the key area features are upsampled and aligned, and the feature information of different scales is fused by weighted summation. The center point position and size offset of the target are predicted through the convolution layer. The center point prediction branch outputs a heat map to represent the probability distribution of the target's existence, and the size prediction branch directly regresses the width and height ratio of the target. The prediction results are subjected to maximum suppression to remove overlapping detection boxes and bounding box adjustments to obtain abnormal detection results.

[0017] The anomaly detection result and the depth information are integrated by a dual-stream structure through a feature fusion network guided by depth information, wherein the two feature branches of the dual-stream structure respectively include multi-layer convolution modules and each convolution module consists of a convolution layer, a normalization layer and an activation function. The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through a decoder network to obtain accurate three-dimensional characterization data using spatial position coordinates, posture angles and geometric dimensions of the target.

[0018] In an optional embodiment,

[0019] The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through the decoder network to obtain accurate three-dimensional representation data including:

[0020] A first feature map corresponding to the anomaly detection result and a second feature map corresponding to the depth information are respectively extracted through a feature fusion network with a dual-stream structure, wherein the first feature map contains 512 feature channels and the second feature map contains 256 feature channels;

[0021] Performing global average pooling and maximum pooling on the first feature map to extract a first channel-level feature descriptor, connecting the first channel-level feature descriptors in series through a fully connected layer to obtain a first channel attention weight, performing global average pooling and maximum pooling on the second feature map to extract a second channel-level feature descriptor, connecting the second channel-level feature descriptors in series through a fully connected layer to obtain a second channel attention weight;

[0022] Expanding the first feature map and the second feature map into a first feature matrix and a second feature matrix according to channels, normalizing the feature vectors at each position in the first feature matrix and the second feature matrix, and calculating the inner product of the first feature matrix and the second feature matrix to obtain a correlation matrix;

[0023] Using a pre-trained classifier to perform classification prediction on the first feature map to obtain a first prediction score, using the variance of the first prediction score as a first confidence indicator, and calculating a second confidence indicator based on the continuity and local consistency of the depth value in the second feature map;

[0024] Calculating adaptive fusion weights in combination with the correlation matrix, the first confidence index, and the second confidence index, setting equal weights for regions where the correlation is higher than 0.5 and where both the first confidence index and the second confidence index are higher than 0.9, and biasing the weights toward the higher of the first confidence index and the second confidence index for regions where the correlation is lower than 0.5, and introducing context information to assist feature fusion for regions where both the first confidence index and the second confidence index are lower than 0.9;

[0025] Weighting the first feature map and the second feature map according to the adaptive fusion weight to obtain a fusion feature, upsampling the fusion feature through a multi-layer decoder network and fusing jump connection features of different scales in each decoding layer;

[0026] The depth map is back-projected into three-dimensional space to obtain a point cloud, the fusion feature is used to predict the position of the target center point in the point cloud to obtain the spatial position coordinates, the rotation quaternion of the target is predicted to obtain the attitude angle, the three-dimensional bounding box of the target is predicted to obtain the geometric size, and the spatial position coordinates, the attitude angle and the geometric size are combined to form accurate three-dimensional representation data.

[0027] In an optional embodiment,

[0028] Based on the precise three-dimensional representation data and the anomaly detection results, a dynamic anomaly association graph is constructed through an improved time-series graph neural network. By adding time coding to the graph attention layer to model long-range temporal dependencies, the anomaly evolution law is obtained and an anomaly development prediction sequence is generated. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan. The optimal motion trajectory is generated through adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm, including:

[0029] The target three-dimensional position coordinates, posture quaternion, geometric dimensions and anomaly detection confidence in the precise three-dimensional characterization data and the anomaly detection results are combined into a node feature vector, the Euclidean distance between the node feature vectors is calculated to obtain a spatial distance term, the cosine similarity of the node feature vectors is calculated to obtain a feature similarity term, the time series correlation term is calculated based on the motion consistency of the target at adjacent moments, and the weighted sum of the spatial distance term, the feature similarity term and the time series correlation term is used as the edge weight to construct a graph structure;

[0030] The node feature vector is concatenated with the time encoding vector to obtain a fused feature, the time encoding vector is obtained by mapping the timestamp through a periodic sine function, the fused feature is input into the multi-head graph attention layer, each attention head independently calculates the dot product of the query vector and the key vector to obtain an attention score, the attention score is normalized and multiplied with the value vector to obtain a weighted feature, and the weighted features of multiple attention heads are concatenated and linearly transformed to obtain an output feature;

[0031] Based on the output features, the target state change is predicted, including the position offset, the posture change, the size change and the abnormal degree change value, and the predicted state is used as a new input to continue predicting the state at the next moment. The teacher forcing mechanism is used to regularly use the real value to replace the predicted value in the prediction process to suppress the error accumulation and obtain the abnormal development prediction sequence;

[0032] The abnormal development prediction sequence is input into a hierarchical reinforcement learning network, the top-level policy network assigns abnormal targets to designated positions to obtain a sub-target sequence, the middle-level policy network plans the execution sequence based on the sub-target sequence to obtain an execution sequence, and the bottom-level policy network generates an action sequence based on the execution sequence, wherein the reward function of each policy network is designed based on the sub-target completion, execution efficiency and trajectory safety respectively;

[0033] The action sequence is subjected to adversarial training and experience replay optimization, wherein the generator adds random perturbations to the trajectory points to generate adversarial samples, the discriminator evaluates the difference between the true trajectory and the adversarial samples, and selects the sample with the largest temporal difference error for parameter update to obtain an optimized action sequence;

[0034] The optimized action sequence is input into a task decomposition network, spatial features are extracted through multi-layer graph convolution and encoded into task representations, a decoder calculates a task relevance score based on the task representation through a bidirectional attention mechanism, an executor number, a target position, a start time and an expected duration are determined according to the task relevance score and the execution sequence, and a task allocation plan is generated;

[0035] The sampling density is divided according to the obstacle distribution in the environment, and the kinematic constraints and obstacle avoidance constraints are evaluated for each sampling point in the task allocation scheme. The sampling point positions are iteratively optimized by gradient descent to satisfy the constraints and minimize the objective function. The optimized trajectory is smoothed by cubic spline interpolation to obtain the final motion trajectory.

[0036] In an optional embodiment,

[0037] Based on the output features, the target state change is predicted, including the position offset, posture change, size change and abnormal degree change value. The predicted state is used as a new input to continue predicting the state at the next moment. The teacher forcing mechanism is used to regularly use the real value to replace the predicted value during the prediction process to suppress the error accumulation to obtain the abnormal development prediction sequence including:

[0038] The target's three-dimensional position coordinates, attitude quaternion values, geometric dimensions, and anomaly detection confidence values ​​are used as initial state data to construct a state sequence cache with a prediction time step length.

[0039] The output features are respectively input into four independent fully connected layer prediction branches, the first prediction branch outputs the position offset of the target in three directions, the second prediction branch outputs the change in attitude angle represented by quaternion, the third prediction branch outputs the size scaling of the target in three dimensions of length, width and height, and the fourth prediction branch outputs the change in the confidence of abnormality detection;

[0040] The target state is updated using a progressive prediction mechanism. The predicted position offset is added to the current position coordinate value to obtain the position coordinate value at the next moment. The predicted quaternion change is multiplied by the current attitude quaternion value to obtain the attitude angle value at the next moment. The predicted size scaling is multiplied by the current geometric size value to obtain the target size value at the next moment. The predicted abnormality change value is added to the current confidence value to obtain the abnormality degree value at the next moment.

[0041] The prediction sequence is divided into multiple prediction intervals of equal length. At the beginning of each prediction interval, the real observation data is used to replace the prediction value as the input state. The real observation data includes the actual observed target position coordinates, attitude quaternion, geometric dimensions and anomaly detection confidence. At the middle moment of the prediction interval, the prediction state of the previous moment is used as input to continue predicting the state of the next moment.

[0042] Physical constraint checks are performed on the predicted state values. The position offset between adjacent time steps does not exceed the maximum motion speed allowed by the physical system. The rotation angle change between adjacent time steps does not exceed the maximum rotation speed of the system. The size scaling is limited to the preset range of the original size. The abnormality value is kept between zero and one. The predicted values ​​that do not meet the constraints are corrected by linear interpolation.

[0043] The state values ​​of each time step in the prediction sequence are sorted in timestamp order to generate an abnormal development prediction sequence containing occurrence time, three-dimensional spatial position, posture angle, geometric size and abnormality degree value.

[0044] In an optional embodiment,

[0045] The optimal motion trajectory is received, an improved dynamic path planning algorithm is executed and real-time obstacle avoidance navigation is performed in combination with a topological map and a sampling tree to obtain robot obstacle avoidance trajectory data, and the robot is accurately positioned by fusing depth information and real-time collected image features through an improved visual servo algorithm to obtain real-time posture data, and key action features are extracted through an improved imitation learning algorithm and a pre-built action value evaluation model, and dynamic skill transfer is performed in combination with an online adaptation module to obtain key operation sequence data and add it to an impedance adaptive control algorithm, and contact force feedback data is obtained by adaptively adjusting impedance parameters through a neural network compensator, and state monitoring is performed in combination with an improved abnormality prediction algorithm, and abnormality processing results are obtained, including:

[0046] Receiving optimal motion trajectory data including position coordinates, attitude angle, motion speed and acceleration information, wherein the optimal motion trajectory data includes multiple path point information;

[0047] Divide the workspace into multiple connected areas and construct an environmental topology map, record the connection relationship and travel cost between the connected areas, generate an initial path based on the environmental topology map, perform local expansion through a sampling tree on the basis of the initial path, dynamically adjust the sampling density according to the complexity of the environment, perform collision detection on the sampling points and calculate the shortest distance to the obstacles, and perform local path replanning when a collision risk is detected to generate robot obstacle avoidance trajectory data;

[0048] Denoising and hole filling processing are performed on the depth image to obtain effective depth information, feature points are extracted from the image sequence and feature descriptors are established, feature matching is performed through a multi-level screening strategy, and mismatched points are eliminated using geometric consistency constraints. The effective depth information and feature points are weightedly fused according to the measurement confidence, and real-time pose data is obtained through pose optimization;

[0049] Based on the rate of change of state variables, the key time points of the teaching data are identified to obtain action segmentation points, and feature vectors containing position, posture and speed information are extracted from the segmented action segments. The importance of the feature vectors is scored through the action value evaluation model to obtain key action features and dynamically adjust action parameters according to the characteristics of the target object and environmental constraints to generate key operation sequence data;

[0050] A neural network compensator is used to learn the dynamic characteristics of the environment in real time, and position tracking and force control are achieved through a multi-layer nested control structure. The controller stiffness matrix and damping parameters are dynamically adjusted according to the force sensor feedback, and the contact force feedback data is recorded;

[0051] Time domain features are extracted from the system state sequence through a sliding window. The time domain features include the statistical characteristics and change trends of the state variables. The time domain features are input into the prediction model to estimate the state evolution. When the prediction result shows an abnormality, the protection strategy is executed and the abnormality handling result is output.

[0052] In an optional embodiment,

[0053] Perform collision detection on the sampling points and calculate the shortest distance to obstacles. When a collision risk is detected, perform local path replanning to generate robot obstacle avoidance trajectory data including:

[0054] The workspace is divided into basic grid units and an octree structure is constructed, and the obstacle area where obstacles exist is finely segmented by recursive partitioning, and the idle state, fully occupied state and partially occupied state are recorded in the leaf nodes of the octree structure, and the state information of the leaf nodes is updated in real time based on sensor data;

[0055] The robot body is simplified into a multi-level sphere combination model consisting of a large-sized sphere corresponding to the trunk part and a small-sized sphere corresponding to the robotic arm part, and the adjacent nodes are searched in the octree structure based on the center position of the sphere of the multi-level sphere combination model to obtain the collision risk space area;

[0056] The obstacle in the collision risk space region is estimated by using a spherical bounding box to obtain a preliminary distance. If the preliminary distance is less than a preset distance threshold, the obstacle surface in the collision risk space region is converted into point cloud data, and the closest point pair is determined by iterative search to obtain the actual shortest distance.

[0057] The position, velocity and acceleration information of the obstacle are acquired by multi-sensor data fusion. The motion state of the obstacle is estimated by Kalman filter and the future position of the obstacle is predicted. If the future position of the obstacle intersects with the planned path, the obstacle avoidance planning is triggered.

[0058] A sampling tree is established with the current position of the robot as the root node, and an adaptive sampling strategy related to the obstacle distribution is adopted to expand the sampling points to the surrounding space, and a cost evaluation is performed on the sampling points, wherein the cost evaluation includes the target distance, obstacle distance, path length and steering smoothness;

[0059] Transition points are inserted between original path points to smooth the path. The positions of the transition points are determined by an optimization algorithm. The optimization algorithm takes minimizing the change in path curvature as the optimization goal, generates a complete motion trajectory including position coordinates, posture angles, linear velocity and angular velocity, detects environmental contact force and monitors position tracking errors through torque sensors, and replans the motion trajectory when the environmental contact force exceeds a safe range or the position tracking error exceeds a limit.

[0060] A second aspect of an embodiment of the present invention provides a robot alarm processing system based on a target detection algorithm and a cloud platform, comprising:

[0061] The first unit is used to collect image information and depth information corresponding to the industrial site based on a preset sensor array and upload them to the cloud platform, divide the image information and the depth information through an adaptive image segmentation algorithm to obtain a multidimensional feature image block, pre-process the multidimensional feature image block based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block, and the cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism, performs abnormal target detection in combination with an improved anchor-free target detection algorithm, obtains an abnormal detection result, and performs feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data;

[0062] The second unit is used to construct a dynamic anomaly association graph through an improved time-series graph neural network based on the precise three-dimensional characterization data and the anomaly detection results, obtain the anomaly evolution law and generate an anomaly development prediction sequence by adding time coding to the graph attention layer to model the long-range time series dependency, and generate an anomaly development prediction sequence. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan, and the optimal motion trajectory is generated through the adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm;

[0063] The third unit is used to receive the optimal motion trajectory, execute the improved dynamic path planning algorithm and perform real-time obstacle avoidance navigation in combination with the topological map and the sampling tree to obtain the robot's obstacle avoidance trajectory data, accurately locate the robot by fusing the depth information and the real-time collected image features through the improved visual servo algorithm to obtain real-time posture data, extract key action features through the improved imitation learning algorithm and the pre-built action value evaluation model, perform dynamic skill transfer in combination with the online adaptation module, obtain key operation sequence data and add it to the impedance adaptive control algorithm, obtain contact force feedback data by adaptively adjusting the impedance parameters through the neural network compensator, perform state monitoring in combination with the improved abnormality prediction algorithm, and obtain abnormality handling results.

[0064] According to a third aspect of the embodiments of the present invention,

[0065] An electronic device is provided, comprising:

[0066] processor;

[0067] a memory for storing processor-executable instructions;

[0068] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0069] A fourth aspect of the embodiments of the present invention is:

[0070] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0071] In the present invention, by combining multi-dimensional feature image blocks, depth information, an improved anchor-free target detection algorithm and a feature fusion network guided by depth information, it is possible to more accurately detect abnormal targets and obtain their three-dimensional characterization data, thereby improving the accuracy and efficiency of abnormality detection. Through a deep deterministic policy gradient algorithm and a task decomposition network, collaborative task planning and optimized task execution strategy are achieved, thereby generating a better robot alarm processing strategy. Through an improved dynamic path planning algorithm, a visual servoing algorithm and an imitation learning algorithm, the robot can perform real-time obstacle avoidance navigation, precise positioning, and dynamic skill transfer. Combined with impedance adaptive control and anomaly prediction algorithms, real-time monitoring of the robot state and anomaly processing can be achieved, thereby enhancing the robot's autonomy and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a flowchart of a robot alarm processing method based on a target detection algorithm and a cloud platform according to an embodiment of the present invention;

[0073] Figure 2It is a structural schematic diagram of a robot alarm processing system based on a target detection algorithm and a cloud platform according to an embodiment of the present invention. DETAILED DESCRIPTION

[0074] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0075] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0076] Figure 1 FIG. 1 is a flow chart of a robot alarm processing method based on a target detection algorithm and a cloud platform according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0077] S1. Based on the preset sensor array, the image information and depth information corresponding to the industrial site are collected and uploaded to the cloud platform. The image information and the depth information are divided by an adaptive image segmentation algorithm to obtain a multidimensional feature image block. The multidimensional feature image block is preprocessed based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block. The cloud platform calls the convolutional neural network target detection model in the pre-trained target detection model library based on the image features of the standard image block, and determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism. The improved anchor-free target detection algorithm is used to perform abnormal target detection, and the abnormal detection results are obtained. Feature fusion is performed through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data;

[0078] The adaptive image segmentation algorithm is an algorithm that can automatically adjust the segmentation strategy according to the characteristics of the input image. By adaptively selecting a suitable segmentation method, it can be dynamically adjusted according to the local characteristics of different images. It is often used in the field of image processing. The Laplace edge enhancement is a technology that uses the Laplace operator to enhance the edge of the image. The convolutional neural network target detection model is a deep learning model designed based on a convolutional neural network, which aims to detect the position and category of the target object from the input image. The anchor-free target detection algorithm is a target detection algorithm that does not rely on a predefined anchor frame. It avoids the selection of the anchor frame by directly predicting the bounding box position of the target from the image, and has good flexibility and adaptability. The feature fusion network guided by depth information is a deep learning network that combines depth information for feature fusion.

[0079] In an optional embodiment,

[0080] Based on the preset sensor array, the image information and depth information corresponding to the industrial site are collected and uploaded to the cloud platform. The image information and the depth information are divided by an adaptive image segmentation algorithm to obtain a multidimensional feature image block. The multidimensional feature image block is preprocessed based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block. The cloud platform calls the convolutional neural network target detection model in the pre-trained target detection model library based on the image features of the standard image block, and determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism. The improved anchor-free target detection algorithm is combined to perform abnormal target detection, obtain abnormal detection results, and perform feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data including:

[0081] The image information and depth information corresponding to the industrial site are collected through a pre-set sensor array and uploaded to the cloud platform, wherein the image information is stored in three channels of RGB and the value range of each pixel is 0-255, and the depth information records the distance value corresponding to each pixel;

[0082] For the image information and the depth information, the contrast and spatial position information of each pixel are calculated by an adaptive image segmentation algorithm to construct a saliency map, color quantization is performed based on the image information and the depth information is normalized, the saliency map is used as a probability weight in combination with the normalized depth gradient distribution to dynamically adjust the seed point distribution density, the seed point density is increased in an area where the saliency is higher than a preset saliency threshold, and the seed point density is reduced in an area where the saliency is lower than the saliency threshold, the color similarity, texture features and the depth information are integrated to construct an affinity relationship between pixels, and the cluster center position and boundary are iteratively optimized until convergence to obtain the multidimensional feature image block;

[0083] Performing a wavelet transform operation on the multidimensional feature image block, extracting frequency domain features by multi-scale decomposition, performing contrast enhancement on low-frequency components by piecewise linear mapping, and dynamically adjusting the enhancement coefficient according to the grayscale distribution of the multidimensional feature image block, performing adaptive threshold denoising on high-frequency components, wherein the threshold parameter is determined according to the local variance of the multidimensional feature image block, performing Laplace edge enhancement, and adaptively adjusting the sub-kernel coefficient according to the gradient amplitude of the multidimensional feature image block, uniformly scaling the enhanced image block to a standard size, and performing pixel value normalization processing to obtain a standard image block;

[0084] The cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, introduces a channel attention mechanism in the residual structure of the convolutional neural network target detection model, and acquires statistical information of the channel dimension through global feature pooling to learn the importance weights of different channel features, wherein the main branch of the residual structure performs conventional convolution operations, the attention branch calculates channel weights, and the short-circuit branch maintains feature identity mapping, and the enhanced feature map is obtained by weighted fusion of the outputs of the three branches, and the multi-scale deep features and the key area features are extracted based on the enhanced feature map;

[0085] Combined with the improved anchor-free target detection algorithm, the multi-scale deep features and the key area features are upsampled and aligned, and the feature information of different scales is fused by weighted summation. The center point position and size offset of the target are predicted through the convolution layer. The center point prediction branch outputs a heat map to represent the probability distribution of the target's existence, and the size prediction branch directly regresses the width and height ratio of the target. The prediction results are subjected to maximum suppression to remove overlapping detection boxes and bounding box adjustments to obtain abnormal detection results.

[0086] The anomaly detection result and the depth information are integrated by a dual-stream structure through a feature fusion network guided by depth information, wherein the two feature branches of the dual-stream structure respectively include multi-layer convolution modules and each convolution module consists of a convolution layer, a normalization layer and an activation function. The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through a decoder network to obtain accurate three-dimensional characterization data using spatial position coordinates, posture angles and geometric dimensions of the target.

[0087] The saliency map is a map used to represent the most significant areas in an image. By calculating the saliency of each area of ​​the image, the saliency map can highlight the most noticeable parts in the human eye and is widely used in tasks such as image segmentation, target detection and image enhancement. The seed point refers to the key point used as an initialization reference in image segmentation, clustering or image processing tasks. The seed point density refers to the distribution density of seed points in an image or data set. The sub-kernel coefficient refers to an adjustment parameter of the kernel function in kernel methods such as support vector machines. By adjusting the sub-kernel coefficient, the shape and calculation range of the kernel function can be controlled, thereby affecting the generalization ability and classification effect of the classifier. The dual-stream structure is a network architecture that contains two branch streams processed in parallel, and each stream processes different information or data.

[0088] A sensor array is pre-set at the industrial site to collect image information and depth information at the site. Image information is stored in three channels of RGB, and the value range of each pixel is 0 to 255. Depth information records the distance value from each pixel to the sensor. The collected image information and depth information are then uploaded to the cloud platform for processing.

[0089] The uploaded image information and depth information are adaptively segmented on the cloud platform, and a saliency map is constructed by calculating the contrast and spatial position information of each pixel. The image information is color quantized, for example, the color space is divided into several discrete color categories using the K-Means clustering algorithm. At the same time, the depth information is normalized, for example, the depth value is scaled to between 0 and 1. The saliency map is used as a probability weight, and the seed point distribution density is dynamically adjusted in combination with the normalized depth gradient distribution. In areas where the saliency is higher than the preset threshold, such as areas where the saliency value is greater than 0.8, the seed point density is increased; in areas where the saliency is lower than the threshold, such as areas where the saliency value is less than 0.2, the seed point density is reduced. The affinity relationship between pixels is constructed by fusing color similarity, texture features (such as local binary patterns LBP) and depth information. The cluster center position and boundaries are iteratively optimized until convergence, and multiple multidimensional feature image blocks are finally obtained. For example, assuming that an image block contains a red object and a blue background, the segmentation algorithm will divide the red object and the blue background into different image blocks.

[0090] The multi-dimensional feature image blocks obtained by segmentation are preprocessed, and the frequency domain features of the image blocks are extracted by wavelet transform operation. The low-frequency components are contrast enhanced by piecewise linear mapping, and the enhancement coefficient is dynamically adjusted according to the grayscale distribution of the image block. For example, the enhancement coefficient is larger in the area with lower grayscale value, and smaller in the area with higher grayscale value. The high-frequency components are adaptively thresholded for denoising, and the threshold parameter is determined according to the local variance of the image block. For example, the threshold is higher in the area with larger variance, and lower in the area with smaller variance. Laplace edge enhancement is performed, and the sub-kernel coefficient is adaptively adjusted according to the gradient amplitude of the image block. For example, the sub-kernel coefficient is larger in the area with larger gradient amplitude, and smaller in the area with smaller gradient amplitude. Finally, the enhanced image blocks are uniformly scaled to a standard size, such as 224x224 pixels, and the pixel values ​​are normalized, such as scaling the pixel values ​​to between 0 and 1, to obtain standard image blocks.

[0091] The cloud platform calls the pre-trained convolutional neural network target detection model in the target detection model library based on the image features of the standard image block. The channel attention mechanism is introduced in the residual structure of the convolutional neural network target detection model. The statistical information of the channel dimension is obtained through global feature pooling, such as the average and maximum value of each channel, and the importance weights of different channel features are learned. The main branch of the residual structure performs conventional convolution operations, the attention branch calculates the channel weights, and the short-circuit branch maintains the feature identity mapping. The enhanced feature map is obtained by weighted fusion of the outputs of the three branches. Multi-scale deep features and key area features are extracted based on the enhanced feature map.

[0092] Combined with the improved anchor-free object detection algorithm, abnormal object detection is performed. Multi-scale deep features and key area features are upsampled and aligned, and feature information of different scales is fused by weighted summation. The center point position and size offset of the target are predicted through the convolution layer. The center point prediction branch outputs a heat map to represent the probability distribution of the target's existence, and the size prediction branch directly regresses the width and height ratio of the target. The prediction results are subjected to maximum suppression to remove overlapping detection boxes and bounding box adjustments, and finally the abnormal detection results are obtained.

[0093] Accurate three-dimensional representation data is obtained by feature fusion based on a feature fusion network guided by depth information. The feature fusion network guided by depth information uses a two-stream structure to integrate anomaly detection results and depth information. The two feature branches of the two-stream structure contain multi-layer convolution modules, each of which consists of a convolution layer, a normalization layer (such as BatchNorm), and an activation function (such as ReLU). The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through the decoder network to reconstruct the spatial position coordinates, posture angles, and geometric dimensions of the target to obtain accurate three-dimensional representation data. For example, the three-dimensional coordinates (x, y, z) of the center point of the target, the rotation angles (α, β, γ), and the length, width, and height (l, w, h) can be obtained.

[0094] In this embodiment, image information and depth information are integrated, which can more accurately detect abnormal targets, reduce false detection rate and missed detection rate, and reconstruct the three-dimensional information of the target, including spatial position, posture and size, to provide more comprehensive target information. The adaptive image segmentation and preprocessing steps can adapt to different industrial field environments and improve the robustness and generalization ability of the method.

[0095] In an optional embodiment,

[0096] The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through the decoder network to obtain accurate three-dimensional representation data including:

[0097] A first feature map corresponding to the anomaly detection result and a second feature map corresponding to the depth information are respectively extracted through a feature fusion network with a dual-stream structure, wherein the first feature map contains 512 feature channels and the second feature map contains 256 feature channels;

[0098] Performing global average pooling and maximum pooling on the first feature map to extract a first channel-level feature descriptor, connecting the first channel-level feature descriptors in series through a fully connected layer to obtain a first channel attention weight, performing global average pooling and maximum pooling on the second feature map to extract a second channel-level feature descriptor, connecting the second channel-level feature descriptors in series through a fully connected layer to obtain a second channel attention weight;

[0099] Expanding the first feature map and the second feature map into a first feature matrix and a second feature matrix according to channels, normalizing the feature vectors at each position in the first feature matrix and the second feature matrix, and calculating the inner product of the first feature matrix and the second feature matrix to obtain a correlation matrix;

[0100] Using a pre-trained classifier to perform classification prediction on the first feature map to obtain a first prediction score, using the variance of the first prediction score as a first confidence indicator, and calculating a second confidence indicator based on the continuity and local consistency of the depth value in the second feature map;

[0101] Calculating adaptive fusion weights in combination with the correlation matrix, the first confidence index, and the second confidence index, setting equal weights for regions where the correlation is higher than 0.5 and where both the first confidence index and the second confidence index are higher than 0.9, and biasing the weights toward the higher of the first confidence index and the second confidence index for regions where the correlation is lower than 0.5, and introducing context information to assist feature fusion for regions where both the first confidence index and the second confidence index are lower than 0.9;

[0102] Weighting the first feature map and the second feature map according to the adaptive fusion weight to obtain a fusion feature, upsampling the fusion feature through a multi-layer decoder network and fusing jump connection features of different scales in each decoding layer;

[0103] The depth map is back-projected into three-dimensional space to obtain a point cloud, the fusion feature is used to predict the position of the target center point in the point cloud to obtain the spatial position coordinates, the rotation quaternion of the target is predicted to obtain the attitude angle, the three-dimensional bounding box of the target is predicted to obtain the geometric size, and the spatial position coordinates, the attitude angle and the geometric size are combined to form accurate three-dimensional representation data.

[0104] The back projection is an image reconstruction technology, which is usually used in medical imaging, image enhancement and other fields. By reverse mapping the projection data back to the original image space, the details of the image can be reconstructed. It is often used in CT scanning, MRI imaging and other technologies. The rotation quaternion is a mathematical tool for representing the rotation of an object in three-dimensional space. It can effectively represent the rotation in three-dimensional space and avoid the gimbal lock problem of traditional Euler angles. Therefore, it is widely used in computer graphics, robotics, virtual reality and other fields.

[0105] A dual-stream feature fusion network is used to extract feature maps corresponding to anomaly detection results and feature maps corresponding to depth information. Assume that the anomaly detection branch outputs a feature map containing 512 feature channels, and the depth information branch outputs a feature map containing 256 feature channels.

[0106] The two feature maps are processed with the channel-level attention mechanism respectively. The anomaly detection feature map is subjected to global average pooling and global maximum pooling, and the pooled results are concatenated and sent to the fully connected layer to obtain the channel attention weights of the anomaly detection branch. Similarly, the depth information feature map is also subjected to global average pooling and global maximum pooling, and the pooled results are concatenated and sent to the fully connected layer to obtain the channel attention weights of the depth information branch. For example, after the feature map of 512 channels is subjected to global average pooling and global maximum pooling, two 1x1x512 vectors are obtained, which are concatenated to obtain a 1x1x1024 vector, and then passed through the fully connected layer to obtain 512 channel attention weights.

[0107] Calculate the correlation between the two feature maps. Expand the anomaly detection feature map and the depth information feature map into feature matrices by channel, and normalize the feature vectors at each position. For example, convert a feature map of size HxWx512 into a feature matrix of size 512xHxW, and normalize each row. Then, calculate the inner product of the two feature matrices to get the correlation matrix. This matrix reflects the feature correlation of the two feature maps at different positions.

[0108] Calculate the confidence index of the two feature maps. Use the pre-trained classifier to classify the anomaly detection feature map and use the variance of the prediction score as the confidence index of the anomaly detection branch. For example, input the anomaly detection feature map into the classifier to get the probability of each category, and then calculate the variance of these probabilities. The confidence index of the depth information branch is calculated based on the continuity and local consistency of the depth values. For example, the continuity of the depth information can be measured by calculating the gradient amplitude of the depth map.

[0109] Based on the correlation matrix and the two confidence indicators, the adaptive fusion weights are calculated. For areas where the correlation is higher than 0.5 and both confidence indicators are higher than 0.9, equal fusion weights are set, such as 0.5 and 0.5. For areas where the correlation is lower than 0.5, the weights are biased towards the higher confidence indicator. For example, if the confidence of anomaly detection is 0.8 and the confidence of depth information is 0.6, the fusion weights can be set to 0.7 and 0.3. For areas where both confidence indicators are lower than 0.9, contextual information is introduced to assist feature fusion. For example, the features of neighborhood pixels can be used to improve confidence.

[0110] The two feature maps are weighted and fused according to the adaptive fusion weights. For example, the anomaly detection feature map and the depth information feature map are multiplied by the corresponding weights and then added to obtain the fused feature.

[0111] The fused features are fed into a multi-layer decoder network for upsampling, and the skip connection features of different scales are fused in each decoding layer. The depth map is back-projected into a three-dimensional space to obtain a point cloud. The fused features are used to predict the position of the target center point in the point cloud to obtain the spatial position coordinates, the target's rotation quaternion is predicted to obtain the attitude angle, and the target's three-dimensional bounding box is predicted to obtain the geometric size. The spatial position coordinates, attitude angle, and geometric size are combined into accurate three-dimensional representation data.

[0112] In this embodiment, by fusing anomaly detection results and depth information and adopting an adaptive weight strategy, information from different sources can be more effectively utilized, thereby improving the accuracy of three-dimensional target representation. The introduction of confidence indicators and contextual information enhances the model's ability to handle noise and uncertainty and improves the robustness of the model. The application of channel-level attention mechanism can reduce the amount of calculation and improve the computational efficiency of the model.

[0113] S2. Based on the precise three-dimensional characterization data and the anomaly detection results, a dynamic anomaly association graph is constructed through an improved time-series graph neural network. By adding time coding to the graph attention layer to model long-range temporal dependencies, the anomaly evolution law is obtained and an anomaly development prediction sequence is generated. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan. The optimal motion trajectory is generated through adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm.

[0114] The time-series graph neural network is a neural network model that combines time-series data and graph structure, and is designed to process graph data with time-series dependencies. The dynamic anomaly association graph is a method that captures abnormal relationships and patterns in data by building a graph model. By dynamically updating the properties of nodes and edges in the graph, the dynamic anomaly association graph can identify and track abnormal patterns that change over time. It is often used in application scenarios such as network security, financial fraud detection, and abnormal behavior analysis. The improved hierarchical reinforcement learning algorithm is an algorithm that is improved on the basis of traditional reinforcement learning, and adopts a hierarchical strategy to handle complex decision-making tasks. The hierarchical option framework is a framework for reinforcement learning, in which complex decision-making tasks are decomposed into multiple levels of subtasks, each subtask has its own goals and strategies. The improved deep deterministic policy gradient algorithm is a reinforcement learning algorithm that is a combination of deep learning and deterministic policy gradient algorithms. It uses deep neural networks to approximate policies and value functions, and uses deterministic policy gradients to optimize the decision-making process.

[0115] In an optional embodiment,

[0116] Based on the precise three-dimensional representation data and the anomaly detection results, a dynamic anomaly association graph is constructed through an improved time-series graph neural network. By adding time coding to the graph attention layer to model long-range temporal dependencies, the anomaly evolution law is obtained and an anomaly development prediction sequence is generated. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan. The optimal motion trajectory is generated through adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm, including:

[0117] The target three-dimensional position coordinates, posture quaternion, geometric dimensions and anomaly detection confidence in the precise three-dimensional characterization data and the anomaly detection results are combined into a node feature vector, the Euclidean distance between the node feature vectors is calculated to obtain a spatial distance term, the cosine similarity of the node feature vectors is calculated to obtain a feature similarity term, the time series correlation term is calculated based on the motion consistency of the target at adjacent moments, and the weighted sum of the spatial distance term, the feature similarity term and the time series correlation term is used as the edge weight to construct a graph structure;

[0118] The node feature vector is concatenated with the time encoding vector to obtain a fused feature, the time encoding vector is obtained by mapping the timestamp through a periodic sine function, the fused feature is input into the multi-head graph attention layer, each attention head independently calculates the dot product of the query vector and the key vector to obtain an attention score, the attention score is normalized and multiplied with the value vector to obtain a weighted feature, and the weighted features of multiple attention heads are concatenated and linearly transformed to obtain an output feature;

[0119] Based on the output features, the target state change is predicted, including the position offset, the posture change, the size change and the abnormal degree change value, and the predicted state is used as a new input to continue predicting the state at the next moment. The teacher forcing mechanism is used to regularly use the real value to replace the predicted value in the prediction process to suppress the error accumulation and obtain the abnormal development prediction sequence;

[0120] The abnormal development prediction sequence is input into a hierarchical reinforcement learning network, the top-level policy network assigns abnormal targets to designated positions to obtain a sub-target sequence, the middle-level policy network plans the execution sequence based on the sub-target sequence to obtain an execution sequence, and the bottom-level policy network generates an action sequence based on the execution sequence, wherein the reward function of each policy network is designed based on the sub-target completion, execution efficiency and trajectory safety respectively;

[0121] The action sequence is subjected to adversarial training and experience replay optimization, wherein the generator adds random perturbations to the trajectory points to generate adversarial samples, the discriminator evaluates the difference between the true trajectory and the adversarial samples, and selects the sample with the largest temporal difference error for parameter update to obtain an optimized action sequence;

[0122] The optimized action sequence is input into a task decomposition network, spatial features are extracted through multi-layer graph convolution and encoded into task representations, a decoder calculates a task relevance score based on the task representation through a bidirectional attention mechanism, an executor number, a target position, a start time and an expected duration are determined according to the task relevance score and the execution sequence, and a task allocation plan is generated;

[0123] The sampling density is divided according to the obstacle distribution in the environment, and the kinematic constraints and obstacle avoidance constraints are evaluated for each sampling point in the task allocation scheme. The sampling point positions are iteratively optimized by gradient descent to satisfy the constraints and minimize the objective function. The optimized trajectory is smoothed by cubic spline interpolation to obtain the final motion trajectory.

[0124] The attitude quaternion is a mathematical representation used to describe the rotation of an object in three-dimensional space. The kinematic constraint refers to some physical rules or restrictions that must be followed when defining the movement of an object or system in the fields of robotics and manipulator control. The degree of satisfaction of the obstacle avoidance constraint refers to the degree to which the control strategy follows the obstacle avoidance rules in systems such as robots or autonomous driving.

[0125] Construct a dynamic anomaly association graph. The three-dimensional position coordinates, attitude quaternion, geometric dimensions and anomaly detection confidence of the target are combined into a vector as the node feature of the graph, and the correlation between nodes in the three dimensions of space, feature and time is calculated. Spatial correlation is obtained by calculating the Euclidean distance between node feature vectors; feature correlation is obtained by calculating the cosine similarity of node feature vectors; and temporal correlation is calculated based on the motion consistency of targets at adjacent moments. The three correlation indicators are weighted and summed to obtain the edge weight of the graph, thereby constructing a dynamic anomaly association graph. For example, suppose there are two targets A and B, whose three-dimensional coordinates are (1, 2, 3) and (2, 3, 4), respectively, and whose anomaly confidences are 0.8 and 0.9, respectively, and whose motion directions are consistent at the previous moment. Then their spatial distance, feature similarity and temporal correlation can be calculated, and these three values ​​are weighted and summed to obtain the edge weight to connect nodes A and B.

[0126] Predict the evolution law of anomalies, concatenate the node feature vector with the time encoding vector to obtain fused features. The time encoding vector maps the timestamp to a high-dimensional space through a periodic sine function. The fused features are input into the multi-head graph attention layer, each attention head calculates the attention score independently, and the results of multiple attention heads are merged to obtain the output features. Based on the output features, the changes in the target state are predicted, including position offset, posture change, size change, and abnormality degree change. The predicted state is used as a new input to continue predicting the state at the next moment to form an abnormal development prediction sequence. To avoid error accumulation, a teacher forcing mechanism is used to periodically replace the predicted value with the true value. For example, based on the position of target A at the current moment (1, 2, 3) and the predicted position offset (0.1, 0.2, 0.3), the predicted position of A at the next moment (1.1, 2.2, 3.3) can be obtained.

[0127] Generate multi-level decision sequences. Input the abnormal development prediction sequence into the hierarchical reinforcement learning network. The network contains top, middle and bottom policy networks. The top policy network assigns the target to the specified location according to the predicted development trend of the abnormal target to form a sub-target sequence. The middle policy network plans the execution order according to the sub-target sequence to form an execution sequence. The bottom policy network generates a specific action sequence based on the execution sequence. The reward function of each policy network is designed based on the sub-target completion, execution efficiency and trajectory safety. For example, the top policy network can assign abnormal target A to a safe area, the middle policy network can plan to process A first and then other targets, and the bottom policy network can generate action sequences such as moving and grasping.

[0128] Optimize the task execution strategy. Perform adversarial training and experience replay optimization on the generated action sequence. The generator adds random perturbations to the trajectory points to generate adversarial samples. The discriminator evaluates the difference between the real trajectory and the adversarial sample, and selects the sample with the largest difference for parameter update, thereby optimizing the action sequence. For example, the generator can add some random noise to the original trajectory to generate a new trajectory. The discriminator determines whether the new trajectory is better and updates the network parameters based on the judgment result.

[0129] Perform collaborative task planning and input the optimized action sequence into the task decomposition network. The network extracts spatial features through graph convolution and encodes them into task representations. The decoder calculates the task relevance score based on the task representation through a bidirectional attention mechanism. According to the task relevance score and execution sequence, the executor number, target location, start time and expected duration are determined, and a task allocation scheme is generated. For example, according to the relevance scores of task A and task B, and their execution order, task A can be assigned to executor 1 and task B to executor 2, and their start time and duration are determined.

[0130] Generate the optimal motion trajectory. Divide the sampling density according to the obstacle distribution in the environment, and evaluate the degree of satisfaction of the kinematic constraints and obstacle avoidance constraints for each sampling point in the task allocation plan. Iterate and optimize the sampling point positions through gradient descent to satisfy the constraints and minimize the objective function. Smooth the optimized trajectory with cubic spline interpolation to obtain the final motion trajectory. For example, according to the position and shape of the obstacle, the density of the sampling points can be adjusted, and the position of the sampling points can be adjusted through the optimization algorithm to avoid the obstacle and satisfy the kinematic constraints.

[0131] In this embodiment, by constructing a dynamic anomaly association graph and using an improved time series graph neural network, the anomaly development trend can be predicted more accurately, providing a reliable basis for subsequent decision-making. The hierarchical reinforcement learning algorithm is combined with the hierarchical option framework to decompose complex tasks into multiple sub-goals, simplifying the decision-making process and improving decision-making efficiency. The introduction of adversarial training and experience replay mechanism enhances the generalization ability of the action value evaluation model, making the task execution strategy more robust and better able to adapt to complex and changing environments.

[0132] In an optional embodiment,

[0133] Based on the output features, the target state change is predicted, including the position offset, posture change, size change and abnormal degree change value. The predicted state is used as a new input to continue predicting the state at the next moment. The teacher forcing mechanism is used to regularly use the real value to replace the predicted value during the prediction process to suppress the error accumulation to obtain the abnormal development prediction sequence including:

[0134] The target's three-dimensional position coordinates, attitude quaternion values, geometric dimensions, and anomaly detection confidence values ​​are used as initial state data to construct a state sequence cache with a prediction time step length.

[0135] The output features are respectively input into four independent fully connected layer prediction branches, the first prediction branch outputs the position offset of the target in three directions, the second prediction branch outputs the change in attitude angle represented by quaternion, the third prediction branch outputs the size scaling of the target in three dimensions of length, width and height, and the fourth prediction branch outputs the change in the confidence of abnormality detection;

[0136] The target state is updated using a progressive prediction mechanism. The predicted position offset is added to the current position coordinate value to obtain the position coordinate value at the next moment. The predicted quaternion change is multiplied by the current attitude quaternion value to obtain the attitude angle value at the next moment. The predicted size scaling is multiplied by the current geometric size value to obtain the target size value at the next moment. The predicted abnormality change value is added to the current confidence value to obtain the abnormality degree value at the next moment.

[0137] The prediction sequence is divided into multiple prediction intervals of equal length. At the beginning of each prediction interval, the real observation data is used to replace the prediction value as the input state. The real observation data includes the actual observed target position coordinates, attitude quaternion, geometric dimensions and anomaly detection confidence. At the middle moment of the prediction interval, the prediction state of the previous moment is used as input to continue predicting the state of the next moment.

[0138] Physical constraint checks are performed on the predicted state values. The position offset between adjacent time steps does not exceed the maximum motion speed allowed by the physical system. The rotation angle change between adjacent time steps does not exceed the maximum rotation speed of the system. The size scaling is limited to the preset range of the original size. The abnormality value is kept between zero and one. The predicted values ​​that do not meet the constraints are corrected by linear interpolation.

[0139] The state values ​​of each time step in the prediction sequence are sorted in timestamp order to generate an abnormal development prediction sequence containing occurrence time, three-dimensional spatial position, posture angle, geometric size and abnormality degree value.

[0140] Initialize the state sequence cache to obtain the initial state information of the target, including 3D position coordinates (for example, the initial position is x=10, y=20, z=30), attitude quaternion (for example, the initial attitude is w=1, x=0, y=0, z=0, indicating no rotation), geometric dimensions (for example, the initial length, width and height are 5, 10, 15) and anomaly detection confidence (for example, the initial confidence is 0.1). Create a state sequence cache based on the length of the prediction time step (for example, predict 100 time steps) to store the target state information for each time step.

[0141] Perform state prediction and update. Input the target state data of the current time step into four independent fully connected layer prediction branches respectively. Assume that the first prediction branch outputs the position offset of the target in three directions as (Δx=0.1, Δy=0.2, Δz=0.3). Add this offset to the current position coordinate to obtain the position coordinate at the next moment (x=10.1, y=20.2, z=30.3). Assume that the second prediction branch outputs the change in attitude angle, expressed as a quaternion (Δw, Δx, Δy, Δz). Perform quaternion multiplication on the quaternion change and the current attitude quaternion to obtain the attitude quaternion at the next moment. Assume that the third prediction branch outputs the size scaling of the target length, width and height in three dimensions as (1.01, 1.02, 1.03). Multiply this scaling by the current geometric size to obtain the target size at the next moment (5.05, 10.2, 15.45). Assume that the change in the confidence of abnormal detection output by the fourth prediction branch is 0.01. Add the change to the current confidence value to obtain the abnormality value of 0.11 at the next moment.

[0142] Apply the teacher forcing mechanism to divide the prediction sequence into multiple prediction intervals of equal length (for example, each interval is 10 in length). At the start of each prediction interval (for example, the 0th, 10th, and 20th time steps), use the actual observed target position coordinates, attitude quaternion, geometric size, and anomaly detection confidence to replace the predicted value as the input state. For example, at the 10th time step, obtain the real observation data, such as the position coordinates are (11, 22, 33), the attitude quaternion is (0.9, 0.1, 0.2, 0.3), the size is (5.5, 11, 16.5), and the confidence is 0.2. Use these real observation data as the state value of the 10th time step and use them in subsequent predictions.

[0143] Physical constraints are checked on the predicted state values. For example, suppose the maximum velocity allowed by the physical system is 1, the maximum rotation velocity is 0.5 radians, the size scaling range is 0.9 to 1.1 times the original size, and the abnormality range is 0 to 1. If the position offset between adjacent time steps exceeds the maximum velocity value, it is corrected by linear interpolation. For example, if the predicted offset is 1.5, it is corrected to 1. Similarly, constraint checks and corrections are performed on rotation angle changes, size scaling, and abnormality values.

[0144] Generate an abnormal development prediction sequence, sort the state values ​​of each time step in the prediction sequence in timestamp order, and generate an abnormal development prediction sequence containing occurrence time, three-dimensional spatial position, posture angle, geometric size and abnormality degree value. For example, generate a state sequence containing 100 time steps, each time step contains timestamp, position, posture, size and abnormality degree value.

[0145] In this embodiment, through the teacher enforcement mechanism, the real value is used regularly to correct the predicted value, which effectively suppresses the error accumulation and improves the prediction accuracy. Through the physical constraint check, it ensures that the prediction result conforms to the physical laws, avoids unreasonable prediction values, and makes the prediction result more practical. It comprehensively considers multiple state dimensions such as the position, posture, size and degree of abnormality of the target, and can more comprehensively predict the abnormal development trend of the target, providing richer information for the early warning and handling of abnormal events.

[0146] S3. Receive the optimal motion trajectory, execute the improved dynamic path planning algorithm and perform real-time obstacle avoidance navigation in combination with the topological map and the sampling tree to obtain the robot's obstacle avoidance trajectory data, perform precise positioning by fusing the depth information and the real-time collected image features through the improved visual servo algorithm, and obtain real-time posture data, extract key action features through the improved imitation learning algorithm and the pre-built action value evaluation model, combine the online adaptation module to perform dynamic skill transfer, obtain key operation sequence data and add it to the impedance adaptive control algorithm, obtain contact force feedback data by adaptively adjusting the impedance parameters through the neural network compensator, and perform state monitoring in combination with the improved abnormality prediction algorithm to obtain abnormality handling results.

[0147] The improved dynamic path planning algorithm is an algorithm for dynamically adjusting the path planning strategy. By acquiring environmental change information in real time, such as the appearance or position change of obstacles, the path is adjusted in time to avoid collision or optimize driving time. The real-time obstacle avoidance navigation is a navigation technology that adjusts the path in real time according to environmental changes. By using sensors to collect surrounding environment data and combining with the real-time path planning algorithm, obstacles can be avoided and the navigation path can be corrected in real time to ensure the safety of robots or autonomous vehicles. The improved visual servo algorithm is an algorithm for controlling motion based on visual feedback. The target features are extracted through technologies such as image processing and deep learning, and the motion trajectory of the object is adjusted according to the gap between the features and the desired target. The imitation learning algorithm is a machine learning method for learning tasks by imitating expert behavior. The intelligent agent is trained by observing the actions of experts performing tasks in the environment and imitating their decision-making process. The neural network compensator is a device that uses a neural network for compensation control. By learning the nonlinear characteristics and disturbances in the system and feeding this information back to the control system, the system is compensated to reduce errors or improve the response speed of the system. The contact force feedback data is a feedback signal from a tactile sensor or a force sensor, indicating the contact force between the object and the environment.

[0148] In an optional embodiment,

[0149] The optimal motion trajectory is received, an improved dynamic path planning algorithm is executed and real-time obstacle avoidance navigation is performed in combination with a topological map and a sampling tree to obtain robot obstacle avoidance trajectory data, and the robot is accurately positioned by fusing depth information and real-time collected image features through an improved visual servo algorithm to obtain real-time posture data, and key action features are extracted through an improved imitation learning algorithm and a pre-built action value evaluation model, and dynamic skill transfer is performed in combination with an online adaptation module to obtain key operation sequence data and add it to an impedance adaptive control algorithm, and contact force feedback data is obtained by adaptively adjusting impedance parameters through a neural network compensator, and state monitoring is performed in combination with an improved abnormality prediction algorithm, and abnormality processing results are obtained, including:

[0150] Receiving optimal motion trajectory data including position coordinates, attitude angle, motion speed and acceleration information, wherein the optimal motion trajectory data includes multiple path point information;

[0151] Divide the workspace into multiple connected areas and construct an environmental topology map, record the connection relationship and travel cost between the connected areas, generate an initial path based on the environmental topology map, perform local expansion through a sampling tree on the basis of the initial path, dynamically adjust the sampling density according to the complexity of the environment, perform collision detection on the sampling points and calculate the shortest distance to the obstacles, and perform local path replanning when a collision risk is detected to generate robot obstacle avoidance trajectory data;

[0152] Denoising and hole filling processing are performed on the depth image to obtain effective depth information, feature points are extracted from the image sequence and feature descriptors are established, feature matching is performed through a multi-level screening strategy, and mismatched points are eliminated using geometric consistency constraints. The effective depth information and feature points are weightedly fused according to the measurement confidence, and real-time pose data is obtained through pose optimization;

[0153] Based on the rate of change of state variables, the key time points of the teaching data are identified to obtain action segmentation points, and feature vectors containing position, posture and speed information are extracted from the segmented action segments. The importance of the feature vectors is scored through the action value evaluation model to obtain key action features and dynamically adjust action parameters according to the characteristics of the target object and environmental constraints to generate key operation sequence data;

[0154] A neural network compensator is used to learn the dynamic characteristics of the environment in real time, and position tracking and force control are achieved through a multi-layer nested control structure. The controller stiffness matrix and damping parameters are dynamically adjusted according to the force sensor feedback, and the contact force feedback data is recorded;

[0155] Time domain features are extracted from the system state sequence through a sliding window. The time domain features include the statistical characteristics and change trends of the state variables. The time domain features are input into the prediction model to estimate the state evolution. When the prediction result shows an abnormality, the protection strategy is executed and the abnormality handling result is output.

[0156] The environment topology map is a graphical model that represents the relationship between objects and obstacles in space. In the environment topology map, nodes usually represent key locations in the environment, while edges represent the connection relationship between these locations. The local path replanning refers to the process of making local adjustments to the existing path in the navigation system when a robot or autonomous driving vehicle encounters a new obstacle or target change. The geometric consistency constraint is a constraint used in path planning or image processing, which requires that the generated path or image be consistent with a given geometric feature.

[0157] Receive optimal motion trajectory data containing position coordinates, attitude angles, motion speed and acceleration information. For example, a trajectory contains 10 path points, each path point contains three-dimensional coordinates (x, y, z), attitude is expressed by Euler angles (roll, pitch, yaw), and speed and acceleration are three-dimensional vectors.

[0158] Divide the workspace into multiple connected areas and construct an environmental topology map. For example, divide a 10mx10m space into 1mx1m grids, each grid representing a connected area. Record the connection relationship and travel cost between connected areas, for example, the travel cost of two adjacent grids is 1, and the travel cost of non-adjacent grids is infinite. Generate an initial path based on the environmental topology map. For example, use the A* algorithm to search for an initial path from the starting point to the target point. Based on the initial path, perform local expansion using the rapidly expanding random tree (RRT) algorithm. Dynamically adjust the sampling density according to the complexity of the environment. For example, increase the sampling density in areas with dense obstacles and reduce the sampling density in open areas. Perform collision detection on the sampling points and calculate the shortest distance to the obstacles. For example, use a spherical bounding box for collision detection. At the same time, perform local path replanning when a collision risk is detected. For example, use the RRT-Connect algorithm to replan a path that avoids obstacles. Finally, generate robot obstacle avoidance trajectory data, for example, trajectory data containing a series of path point coordinates, velocity, and acceleration information.

[0159] Denoise and hole filling are performed on the depth image to obtain effective depth information. For example, bilateral filtering is used for noise reduction, and a hole filling algorithm is used to fill holes in the depth image. Feature points are extracted from the image sequence, for example, using the ORB algorithm to extract feature points, and feature descriptors are established. Feature matching is performed through a multi-level screening strategy. For example, the fast approximate nearest neighbor search (FLANN) is first used for coarse matching, and then the random sampling consensus (RANSAC) algorithm is used for fine matching. Geometric consistency constraints are used to eliminate mismatched points. The effective depth information and feature points are weighted and fused according to the measurement confidence. For example, areas with high confidence in depth information are given a greater weight. Real-time pose data, such as six-degree-of-freedom pose data containing the robot's position and posture, is obtained through a pose optimization algorithm, for example, using the iterative closest point (ICP) algorithm.

[0160] Identify key time points of the teaching data based on the rate of change of the state variables to obtain action segmentation points. For example, when the rate of change of speed exceeds a certain threshold, it is considered to be an action segmentation point. Extract feature vectors containing position, posture and speed information from the segmented action segments. For example, the position and speed information of the starting point, middle point and end point of each action segment are combined into a feature vector. Score the importance of the feature vectors through the action value evaluation model. For example, use a pre-trained neural network model to score the feature vectors. Dynamically adjust the action parameters according to the characteristics of the target object and environmental constraints. For example, adjust the posture and strength of the grasping action according to the shape and size of the target object. Generate key operation sequence data, for example, data containing a series of robot action instructions.

[0161] Use a neural network compensator to learn the dynamic characteristics of the environment in real time. For example, use a radial basis function (RBF) neural network to learn the friction and damping forces in the environment. Implement position tracking and force control through a multi-layer nested control structure. For example, the outer loop controls the position and the inner loop controls the force. Dynamically adjust the controller stiffness matrix and damping parameters based on force sensor feedback. For example, when the contact force is too large, reduce the stiffness and damping parameters. Record contact force feedback data, for example, data containing three-dimensional force vectors and torque vectors.

[0162] Extract time domain features from the system state sequence through a sliding window. For example, use a time window of length 10 to calculate statistical characteristics such as the mean, variance, and rate of change of the state variables in the window. Input the time domain features into the prediction model to estimate the state evolution. For example, use a long short-term memory (LSTM) network to predict the future system state. When the prediction result shows an abnormality, for example, the predicted contact force exceeds the safety threshold, execute a protection strategy, for example, stop the robot movement or issue an alarm. Output the abnormality handling result, for example, including information about the abnormality type and handling measures.

[0163] In this embodiment, by combining the topological map and the sampling tree for path planning, the local optimal solution can be effectively avoided, and the path can be dynamically adjusted according to environmental changes, thereby improving the safety and smoothness of the robot's navigation. Through the improved imitation learning algorithm and the online adaptation module, the robot can dynamically adjust the action parameters according to the characteristics of the target object and the environmental constraints, thereby enhancing the flexibility and adaptability of the robot's operation. Through the improved visual servo algorithm and the impedance adaptive control algorithm, the robot can accurately perceive the environmental information and perform fine operations. At the same time, the abnormality prediction algorithm can timely detect and handle system abnormalities, thereby improving the robustness and reliability of the robot system.

[0164] In an optional embodiment,

[0165] Perform collision detection on the sampling points and calculate the shortest distance to obstacles. When a collision risk is detected, perform local path replanning to generate robot obstacle avoidance trajectory data including:

[0166] The workspace is divided into basic grid units and an octree structure is constructed, and the obstacle area where obstacles exist is finely segmented by recursive partitioning, and the idle state, fully occupied state and partially occupied state are recorded in the leaf nodes of the octree structure, and the state information of the leaf nodes is updated in real time based on sensor data;

[0167] The robot body is simplified into a multi-level sphere combination model consisting of a large-sized sphere corresponding to the trunk part and a small-sized sphere corresponding to the robotic arm part, and the adjacent nodes are searched in the octree structure based on the center position of the sphere of the multi-level sphere combination model to obtain the collision risk space area;

[0168] The obstacle in the collision risk space region is estimated by using a spherical bounding box to obtain a preliminary distance. If the preliminary distance is less than a preset distance threshold, the obstacle surface in the collision risk space region is converted into point cloud data, and the closest point pair is determined by iterative search to obtain the actual shortest distance.

[0169] The position, velocity and acceleration information of the obstacle are acquired by multi-sensor data fusion. The motion state of the obstacle is estimated by Kalman filter and the future position of the obstacle is predicted. If the future position of the obstacle intersects with the planned path, the obstacle avoidance planning is triggered.

[0170] A sampling tree is established with the current position of the robot as the root node, and an adaptive sampling strategy related to the obstacle distribution is adopted to expand the sampling points to the surrounding space, and a cost evaluation is performed on the sampling points, wherein the cost evaluation includes the target distance, obstacle distance, path length and steering smoothness;

[0171] Transition points are inserted between original path points to smooth the path. The positions of the transition points are determined by an optimization algorithm. The optimization algorithm takes minimizing the change in path curvature as the optimization goal, generates a complete motion trajectory including position coordinates, posture angles, linear velocity and angular velocity, detects environmental contact force and monitors position tracking errors through torque sensors, and replans the motion trajectory when the environmental contact force exceeds a safe range or the position tracking error exceeds a limit.

[0172] The fine-grained segmentation is a technique in image processing used to divide an object or region in an image into smaller, finer parts. The transition point refers to a key point in path planning or image segmentation that represents the transition from one region to another. The path smoothing is a technique used in path planning that aims to smooth the planned path to make the path smoother and more natural.

[0173] Build an environmental model. Divide the robot workspace into a series of uniform basic grid units and build an octree structure on this basis. For areas containing obstacles, recursively subdivide them into smaller child nodes until the preset minimum size is reached. Each leaf node records three states: idle, fully occupied, and partially occupied, indicating that the area is completely obstacle-free, completely occupied by obstacles, and partially occupied by obstacles. Using real-time environmental data obtained by robot sensors (such as lidar, depth cameras, etc.), dynamically update the status information of the octree leaf nodes to ensure that the environmental model is consistent with the actual environment. For example, assuming that the workspace is a 10mx10mx10m cube and the basic grid unit size is 0.1m, there are initially 1 million leaf nodes.

[0174] Build a robot model. Simplify the robot body into a multi-level sphere combination model composed of multiple spheres. The robot trunk is represented by a large sphere, and the robotic arms and other parts are represented by multiple small spheres. Each sphere has its center position and radius. For example, a mobile robot can be simplified into a trunk sphere with a radius of 0.3m and two robotic arm spheres with a radius of 0.1m.

[0175] Perform collision detection. According to the center position of each sphere in the multi-level sphere combination model, search for its neighboring nodes in the octree structure. Mark these neighboring nodes as collision risk space areas. For obstacles in the collision risk space area, first use the sphere bounding box to make a preliminary distance estimate. If the preliminary distance is less than the preset distance threshold (for example, 0.5m), the obstacle surface is represented as point cloud data, and the actual shortest distance between the robot and the obstacle and the closest point pair are determined through an iterative search algorithm (for example, k-dtree search).

[0176] Predict obstacle motion. Use multi-sensor data fusion technology (such as fusing LiDAR data with visual sensor data) to obtain the position, velocity, and acceleration information of the obstacle. Use the Kalman filter to estimate and predict the motion state of the obstacle to obtain the position of the obstacle in the future. If the predicted obstacle position intersects with the planned robot path, the obstacle avoidance planning is triggered.

[0177] Perform local path replanning. Establish a sampling tree with the robot's current position as the root node. Use an adaptive sampling strategy related to the obstacle distribution to expand the sampling points in the surrounding space. For example, increase the density of sampling points in areas with high obstacle density. Evaluate the cost of each sampling point, and the cost function includes factors such as target distance, obstacle distance, path length, and steering smoothness. For example, a penalty term can be set based on the obstacle distance, and the closer the distance, the greater the penalty.

[0178] The transition points are inserted between the original path points, and the positions of the transition points are determined using an optimization algorithm (such as the gradient descent method), with the optimization goal of minimizing the change in path curvature, to generate a complete motion trajectory containing position coordinates, attitude angles, linear velocity, and angular velocity. For example, three transition points can be inserted between adjacent path points.

[0179] Track and monitor the trajectory. The robot moves according to the generated trajectory and uses the torque sensor to detect the environmental contact force while monitoring the position tracking error. If the environmental contact force exceeds the safety range (e.g. 10N) or the position tracking error exceeds the limit (e.g. 0.1m), the motion trajectory is replanned.

[0180] In this embodiment, an octree structure and a multi-level sphere model are used to quickly identify potential collision areas, avoiding complex distance calculations for all obstacles, thereby improving the efficiency of collision detection. An adaptive sampling strategy and a cost evaluation based on multiple factors can find a better obstacle avoidance path in a complex environment, and a local path replanning mechanism can effectively avoid dynamic obstacles. By inserting transition points and a smoothing optimization algorithm, the generated motion trajectory is smoother, reducing the acceleration, deceleration and turning actions of the robot, and improving the stability and safety of the motion. At the same time, through the monitoring of torque sensors and position tracking errors, abnormal situations can be discovered and handled in a timely manner to ensure the safe operation of the robot.

[0181] Figure 2 FIG. 1 is a schematic diagram of the structure of a robot alarm processing system based on a target detection algorithm and a cloud platform according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0182] The first unit is used to collect image information and depth information corresponding to the industrial site based on a preset sensor array and upload them to the cloud platform, divide the image information and the depth information through an adaptive image segmentation algorithm to obtain a multidimensional feature image block, pre-process the multidimensional feature image block based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block, and the cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism, performs abnormal target detection in combination with an improved anchor-free target detection algorithm, obtains an abnormal detection result, and performs feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data;

[0183] The second unit is used to construct a dynamic anomaly association graph through an improved time-series graph neural network based on the precise three-dimensional characterization data and the anomaly detection results, obtain the anomaly evolution law and generate an anomaly development prediction sequence by adding time coding to the graph attention layer to model the long-range time series dependency, and generate an anomaly development prediction sequence. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan, and the optimal motion trajectory is generated through the adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm;

[0184] The third unit is used to receive the optimal motion trajectory, execute the improved dynamic path planning algorithm and perform real-time obstacle avoidance navigation in combination with the topological map and the sampling tree to obtain the robot's obstacle avoidance trajectory data, accurately locate the robot by fusing the depth information and the real-time collected image features through the improved visual servo algorithm to obtain real-time posture data, extract key action features through the improved imitation learning algorithm and the pre-built action value evaluation model, perform dynamic skill transfer in combination with the online adaptation module, obtain key operation sequence data and add it to the impedance adaptive control algorithm, obtain contact force feedback data by adaptively adjusting the impedance parameters through the neural network compensator, perform state monitoring in combination with the improved abnormality prediction algorithm, and obtain abnormality handling results.

[0185] According to a third aspect of the embodiments of the present invention,

[0186] An electronic device is provided, comprising:

[0187] processor;

[0188] a memory for storing processor-executable instructions;

[0189] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0190] According to a fourth aspect of the embodiments of the present invention,

[0191] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0192] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A robot alarm processing method based on target detection algorithm and cloud platform, characterized in that: include: Based on a preset sensor array, image information and depth information corresponding to the industrial site are collected and uploaded to the cloud platform. The image information and depth information are divided by an adaptive image segmentation algorithm to obtain a multidimensional feature image block. The multidimensional feature image block is preprocessed based on a wavelet transform operation and a Laplace edge enhancement to obtain a standard image block. The cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, and determines the multi-scale deep features and key area features corresponding to the standard image block in combination with a residual structure and a channel attention mechanism. Abnormal target detection is performed in combination with an improved anchor-free target detection algorithm to obtain an abnormal detection result and perform feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data; Based on the precise three-dimensional representation data and the anomaly detection results, a dynamic anomaly association graph is constructed through an improved time-series graph neural network. By adding time encoding to the graph attention layer to model long-range temporal dependencies, the anomaly evolution law is obtained and an anomaly development prediction sequence is generated. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan. The optimal motion trajectory is generated through adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm, wherein the anomaly development prediction sequence includes the occurrence time, three-dimensional spatial position, posture angle, geometric size and anomaly degree value; Receive the optimal motion trajectory, execute the improved dynamic path planning algorithm and combine the topological map with the sampling tree to perform real-time obstacle avoidance navigation, obtain the robot's obstacle avoidance trajectory data, use the improved visual servo algorithm to fuse the depth information and the real-time collected image features for precise positioning, and obtain real-time posture data, extract key action features through the improved imitation learning algorithm and the pre-built action value evaluation model, combine the online adaptation module to perform dynamic skill transfer, obtain key operation sequence data and add it to the impedance adaptive control algorithm, adaptively adjust the impedance parameters through the neural network compensator to obtain contact force feedback data, combine the improved abnormality prediction algorithm to perform state monitoring, and obtain abnormality handling results, wherein the real-time posture data includes six-degree-of-freedom posture data of the robot's position and posture, and the contact force feedback data includes three-dimensional force vector and torque vector data.

2. The method according to claim 1, characterized in that: Based on the preset sensor array, the image information and depth information corresponding to the industrial site are collected and uploaded to the cloud platform. The image information and the depth information are divided by an adaptive image segmentation algorithm to obtain a multidimensional feature image block. The multidimensional feature image block is preprocessed based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block. The cloud platform calls the convolutional neural network target detection model in the pre-trained target detection model library based on the image features of the standard image block, and determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism. The improved anchor-free target detection algorithm is combined to perform abnormal target detection, obtain abnormal detection results, and perform feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data including: The image information and depth information corresponding to the industrial site are collected through a pre-set sensor array and uploaded to the cloud platform, wherein the image information is stored in three channels of RGB and the value range of each pixel is 0-255, and the depth information records the distance value corresponding to each pixel; For the image information and the depth information, the contrast and spatial position information of each pixel are calculated by an adaptive image segmentation algorithm to construct a saliency map, color quantization is performed based on the image information and the depth information is normalized, the saliency map is used as a probability weight in combination with the normalized depth gradient distribution to dynamically adjust the seed point distribution density, the seed point density is increased in an area where the saliency is higher than a preset saliency threshold, and the seed point density is reduced in an area where the saliency is lower than the saliency threshold, the color similarity, texture features and the depth information are integrated to construct an affinity relationship between pixels, and the cluster center position and boundary are iteratively optimized until convergence to obtain the multidimensional feature image block; Performing a wavelet transform operation on the multidimensional feature image block, extracting frequency domain features by multi-scale decomposition, performing contrast enhancement on low-frequency components by piecewise linear mapping, and dynamically adjusting the enhancement coefficient according to the grayscale distribution of the multidimensional feature image block, performing adaptive threshold denoising on high-frequency components, wherein the threshold parameter is determined according to the local variance of the multidimensional feature image block, performing Laplace edge enhancement, and adaptively adjusting the sub-kernel coefficient according to the gradient amplitude of the multidimensional feature image block, uniformly scaling the enhanced image block to a standard size, and performing pixel value normalization processing to obtain a standard image block; The cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, introduces a channel attention mechanism in the residual structure of the convolutional neural network target detection model, and acquires statistical information of the channel dimension through global feature pooling to learn the importance weights of different channel features, wherein the main branch of the residual structure performs conventional convolution operations, the attention branch calculates channel weights, and the short-circuit branch maintains feature identity mapping, and an enhanced feature map is obtained by weighted fusion of the outputs of the three branches, and the multi-scale deep features and the key area features are extracted based on the enhanced feature map; Combined with the improved anchor-free target detection algorithm, the multi-scale deep features and the key area features are upsampled and aligned, and the feature information of different scales is fused by weighted summation. The center point position and size offset of the target are predicted through the convolution layer. The center point prediction branch outputs a heat map to represent the probability distribution of the target's existence, and the size prediction branch directly regresses the width and height ratio of the target. The prediction results are subjected to maximum suppression to remove overlapping detection boxes and bounding box adjustments to obtain abnormal detection results. The anomaly detection result and the depth information are integrated by a dual-stream structure through a feature fusion network guided by depth information, wherein the two feature branches of the dual-stream structure respectively include multi-layer convolution modules and each convolution module consists of a convolution layer, a normalization layer and an activation function. The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through a decoder network to obtain accurate three-dimensional characterization data using spatial position coordinates, posture angles and geometric dimensions of the target.

3. The method according to claim 2, characterized in that The features of the two branches are fused through adaptive weights determined by feature correlation and confidence. The fused features are reconstructed through the decoder network to obtain accurate three-dimensional representation data including: A first feature map corresponding to the anomaly detection result and a second feature map corresponding to the depth information are respectively extracted through a feature fusion network with a dual-stream structure, wherein the first feature map contains 512 feature channels and the second feature map contains 256 feature channels; Performing global average pooling and maximum pooling on the first feature map to extract a first channel-level feature descriptor, connecting the first channel-level feature descriptors in series through a fully connected layer to obtain a first channel attention weight, performing global average pooling and maximum pooling on the second feature map to extract a second channel-level feature descriptor, connecting the second channel-level feature descriptors in series through a fully connected layer to obtain a second channel attention weight; Expanding the first feature map and the second feature map into a first feature matrix and a second feature matrix according to channels, normalizing the feature vectors at each position in the first feature matrix and the second feature matrix, and calculating the inner product of the first feature matrix and the second feature matrix to obtain a correlation matrix; Using a pre-trained classifier to perform classification prediction on the first feature map to obtain a first prediction score, using the variance of the first prediction score as a first confidence indicator, and calculating a second confidence indicator based on the continuity and local consistency of the depth value in the second feature map; Calculating adaptive fusion weights in combination with the correlation matrix, the first confidence index, and the second confidence index, setting equal weights for regions where the correlation is higher than 0.5 and where both the first confidence index and the second confidence index are higher than 0.9, and biasing the weights toward the higher of the first confidence index and the second confidence index for regions where the correlation is lower than 0.5, and introducing context information to assist feature fusion for regions where both the first confidence index and the second confidence index are lower than 0.9; Weighting the first feature map and the second feature map according to the adaptive fusion weight to obtain a fusion feature, upsampling the fusion feature through a multi-layer decoder network and fusing jump connection features of different scales in each decoding layer; The depth map is back-projected into three-dimensional space to obtain a point cloud, the fusion feature is used to predict the position of the target center point in the point cloud to obtain the spatial position coordinates, the rotation quaternion of the target is predicted to obtain the attitude angle, the three-dimensional bounding box of the target is predicted to obtain the geometric size, and the spatial position coordinates, the attitude angle and the geometric size are combined to form accurate three-dimensional representation data.

4. The method according to claim 1, characterized in that: Based on the precise three-dimensional representation data and the anomaly detection results, a dynamic anomaly association graph is constructed through an improved time-series graph neural network. By adding time coding to the graph attention layer to model long-range temporal dependencies, the anomaly evolution law is obtained and an anomaly development prediction sequence is generated. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan. The optimal motion trajectory is generated through adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm, including: The target three-dimensional position coordinates, posture quaternion, geometric dimensions and anomaly detection confidence in the precise three-dimensional characterization data and the anomaly detection results are combined into a node feature vector, the Euclidean distance between the node feature vectors is calculated to obtain a spatial distance term, the cosine similarity of the node feature vectors is calculated to obtain a feature similarity term, the time series correlation term is calculated based on the motion consistency of the target at adjacent moments, and the weighted sum of the spatial distance term, the feature similarity term and the time series correlation term is used as the edge weight to construct a graph structure; The node feature vector is concatenated with the time encoding vector to obtain a fused feature, the time encoding vector is obtained by mapping the timestamp through a periodic sine function, the fused feature is input into the multi-head graph attention layer, each attention head independently calculates the dot product of the query vector and the key vector to obtain an attention score, the attention score is normalized and multiplied with the value vector to obtain a weighted feature, and the weighted features of multiple attention heads are concatenated and linearly transformed to obtain an output feature; Based on the output features, the target state change is predicted, including the position offset, the posture change, the size change and the abnormal degree change value, and the predicted state is used as a new input to continue predicting the state at the next moment. The teacher forcing mechanism is used to regularly use the real value to replace the predicted value in the prediction process to suppress the error accumulation and obtain the abnormal development prediction sequence; The abnormal development prediction sequence is input into a hierarchical reinforcement learning network, the top-level policy network assigns abnormal targets to designated positions to obtain a sub-target sequence, the middle-level policy network plans the execution sequence based on the sub-target sequence to obtain an execution sequence, and the bottom-level policy network generates an action sequence based on the execution sequence, wherein the reward function of each policy network is designed based on the sub-target completion, execution efficiency and trajectory safety respectively; The action sequence is subjected to adversarial training and experience replay optimization, wherein the generator adds random perturbations to the trajectory points to generate adversarial samples, the discriminator evaluates the difference between the true trajectory and the adversarial samples, and selects the sample with the largest temporal difference error for parameter update to obtain an optimized action sequence; The optimized action sequence is input into a task decomposition network, spatial features are extracted through multi-layer graph convolution and encoded into task representations, a decoder calculates a task relevance score based on the task representation through a bidirectional attention mechanism, an executor number, a target position, a start time and an expected duration are determined according to the task relevance score and the execution sequence, and a task allocation plan is generated; The sampling density is divided according to the obstacle distribution in the environment, and the kinematic constraints and obstacle avoidance constraints are evaluated for each sampling point in the task allocation scheme. The sampling point positions are iteratively optimized by gradient descent to satisfy the constraints and minimize the objective function. The optimized trajectory is smoothed by cubic spline interpolation to obtain the final motion trajectory.

5. The method according to claim 4, characterized in that Based on the output features, the target state change is predicted, including the position offset, posture change, size change and abnormal degree change value. The predicted state is used as a new input to continue predicting the state at the next moment. The teacher forcing mechanism is used to regularly use the real value to replace the predicted value during the prediction process to suppress the error accumulation to obtain the abnormal development prediction sequence including: The target's three-dimensional position coordinates, attitude quaternion values, geometric dimensions, and anomaly detection confidence values ​​are used as initial state data to construct a state sequence cache with a prediction time step length. The output features are respectively input into four independent fully connected layer prediction branches, the first prediction branch outputs the position offset of the target in three directions, the second prediction branch outputs the change in attitude angle represented by quaternion, the third prediction branch outputs the size scaling of the target in three dimensions of length, width and height, and the fourth prediction branch outputs the change in the confidence of abnormality detection; The target state is updated using a progressive prediction mechanism. The predicted position offset is added to the current position coordinate value to obtain the position coordinate value at the next moment. The predicted quaternion change is multiplied by the current attitude quaternion value to obtain the attitude angle value at the next moment. The predicted size scaling is multiplied by the current geometric size value to obtain the target size value at the next moment. The predicted abnormality change value is added to the current confidence value to obtain the abnormality degree value at the next moment. The prediction sequence is divided into multiple prediction intervals of equal length. At the beginning of each prediction interval, the real observation data is used to replace the prediction value as the input state. The real observation data includes the actual observed target position coordinates, attitude quaternion, geometric dimensions and anomaly detection confidence. At the middle moment of the prediction interval, the prediction state of the previous moment is used as input to continue predicting the state of the next moment. Physical constraint checks are performed on the predicted state values. The position offset between adjacent time steps does not exceed the maximum motion speed allowed by the physical system. The rotation angle change between adjacent time steps does not exceed the maximum rotation speed of the system. The size scaling is limited to the preset range of the original size. The abnormality value is kept between zero and one. The predicted values ​​that do not meet the constraints are corrected by linear interpolation. The state values ​​of each time step in the prediction sequence are sorted in timestamp order to generate an abnormal development prediction sequence containing occurrence time, three-dimensional spatial position, posture angle, geometric size and abnormality degree value.

6. The method according to claim 1, characterized in that The optimal motion trajectory is received, an improved dynamic path planning algorithm is executed and real-time obstacle avoidance navigation is performed in combination with a topological map and a sampling tree to obtain robot obstacle avoidance trajectory data, and the robot is accurately positioned by fusing depth information and real-time collected image features through an improved visual servo algorithm to obtain real-time posture data, and key action features are extracted through an improved imitation learning algorithm and a pre-built action value evaluation model, and dynamic skill transfer is performed in combination with an online adaptation module to obtain key operation sequence data and add it to an impedance adaptive control algorithm, and contact force feedback data is obtained by adaptively adjusting impedance parameters through a neural network compensator, and state monitoring is performed in combination with an improved abnormality prediction algorithm, and abnormality processing results are obtained, including: Receiving optimal motion trajectory data including position coordinates, attitude angle, motion speed and acceleration information, wherein the optimal motion trajectory data includes multiple path point information; Divide the workspace into multiple connected areas and construct an environmental topology map, record the connection relationship and travel cost between the connected areas, generate an initial path based on the environmental topology map, perform local expansion through a sampling tree on the basis of the initial path, dynamically adjust the sampling density according to the complexity of the environment, perform collision detection on the sampling points and calculate the shortest distance to the obstacles, and perform local path replanning when a collision risk is detected to generate robot obstacle avoidance trajectory data; Denoising and hole filling processing are performed on the depth image to obtain effective depth information, feature points are extracted from the image sequence and feature descriptors are established, feature matching is performed through a multi-level screening strategy, and mismatched points are eliminated using geometric consistency constraints. The effective depth information and feature points are weightedly fused according to the measurement confidence, and real-time pose data is obtained through pose optimization; Based on the rate of change of state variables, the key time points of the teaching data are identified to obtain action segmentation points, and feature vectors containing position, posture and speed information are extracted from the segmented action segments. The importance of the feature vectors is scored through the action value evaluation model to obtain key action features and dynamically adjust action parameters according to the characteristics of the target object and environmental constraints to generate key operation sequence data; A neural network compensator is used to learn the dynamic characteristics of the environment in real time, and position tracking and force control are achieved through a multi-layer nested control structure. The controller stiffness matrix and damping parameters are dynamically adjusted according to the force sensor feedback, and the contact force feedback data is recorded; Time domain features are extracted from the system state sequence through a sliding window. The time domain features include the statistical characteristics and change trends of the state variables. The time domain features are input into the prediction model to estimate the state evolution. When the prediction result shows an abnormality, the protection strategy is executed and the abnormality handling result is output.

7. The method according to claim 6, characterized in that Perform collision detection on the sampling points and calculate the shortest distance to obstacles. When a collision risk is detected, perform local path replanning to generate robot obstacle avoidance trajectory data including: The workspace is divided into basic grid units and an octree structure is constructed, and the obstacle area where obstacles exist is finely segmented by recursive partitioning, and the idle state, fully occupied state and partially occupied state are recorded in the leaf nodes of the octree structure, and the state information of the leaf nodes is updated in real time based on sensor data; The robot body is simplified into a multi-level sphere combination model consisting of a large-sized sphere corresponding to the trunk part and a small-sized sphere corresponding to the robotic arm part, and the adjacent nodes are searched in the octree structure based on the center position of the sphere of the multi-level sphere combination model to obtain the collision risk space area; The obstacle in the collision risk space region is estimated by using a spherical bounding box to obtain a preliminary distance. If the preliminary distance is less than a preset distance threshold, the obstacle surface in the collision risk space region is converted into point cloud data, and the closest point pair is determined by iterative search to obtain the actual shortest distance. The position, velocity and acceleration information of the obstacle are acquired by multi-sensor data fusion. The motion state of the obstacle is estimated by Kalman filter and the future position of the obstacle is predicted. If the future position of the obstacle intersects with the planned path, the obstacle avoidance planning is triggered. A sampling tree is established with the current position of the robot as the root node, and an adaptive sampling strategy related to the obstacle distribution is adopted to expand the sampling points to the surrounding space, and a cost evaluation is performed on the sampling points, wherein the cost evaluation includes the target distance, obstacle distance, path length and steering smoothness; Transition points are inserted between original path points to smooth the path. The positions of the transition points are determined by an optimization algorithm. The optimization algorithm takes minimizing the change in path curvature as the optimization goal, generates a complete motion trajectory including position coordinates, posture angles, linear velocity and angular velocity, detects environmental contact force and monitors position tracking errors through torque sensors, and replans the motion trajectory when the environmental contact force exceeds a safe range or the position tracking error exceeds a limit.

8. A robot alarm processing system based on a target detection algorithm and a cloud platform, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect image information and depth information corresponding to the industrial site based on a preset sensor array and upload them to the cloud platform, divide the image information and the depth information through an adaptive image segmentation algorithm to obtain a multidimensional feature image block, pre-process the multidimensional feature image block based on wavelet transform operation and Laplace edge enhancement to obtain a standard image block, and the cloud platform calls a convolutional neural network target detection model in a pre-trained target detection model library based on the image features of the standard image block, determines the multi-scale deep features and key area features corresponding to the standard image block in combination with the residual structure and channel attention mechanism, performs abnormal target detection in combination with an improved anchor-free target detection algorithm, obtains an abnormal detection result, and performs feature fusion through a feature fusion network guided by depth information to obtain accurate three-dimensional representation data; The second unit is used to construct a dynamic anomaly association graph through an improved time-series graph neural network based on the precise three-dimensional characterization data and the anomaly detection results, obtain the anomaly evolution law and generate an anomaly development prediction sequence by adding time coding to the graph attention layer to model the long-range time series dependency, and generate an anomaly development prediction sequence. Based on the anomaly development prediction sequence and the improved hierarchical reinforcement learning algorithm, hierarchical sub-goals are obtained by combining the hierarchical option framework decomposition and a multi-level decision sequence is generated by combining the reward mechanism. For the multi-level decision sequence, an action value evaluation model is constructed and the task execution strategy is optimized by combining the improved deep deterministic policy gradient algorithm with adversarial training and experience replay. The optimized task execution strategy is collaboratively planned through the task decomposition network through the bidirectional attention mechanism and the graph convolution layer to obtain a task allocation plan, and the optimal motion trajectory is generated through the adaptive sampling and dynamic constraint processing in the trajectory optimization algorithm; The third unit is used to receive the optimal motion trajectory, execute the improved dynamic path planning algorithm and perform real-time obstacle avoidance navigation in combination with the topological map and the sampling tree to obtain the robot's obstacle avoidance trajectory data, accurately locate the robot by fusing the depth information and the real-time collected image features through the improved visual servo algorithm to obtain real-time posture data, extract key action features through the improved imitation learning algorithm and the pre-built action value evaluation model, perform dynamic skill transfer in combination with the online adaptation module, obtain key operation sequence data and add it to the impedance adaptive control algorithm, obtain contact force feedback data by adaptively adjusting the impedance parameters through the neural network compensator, perform state monitoring in combination with the improved abnormality prediction algorithm, and obtain abnormality handling results.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Track planning method based on vision and radar feature fusion

    CN117944059A

  • Intelligent machine vision detection method and system based on image processing and storage medium

    CN119205719A