Chemical laboratory dynamic obstacle detection method and robot dynamic obstacle avoidance method
Through the detection models of feature extraction network, optical flow and stereo matching layer, propagation layer and iterative refinement layer, the problem of robotic arms detecting dynamic obstacles in chemical laboratories is solved, and efficient and real-time dynamic obstacle avoidance is achieved in low-texture environments.
Patent Information
- Application Number
- CN202510569180.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
AI Technical Summary
It is difficult for robotic robotic arms to accurately detect dynamic obstacles in chemical laboratories, especially in low-texture or single-color surfaces and high reflectivity materials. The existing methods increase model complexity and computational overhead, affecting real-time and applicability.
The detection model consisting of a feature extraction network, optical flow and stereo matching layer, propagation layer and iterative refinement layer is adopted to detect dynamic obstacles through the fusion of optical flow and parallax, including image acquisition, multi-scale feature extraction, optical flow and stereo matching, adaptive propagation and iterative refinement, and optimize prediction with weighted loss function.
It realizes efficient detection of dynamic obstacles in low-texture chemistry laboratory scenarios, reduces model updates and deployment complexity, and has real-time performance, and is suitable for dynamic obstacle avoidance of chemical robot robotic arms.
Smart Images

Figure CN120388355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dynamic obstacle avoidance for robots, and particularly to a method for dynamically detecting obstacles in a chemical laboratory based on the fusion of optical flow and parallax, and a method for dynamically avoiding obstacles for a robotic arm. Background Art
[0002] When the robotic arm carried by a robot is performing chemical experiment operations, there may be researchers beside it taking and placing chemical vessels, adding chemical reagents, or operating other instrument devices. If it operates according to the pre-planned trajectory, it may have a spatial conflict with the operations of the researchers. Therefore, the robot needs to have the ability to detect dynamic obstacles to ensure safety and the smooth execution of tasks. However, most of the instrument devices and walls in a chemical laboratory have low-texture or single-color surfaces, and the researchers wear white laboratory coats, which are highly similar to the laboratory background color and are not easy to distinguish. In addition, most of the surfaces of chemical instruments are made of materials with high reflectivity or transmittance, such as metals, glasses, etc. These materials make it difficult to directly apply 3D vision perception methods based on depth images or 3D point clouds. Although an additional depth recovery network can be used to complete the information and then fuse it with the optical flow to achieve the detection of dynamic obstacles, this also increases the complexity of the entire model and introduces additional computational overhead during the actual deployment and update process, thus having a certain impact on the real-time performance and applicability of the system.
[0003] In view of this, the present invention is specifically proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for dynamically detecting obstacles in a chemical laboratory and a method for dynamically avoiding obstacles for a robot, which can generate optical flow or parallax predictions based on a scene image, fuse the optical flow and parallax to accurately detect dynamic obstacles in the chemical laboratory, and thus well solve the above problems existing in the prior art.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A method for dynamically detecting obstacles in a chemical laboratory uses a detection model formed by sequentially connecting a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer for detection, including:
[0007] Step 1, image acquisition: Use a color camera to acquire two consecutive frame scene images during the movement of the robotic arm of the robot;
[0008] Step 2, feature extraction: Use the feature extraction network of the detection model to perform multi-scale feature extraction on the two frame scene images acquired in Step 1 and then perform relative position encoding to obtain two sets of multi-scale feature maps;
[0009] Step 3, optical flow matching and stereo matching:
[0010] After detecting the optical flow of the model and using the cross-attention mechanism in the stereo matching layer to perform feature fusion on the two sets of multi-scale feature maps obtained in step 2, calculate the feature similarity matrix and normalize it to obtain the matching distribution, calculate the matching coordinates of the pixels using the matching distribution, and obtain the preliminary optical flow prediction through coordinate difference to complete the optical flow matching;
[0011] Calculate the matching cost matrix of the two sets of multi-scale feature maps through the cross-attention mechanism of the optical flow and stereo matching layer, constrain the matching probability distribution through the entropy-regularized optimal transport method, solve the optimal transport matrix that satisfies the marginal distribution consistency from the matching cost matrix using the Sinkhorn-Knopp algorithm, and calculate the preliminary disparity prediction from the optimal transport matrix using the winner-takes-all strategy to complete the stereo matching;
[0012] Step 4, Adaptive Propagation: Using the self-similarity of the image features, through the propagation layer of the detection model, the prediction results of the high-quality parts in the preliminary optical flow prediction and the preliminary disparity prediction are propagated from the matched pixels to the unmatched pixel regions in the two sets of multi-scale feature maps to obtain a new high-quality optical flow prediction and a new high-quality disparity prediction;
[0013] Step 5, Iterative Refinement: Process the high-resolution features of the two sets of multi-scale feature maps obtained in step 2 input through the iterative refinement layer of the detection model, use the feature connection module of the iterative refinement layer to extract and integrate key context information, generate residual update amounts round by round starting from the initial prediction through the cascaded GRU module of the iterative refinement layer, and superimpose the residual update amounts on the refined matching results to obtain the refined optical flow prediction and disparity prediction;
[0014] Step 6, Upsampling Restoration: Perform 8-fold upsampling on the refined optical flow prediction and disparity prediction obtained in step 5 to restore the optical flow prediction and disparity prediction at the original resolution;
[0015] Step 7, Training Supervision: Use a weighted loss function with multi-stage supervision to perform training supervision on the optical flow prediction and disparity prediction at the original resolution obtained in step 6 through the detection model to obtain the optical flow prediction and disparity prediction with optimized accuracy;
[0016] Step 8, Dynamic Obstacle Detection: Screen out the dynamic obstacle regions from the first frame of the two-frame scene images through the optical flow prediction with optimized accuracy, and combine the depth information calculated from the disparity prediction with optimized accuracy to determine whether there is a collision risk with the dynamic obstacle regions.
[0017] A dynamic obstacle avoidance method for a chemical laboratory robot deploys a detection model formed by sequentially connecting a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer to the robot's controller. The detection model performs dynamic obstacle detection in the chemical laboratory according to the detection method described in the present invention. When it is confirmed that there is a potential collision risk with the detected dynamic obstacle area, an emergency stop operation is immediately executed to achieve dynamic obstacle avoidance.
[0018] Compared with the prior art, the dynamic obstacle detection method for a chemical laboratory and the dynamic obstacle avoidance method for a robot provided by the present invention have the following beneficial effects:
[0019] By adopting a detection model formed by sequentially connecting a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer as a unified detection framework applicable to optical flow prediction and disparity prediction, it can exceed the performance of the pure iterative structure RAFT and its variants after dozens of refinements with only 3 - 4 small iterations. In addition, the trained detection model can directly perform cross - task migration between the two tasks, significantly reducing the complexity of model update and deployment. The method of the present invention has a certain real - time performance and has been successfully verified in the chemical laboratory scenario with low texture and single color, and can be effectively used for dynamic obstacle avoidance of the mechanical arm of a chemical robot. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of the dynamic obstacle detection method for a chemical laboratory provided by the embodiment of the present invention.
[0022] Figure 2 It is a detection flowchart of the detection model of the dynamic obstacle detection method for a chemical laboratory provided by the embodiment of the present invention. Detailed Embodiments
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the specific content of the present invention; obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments, which do not constitute a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0024] First, the following explanations are given for the terms that may be used in this article:
[0025] The term "and / or" means that either or both of the two can be achieved. For example, X and / or Y means that it includes three cases: the case of "X" or "Y" and the case of "X and Y".
[0026] Descriptions with terms such as "comprising", "including", "containing", "having" or other similar semantics should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction condition, processing condition, parameter, algorithm, signal, data, product or article, etc.) should be interpreted as not only including the explicitly listed technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.
[0027] The term "consisting of" means excluding any technical feature element that is not explicitly listed. If this term is used in a claim, it will make the claim a closed type, so that it does not include technical feature elements other than the explicitly listed ones, except for conventional impurities related thereto. If this term only appears in a sub-clause of a claim, then it only limits the elements explicitly listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.
[0028] Unless otherwise clearly specified or limited, terms such as "installed", "connected", "joined", "fixed", etc. should be understood in a broad sense. For example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this article can be understood according to specific circumstances.
[0029] When a concentration, temperature, pressure, size or other parameter is expressed in the form of a numerical range, this numerical range should be understood as specifically disclosing all ranges formed by the pairing of any upper limit value, lower limit value, and preferred value within this numerical range, regardless of whether the range is explicitly recorded; for example, if the numerical range "2 - 8" is recorded, then this numerical range should be interpreted as including ranges such as "2 - 7", "2 - 6", "5 - 7", "3 - 4 and 6 - 7", "3 - 5 and 7", "2 and 5 - 7", etc. Unless otherwise stated, the numerical ranges recorded in this article include both their end values and all integers and fractions within this numerical range.
[0030] The following provides a detailed description of the solution provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. For those conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the reagents or instruments not specified in the embodiments of the present invention for the manufacturer, they are all conventional products that can be obtained through commercial purchase.
[0031] As Figure 1 and Figure 2 shown, an embodiment of the present invention provides a method for detecting dynamic obstacles in a chemical laboratory. A detection model formed by sequentially connecting a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer is used for detection, including:
[0032] Step 1, image acquisition: Use a color camera to collect two consecutive frame scene images during the movement of the robotic arm of the robot;
[0033] Step 2, feature extraction: Use the feature extraction network of the detection model to perform multi-scale feature extraction on the two frame scene images collected in Step 1, and then perform relative position encoding to obtain two sets of multi-scale feature maps;
[0034] Step 3, optical flow matching and stereo matching:
[0035] After using the cross-attention mechanism of the optical flow and stereo matching layer of the detection model to perform feature fusion on the two sets of multi-scale feature maps obtained in Step 2, calculate the feature similarity matrix and perform normalization to obtain the matching distribution, calculate the matching coordinates of the pixels using the matching distribution, and obtain the preliminary optical flow prediction through coordinate difference to complete the optical flow matching;
[0036] Calculate the matching cost matrix of the two sets of multi-scale feature maps through the cross-attention mechanism of the optical flow and stereo matching layer, constrain the matching probability distribution through the entropy-regularized optimal transport method, solve the optimal transport matrix that satisfies the marginal distribution consistency from the matching cost matrix using the Sinkhorn-Knopp algorithm, and calculate the preliminary disparity prediction from the optimal transport matrix using the winner-takes-all strategy to complete the stereo matching;
[0037] Step 4, adaptive propagation: Using the self-similarity of the image features, through the propagation layer of the detection model, the prediction results of the high-quality parts in the preliminary optical flow prediction and the preliminary disparity prediction are propagated from the matched pixels to the unmatched pixel regions in the two sets of multi-scale feature maps to obtain a new high-quality optical flow prediction and a new high-quality disparity prediction;
[0038] Step 5, Iterative Refinement: Process the high-resolution features of the two sets of multi-scale feature maps obtained in Step 2 through the iterative refinement layer of the detection model. Use the feature connection module of the iterative refinement layer to extract and integrate key context information. Starting from the initial prediction, generate residual update amounts round by round through the cascaded GRU module of the iterative refinement layer, and superimpose the residual update amounts on the refined matching results to obtain refined optical flow prediction and disparity prediction;
[0039] Step 6, Upsampling Restoration: Perform 8-fold upsampling on the refined optical flow prediction and disparity prediction obtained in Step 5 to restore the optical flow prediction and disparity prediction at the original resolution;
[0040] Step 7, Training Supervision: Use a weighted loss function for multi-stage supervision to perform training supervision on the optical flow prediction and disparity prediction at the original resolution obtained in Step 6 through the detection model, and obtain optical flow prediction and disparity prediction with optimized accuracy;
[0041] Step 8, Dynamic Obstacle Detection: Screen out the dynamic obstacle areas from the first frame of the two-frame scene images through the optical flow prediction with optimized accuracy, and combine the depth information calculated from the disparity prediction with optimized accuracy to determine whether there is a collision risk with the dynamic obstacle areas.
[0042] Preferably, in the feature extraction of Step 2 of the above method, the feature extraction network of the detection model is used to perform multi-scale feature extraction on the two-frame scene images collected in Step 1 in the following manner and then perform relative position encoding to obtain two sets of multi-scale feature maps, including:
[0043] Step 21, Split the two-frame scene images collected in Step 1 into multiple non-overlapping and equal-sized patches through the Patch division layer, and take the multiple patches divided from each frame of the scene image as a group;
[0044] Step 22, Use four Swin-T networks as the feature extraction network. Take every two Swin-T networks as a group to perform the shifted window multi-head self-attention mechanism and the window multi-head self-attention mechanism respectively, and alternately process the group of multiple patches corresponding to each frame of the scene image obtained in Step 21 to achieve local feature extraction and global information interaction to extract multi-scale features;
[0045] Step 23, Perform relative position encoding on the shifted window multi-head self-attention mechanism and the window multi-head self-attention mechanism in Step 22 respectively to obtain two sets of multi-scale feature maps.
[0046] Preferably, in Step 22 of the above method, a feed-forward network is provided inside the feature extraction network, and the representation ability of the Swin-T network is further enhanced through residual connection and layer normalization;
[0047] A Swin-T network only acts on the same frame of image. The queries, keys, and values of the self-attention mechanism of the Swin-T network all originate from the features of the same image, and the formula is as follows:
[0048] W = XW Q ;
[0049] K = XW K ;
[0050] V = XW V ;
[0051]
[0052] Among them, Q, K, and V are the queries, keys, and values in the self-attention mechanism of the Swin-T network respectively; is the input feature, N is the number of patches, d is the feature dimension of each patch; W Q , W K , W V are learnable weight matrices with dimensions of d×d k , d k is the feature dimension of each attention head; B is the relative position bias introduced additionally;
[0053] The two sets of multi-scale feature maps finally obtained and are respectively expressed as:
[0054]
[0055] Among them, are respectively the patches obtained by splitting the first frame image of the two-frame scene image after passing through the Patch division layer and the patches obtained by splitting the second frame image of the two-frame scene image after passing through the Patch division layer;
[0056] In step 23, the relative position encoding is performed on the shifted window multi-head self-attention mechanism and the window multi-head self-attention mechanism in step 22 respectively in the following manner to obtain two sets of multi-scale feature maps, including:
[0057] In the shifted window multi-head self-attention mechanism, construct a relative position bias table Index the relative position bias table with relative displacements in two directions to obtain the relative position bias B; for the data in the shifted window, the relative position bias table realizes relative position encoding through global coordinate mapping to relative indices;
[0058] In the window multi-head self-attention mechanism, rotational position embeddings are introduced. By designing a function f(x, p), a rotational transformation is performed according to the absolute position p of the feature x, so that the similarity between the transformed query and key features is only a function of q and k, and their relative positions s - t;
[0059] The relative position encoding in the window multi-head self-attention mechanism is implemented in the following way, including:
[0060] This feature is achieved by using a series of rotation matrices with different frequencies to group and rotate the feature dimensions; for two-dimensional image signals, the features are divided into two parts, and position embeddings in the x direction are applied to the first part, and position embeddings in the y direction are applied to the second part.
[0061] Preferably, in the optical flow matching of step 3 of the above method, the matching cost matrix of two groups of multi-scale feature maps is calculated by the cross-attention mechanism of the optical flow and stereo matching layer of the detection model in the following way. The matching probability distribution is constrained by the entropy-regularized optimal transport method, and the optimal transport matrix satisfying the marginal distribution consistency is solved from the matching cost matrix by the Sinkhorn-Knopp algorithm. The initial disparity prediction is calculated from the optimal transport matrix by the winner-takes-all strategy to complete stereo matching, including:
[0062] Step 311, obtain the bidirectional cross-attention features between two groups of multi-scale feature maps F1 and F2 obtained in step 2 through two cross-attention Transformer deployed by the optical flow and stereo matching layer. This process is symmetrically executed for the two groups of multi-scale feature maps F1 and F2. The query of the cross-attention comes from one of the multi-scale feature maps, and the key and value come from the other multi-scale feature map, and the first feature is obtained respectively and the second feature
[0063]
[0064] where, T cross is the cross-attention Transformer;
[0065] Step 312, use the dot-product-based correlation calculation method to obtain the feature similarity C between each pixel in the first feature and all pixels in the second feature :
[0066]
[0067] where, d F is the number of channels of the first feature and the second feature ; A normalization coefficient to avoid excessive values after dot product operations; is the set of real numbers, H is the height of the original resolution of the two input frames of scene images, W is the width of the original resolution of the two input frames of scene images, and the superscript T represents the transpose of the matrix;
[0068] Step 313, perform softmax differentiable processing through the following formula to obtain a weighted average matching distribution
[0069]
[0070] Step 314, through the weighted average matching distribution Calculate the correspondence with the 2D coordinates G of the pixel grid:
[0071]
[0072] Use the difference between the front and back coordinates to represent the preliminary optical flow prediction of the final sub-pixel accuracy as:
[0073]
[0074] In the disparity matching of step 3, the matching cost matrix of two groups of multi-scale feature maps is calculated by the cross-attention mechanism through the optical flow and the stereo matching layer in the following manner. The matching probability distribution is constrained by the entropy-regularized optimal transport method, and the optimal transport matrix satisfying the marginal distribution consistency is solved from the matching cost matrix using the Sinkhorn-Knopp algorithm. The preliminary disparity prediction is calculated from the optimal transport matrix using the winner-takes-all strategy to complete the stereo matching, including:
[0075] Step 321, calculate the one-way cross-attention scores of two groups of multi-scale feature maps F1 and F2 through two cross-attention Transformers deployed by the optical flow and the stereo matching layer
[0076]
[0077] where, T cross is the cross-attention Transformer;
[0078] Step 322, use the entropy-regularized optimal transport method to find the optimal transport matrix from the marginal distribution a of the first frame image to the marginal distribution b of the second frame image through the following optimal transport equation The optimal transport equation is:
[0079]
[0080] The constraint conditions of the optimal transport equation are:
[0081]
[0082] Among them, R + represents an I×I matrix with all non-negative elements; I w ×I w is the width of the edge distribution a of the first frame image and the width of the edge distribution b of the second frame image; w represents the matching probability between the i-th pixel in the first frame image and the j-th pixel in the second frame image among two frame scene images; M ij represents the cost of matching the i-th pixel in the first frame image with the j-th pixel in the second frame image, taking the negative value of the cross-attention score calculated in step 321; γ is a regularization parameter used to control the regularization intensity; is an entropy regularization term that makes the value of the optimal transport matrix more uniform to increase the distribution smoothness, defined as
[0083] Use the Sinkhorn-Knopp algorithm to iteratively solve the solution of the above optimal transport equation to obtain the optimal transport matrix
[0084] Step 323, use the following improved winner-takes-all method to obtain the initial disparity prediction for each pixel from the optimal transport matrix Specifically, determine the pixel index position k with the highest matching probability in the second frame image for the i-th pixel in the first frame image, which is:
[0085]
[0086] Construct a window N3(k) with a width of 3×3 pixels centered on the index position k, and renormalize the probabilities within the window N3(k) to obtain new probabilities
[0087]
[0088] Among them, represents the matching probability between the i-th pixel in the first frame image and the l-th pixel in the second frame image;
[0089] The new probabilities after renormalization reflect the relative matching strength of each pixel within the window N3(k);
[0090] Step 324, calculate the weighted pixel positions within the window N3(k) to obtain the preliminary disparity prediction
[0091] Preferably, in step 4 of the above method, the high-quality prediction results are propagated from the matched pixels to the unmatched regions in the two groups of multi-scale feature maps through the propagation layer in the following manner, including:
[0092] The high-quality prediction results are propagated from the matched pixels to the unmatched regions through the self-attention mechanism of the propagation layer to obtain new high-quality optical flow predictions and disparity predictions as follows:
[0093]
[0094] Preferably, in step 5 of the above method, the high-resolution features of the two groups of multi-scale feature maps obtained in step 2 are processed through the iterative refinement layer of the detection model in the following manner. The key context information is extracted and integrated by the feature connection module of the iterative refinement layer, and the cascade GRU module of the iterative refinement layer generates the residual update amount round by round starting from the initial prediction, and the residual update amount is superimposed on the refined matching result to obtain the refined disparity prediction, including:
[0095] Step 51, the iterative refinement layer consists of a feature connection module FCM cascaded with multiple GRU modules. Taking the high-resolution features in the feature pyramid of the two groups of multi-scale feature maps obtained in step 2 as the input, the context features required for each GRU module are generated through the processing of the feature connection module FCM. Starting from the initial disparity prediction f0, the GRU module calculates the residual update amount Δf i to gradually refine the prediction result. In each round of iteration, these residual update amounts are added to the disparity prediction result f of the previous round i-1 to generate the refined disparity prediction f i . Subsequently, the updated refined disparity prediction f i is passed to the next cascaded FCM-GRU module;
[0096] Step 52, for the feature connection module FCM described in step 51, it first calculates the bidirectional cross-attention scores of the high-resolution features f left of the first-frame multi-scale feature map and the high-resolution features F right of the second-frame multi-scale feature map in the two groups of multi-scale feature maps, extracts the feature similarity through the dot-product-based correlation calculation method according to the bidirectional cross-attention scores, and splices the extracted feature similarity with the disparity prediction result f i refined in the previous round of iteration to obtain the input feature x i of the GRU module.
[0097] Preferably, in the upsampling restoration of step 6 of the above method, the refined optical flow prediction and disparity prediction are upsampled 8 times in the following manner to restore to the original resolution to obtain the optical flow prediction and disparity prediction at the original resolution, including:
[0098] Both the refined optical flow prediction and disparity prediction are at 1 / 8 resolution. The restoration methods for the optical flow prediction and disparity prediction are the same. The full-resolution flow of each pixel is calculated as a weighted combination of the 3×3 grid of its coarse-resolution neighborhood. The weighted mask is predicted by a small 2-layer convolutional network with a size of H / 8×W / 8×8×8×9 for 8-fold upsampling;
[0099] After the above upsampling process, it is restored to the original resolution to obtain the optical flow prediction and disparity prediction at the original resolution.
[0100] Preferably, in the training supervision of step 7 of the above method, the multi-stage supervised optical flow weighted loss function L flow and the disparity weighted loss function L disp are respectively:
[0101]
[0102] where i represents the i-th stage of the model iteration; N is the number of iteration stages of the model; γ is the decay factor, set to 0.9; is the optical flow prediction loss, is the optical flow prediction result at the i-th stage of the model iteration, V gt is the optical flow ground truth; is the disparity prediction loss, is the disparity prediction result at the i-th stage of the model iteration, d gt is the disparity ground truth;
[0103] For the optical flow prediction loss at each iteration stage, it is expressed as the sum of the endpoint error and smoothness loss of the optical flow, which is:
[0104]
[0105] where n is the total number of pixels; are the predicted optical flow components at pixel p respectively; u gt,i (p), v gt,i (p) are the corresponding ground truth optical flow components at pixel p respectively; represent the gradient in the horizontal direction and the gradient in the vertical direction respectively;
[0106] Select the l1 loss as the disparity prediction loss at each stage which is:
[0107]
[0108] Among them, is the disparity predicted at the i-th round stage for point p, and d gt,i (p) is the ground truth disparity corresponding to point p.
[0109] Preferably, in the dynamic obstacle detection in step 8 of the above method, the dynamic obstacle region is screened out from the first frame image of two-frame scene images by optical flow prediction with optimized accuracy in the following manner, and whether there is a collision risk in the dynamic obstacle region is judged by combining the depth information calculated from the disparity prediction with optimized accuracy, including:
[0110] The region whose displacement amount differs from the background by a predetermined value is screened out by the obtained optical flow prediction with optimized accuracy, marked as the dynamic obstacle region, and the dynamic obstacle region is extracted as the optical flow mask;
[0111] Combined with the disparity prediction information, the region covered by the optical flow mask is selected as the region of interest in the disparity map, and the depth value D of the region of interest is calculated from the disparity prediction with optimized accuracy combined with the camera parameters according to the following formula as the depth value of the dynamic obstacle region, which is: where B is the baseline of the color camera; f is the focal length of the color camera; d is the upsampled disparity;
[0112] If there is a dynamic obstacle region in the current frame and the depth value of the dynamic obstacle region exceeds a predetermined threshold, it is confirmed that there is a potential collision risk with the dynamic obstacle region.
[0113] The embodiment of the present invention also provides a method for a robot in a chemical laboratory to dynamically avoid obstacles. The detection model formed by sequentially connecting a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer is deployed on the controller of the robot. The detection model performs dynamic obstacle detection in the chemical laboratory according to the above detection method. When it is confirmed that there is a potential collision risk with the detected dynamic obstacle region, an emergency stop operation is immediately executed to achieve dynamic obstacle avoidance.
[0114] In summary, the detection method of the embodiment of the present invention, by using the detection model formed by sequentially connecting a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer as a unified detection framework applicable to optical flow prediction and disparity prediction, can exceed the performance of the pure iterative structure RAFT and its variants after dozens of refinements with only 3 - 4 small iterations. In addition, the trained detection model can directly perform cross-task migration between the two tasks, significantly reducing the complexity of model update and deployment. The method of the present invention has a certain real-time performance and has been successfully verified in the chemical laboratory scene with low texture and single color, and can be effectively used for dynamic obstacle avoidance of the robotic arm of a chemical robot.
[0115] To more clearly demonstrate the technical solution provided by the present invention and the resulting technical effects, the following uses specific embodiments to describe in detail the solution provided by the embodiments of the present invention.
[0116] Embodiment 1
[0117] As shown in Figure 1 and Figure 2 This embodiment provides a method for dynamically detecting obstacles in a chemical laboratory, which is a dynamic obstacle avoidance method for a robotic arm based on the fusion of optical flow and parallax, and includes the following steps:
[0118] Step 1, data acquisition: Use a color camera to collect scene images during the movement of the robotic arm;
[0119] Step 2, feature extraction and relative position encoding, including the following steps:
[0120] Step 21, first split the two input RGB images into non-overlapping equal-sized patches through a Patch division layer, regarded as patch tokens.
[0121] Step 22, then the tokens obtained in Step 21 (the tokens from the same frame of scene image are processed as a group) are processed through four Swin-T structures to generate multi-scale features. Among them, every two Swin-T structures form a group and respectively execute the window multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism to alternately achieve local feature extraction and global information interaction. In addition, a feed-forward network is also included inside the structure to further enhance the representation ability of the network through residual connection and layer normalization. The self-attention mechanism only acts on the same image, and its query, key, and value all originate from the features of this image. The specific formula is as follows:
[0122] Q = XW Q ;
[0123] K = XW K ;
[0124] V = XW V ;
[0125]
[0126] Among them, Q, K, and V are respectively the query, key, and value in the attention mechanism, is the input feature, N represents the number of patches, d represents the feature dimension of each patch, is the learnable weight matrix, d k is the feature dimension of each head, and B is the relative position bias introduced additionally.
[0127] The finally obtained two multi-scale features and is expressed as:
[0128]
[0129] wherein, are the patches obtained by splitting the first frame image and the second frame image respectively after passing through the Patch division layer.
[0130] Step 23, perform relative position encoding on the two multi-head self-attention mechanisms in Step 22. In the shifted window multi-head self-attention mechanism, a relative position bias table is constructed The relative position bias B is obtained by indexing the table using relative displacements in two directions. For the data in the shifted window, the relative position bias table is mapped to relative indices through global coordinates, thus achieving effective encoding. In the window multi-head self-attention mechanism, Rotary Positional Embedding (RoPE) is introduced. By designing a function f(x,p), a rotation transformation is performed according to the absolute position p of the feature x, so that the similarity between the transformed query and key features is only a function of q, k, and their relative position s - t. Specifically, a series of rotation matrices with different frequencies are used to group and rotate the feature dimensions to achieve this property. For two-dimensional image signals, the features are divided into two parts, and position embedding in the x direction is applied to the first part, and position embedding in the y direction is applied to the second part.
[0131] Step 3, optical flow and stereo matching, includes the following steps:
[0132] Step 31, for optical flow matching, first deploy two cross-attention Transformers to obtain the bidirectional cross-attention features between the two sets of features F1 and F2 obtained in Step 22. This process is executed symmetrically for F1 and F2. The query of the cross-attention comes from one of the images, while the key and value come from the other image, obtaining the features and
[0133]
[0134] wherein, T cross is the cross-attention Transformer.
[0135] Step 32, use the dot product-based correlation calculation method to obtain the feature similarity C of each pixel in with all pixels in
[0136]
[0137] wherein, dF as a feature and the number of channels, is the normalization coefficient to avoid excessive values after dot product operations.
[0138] Step 33, then perform softmax differentiable processing to obtain the matching distribution
[0139]
[0140] Step 34, calculate the correspondence with the 2D coordinates G of the pixel grid by weighted averaging the matching distribution :
[0141]
[0142] Final rough optical flow (sub-pixel accuracy) is expressed as the difference in the front and back coordinates:
[0143]
[0144] Step 35, for stereo matching, first calculate the one-way cross-attention scores for F1 and F2:
[0145]
[0146] Step 36, then use the entropy-regularized optimal transport method to find the optimal transport matrix from the marginal distribution a of the first frame image to the marginal distribution b of the second frame image
[0147]
[0148] where represents the matching probability between the i-th pixel in the first frame image and the j-th pixel in the second frame image, M ij represents the cost of matching the i-th pixel in the first frame image with the j-th pixel in the second frame image, here taking the negative value of the cross-attention score . γ is the regularization parameter used to control the regularization strength, is the entropy regularization term, defined as Entropy regularization makes the value of the optimal transport matrix more uniform by increasing the smoothness of the distribution, thereby achieving soft matching. I w is the width of the marginal distributions a and b.
[0149] Use the Sinkhorn-Knopp algorithm to solve the solution of the optimal transport equation, which efficiently calculates the optimal transport matrix through an iterative method
[0150] Step 37. To obtain the initial disparity of each pixel from the optimal transport matrix , an improved winner-takes-all method is adopted. For the \(i\)-th pixel in the first frame image, the position \(k\) of the pixel in the second frame image with the highest matching probability is:
[0151]
[0152] Then, a window \(N_3(k)\) with a width of \(3\times3\) pixels is constructed around the index \(k\) to capture the matching probabilities within this window. At the same time, local weighting is introduced so that the final disparity prediction can better handle the problem of multiple similar matching patterns existing simultaneously. The probabilities within the window \(N_3(k)\) are renormalized to obtain the new probabilities
[0153]
[0154] The renormalized probabilities reflect the relative matching strengths of the pixels within the window.
[0155] Step 38. The rough disparity is calculated by weighting the disparity values of the pixels within the window
[0156]
[0157] Step 4. Adaptive propagation: It is implemented using a simple self-attention mechanism to propagate the high-quality prediction results from the matched pixels to the unmatched regions. The new high-quality optical flow and disparity are:
[0158]
[0159] Note that here there is no explicit distinction between matched and unmatched pixels, but only the process of this propagation is supervised and learned through the ground truth.
[0160] Step 5. Iterative refinement, which includes the following steps:
[0161] Step 51. An iterative refinement layer is constructed based on the GRU module and the Feature Connection Module (FCM) as a post-processing step of the model to make full use of the high-resolution features in the multi-scale feature pyramid. First, the high-resolution features in the feature pyramid obtained in Step 22 are selected as the input, and after being processed by the Feature Connection Module (FCM), the context features required for each GRU module are generated. Starting from the initial prediction \(f_0\), the GRU module updates the prediction by calculating the residual \(\Delta f\) i to gradually refine the prediction result. In each round of iteration, these residual update amounts are added to the prediction result \(f\) of the previous roundi-1 above, thus generating a refined prediction f i . Subsequently, the updated prediction f i will be passed to the FCM-GRU module in the next cascade.
[0162] Step 52, for the feature connection module FCM described in Step 51, it first calculates the bidirectional cross-attention scores of the high-resolution features F left , F rifht , then extracts the feature similarity through a dot-product-based correlation calculation method, and finally concatenates it with the prediction result f i refined in the previous iteration to obtain the input feature x i of the GRU.
[0163] Step 6, upsampling recovery, includes the following steps:
[0164] Both the iteratively refined optical flow prediction and the visual prediction are at 1 / 8 resolution. To restore the original resolution, the same method is used for both, which is to calculate the full-resolution flow of each pixel as a weighted combination of the 3×3 grid of its coarse-resolution neighborhood. The weighted mask is predicted by a small 2-layer convolutional network, with a size of H / 8*W / 8*8*8*9 for 8-fold upsampling;
[0165] After the above upsampling, the optical flow prediction and the disparity prediction at the original resolution are obtained.
[0166] Step 7, training supervision: To optimize the accuracy of the optical flow and disparity predictions, the multi-stage supervised optical flow weighted loss function L flow and the disparity weighted loss function L disp are used, respectively:
[0167]
[0168] where N is the number of iterative stages of the model, and γ is the decay factor (set to 0.9) to ensure that the weights of the later predictions are larger.
[0169] Specifically, for the optical flow prediction loss at each iterative stage, it is expressed as the sum of the endpoint error and the smoothness loss of the optical flow:
[0170]
[0171] where n is the total number of pixels, is the optical flow component predicted at point p, and u gt,i (p), v gt,i (p) are the ground truth optical flow components corresponding to point p. and Indicates the gradients in the horizontal and vertical directions.
[0172] Select the l1 loss as the disparity prediction loss for each stage
[0173]
[0174] Step 8, dynamic obstacle detection: The complete dynamic obstacle avoidance method is divided into two stages. In the first stage, the regions with a large displacement difference from the background are screened out using the obtained optical flow prediction, marked as dynamic obstacle regions, and extracted as masks. In the second stage, further combining the disparity prediction information, the regions of interest covered by the optical flow mask are selected in the disparity map, and the depth value D is calculated in combination with the camera parameters. Specifically:
[0175] Where B is the baseline of the camera, f is the focal length of the camera, and d is the upsampled disparity.
[0176] If there is a dynamic obstacle region in the current frame and the depth value of this region exceeds a certain threshold, it indicates a potential collision risk, and the manipulator is controlled to immediately perform an emergency stop operation.
[0177] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0178] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art section of this article is only intended to deepen the understanding of the overall background art of the present invention and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.
Claims
1. A method for dynamically detecting obstacles in a chemical laboratory, characterized in that, Detection is performed using a detection model composed of a feature extraction network, an optical flow and stereo matching layer, a propagation layer, and an iterative refinement layer connected in sequence, including: Step 1, Image acquisition: Use a color camera to acquire two consecutive frame scene images during the movement of the robotic arm of the robot; Step 2, Feature extraction: Use the feature extraction network of the detection model to perform multi-scale feature extraction on the two frame scene images acquired in Step 1 respectively, and then perform relative position encoding to obtain two sets of multi-scale feature maps; Step 3, Optical flow matching and stereo matching: Through the optical flow and stereo matching layer of the detection model, use the cross-attention mechanism to fuse the features of the two sets of multi-scale feature maps obtained in Step 2, calculate the feature similarity matrix and normalize it to obtain the matching distribution, calculate the matching coordinates of the pixels using the matching distribution, and obtain the preliminary optical flow prediction through coordinate difference to complete the optical flow matching; Through the optical flow and stereo matching layer, use the cross-attention mechanism to calculate the matching cost matrix of the two sets of multi-scale feature maps, constrain the matching probability distribution through the entropy-regularized optimal transport method, use the Sinkhorn-Knopp algorithm to solve the optimal transport matrix that satisfies the marginal distribution consistency from the matching cost matrix, and use the winner-takes-all strategy to calculate the preliminary disparity prediction from the optimal transport matrix to complete the stereo matching; Step 4, Adaptive propagation: Use the self-similarity of the image features, and through the propagation layer of the detection model, propagate the prediction results of the high-quality parts in the preliminary optical flow prediction and the preliminary disparity prediction from the matched pixels to the unmatched pixel regions in the two sets of multi-scale feature maps to obtain a new high-quality optical flow prediction and a new high-quality disparity prediction; Step 5, Iterative refinement: Process the high-resolution features of the two sets of multi-scale feature maps obtained in Step 2 input through the iterative refinement layer of the detection model, use the feature connection module of the iterative refinement layer to extract and integrate key context information, generate residual update amounts round by round starting from the initial prediction through the cascaded GRU module of the iterative refinement layer, and superimpose the residual update amounts on the refined matching results to obtain the refined optical flow prediction and disparity prediction; Step 6, Upsampling restoration: Perform 8-fold upsampling on the refined optical flow prediction and disparity prediction obtained in Step 5 to restore the optical flow prediction and disparity prediction at the original resolution; Step 7, Training supervision: Use a weighted loss function with multi-stage supervision to perform training supervision on the optical flow prediction and disparity prediction at the original resolution obtained in Step 6 through the detection model to obtain the optical flow prediction and disparity prediction with optimized accuracy; Step 8, Dynamic obstacle detection: Screen out the dynamic obstacle area from the first frame image of the two frame scene images through the optical flow prediction with optimized accuracy, and combine the depth information calculated from the disparity prediction with optimized accuracy to determine whether there is a collision risk with the dynamic obstacle area.
2. The chemical laboratory dynamic obstacle detection method according to claim 1, wherein In the feature extraction of Step 2, the feature extraction network of the detection model is used to perform multi-scale feature extraction on the two frame scene images acquired in Step 1 respectively in the following way, and then relative position encoding is performed to obtain two sets of multi-scale feature maps, including: Step 21: Split the two frames of scene images collected in Step 1 into multiple non-overlapping and equal-sized patches through a Patch division layer, and take the multiple patches divided from each frame of scene image as a group. Step 22: Use four Swin-T networks as feature extraction networks. Take every two Swin-T networks as a group and perform the shifted window multi-head self-attention mechanism and the window multi-head self-attention mechanism respectively to alternately process the multiple patches corresponding to each frame of scene image obtained in Step 21 to achieve local feature extraction and global information interaction and extract multi-scale features. Step 23: Perform relative position encoding on the shifted window multi-head self-attention mechanism and the window multi-head self-attention mechanism in Step 22 respectively to obtain two groups of multi-scale feature maps.
3. The chemical laboratory dynamic obstacle detection method according to claim 2, wherein In Step 22, a feed-forward network is provided in the feature extraction network, and the representation ability of the Swin-T network is further enhanced through residual connection and layer normalization. One Swin-T network only acts on the same frame of image. The query, key, and value of the self-attention mechanism of this Swin-T network all originate from the features of the same image. The formula is as follows: W = XW Q ; K = XW K ; V = XW V ; Among them, Q, K, and V are the query, key, and value in the self-attention mechanism of the Swin-T network respectively; is the input feature, N is the number of patches, d is the feature dimension of each patch; W Q 、W K 、W V are learnable weight matrices with dimensions of d×d k , d k is the feature dimension of each attention head; B is the relative position bias introduced additionally; The finally obtained two groups of multi-scale feature maps and are respectively expressed as: Among them, are respectively the patches obtained by splitting the first frame image of two-frame scene images after passing through the Patch division layer and the patches obtained by splitting the second frame image of two-frame scene images after passing through the Patch division layer; In Step 23, perform relative position encoding on the shifted window multi-head self-attention mechanism and the window multi-head self-attention mechanism in Step 22 respectively to obtain two groups of multi-scale feature maps in the following way, including: In the shifted window multi-head self-attention mechanism, a relative position bias table is constructed The relative position bias table is indexed with relative displacements in two directions to obtain the relative position bias B; for the data in the shifted window, the relative position bias table is mapped to relative indices through global coordinates to achieve relative position encoding; In the window multi-head self-attention mechanism, introduce rotational position embeddings. By designing a function f(x,p), perform a rotational transformation according to the absolute position p of the feature x, so that the similarity between the transformed query and key features is only a function of q and k, and their relative position s - t. Implement the relative position encoding in the window multi-head self-attention mechanism in the following way, including: Use a series of rotation matrices with different frequencies to perform grouping and rotation operations on the feature dimensions to achieve this property; for two-dimensional image signals, divide the features into two parts, apply the position embedding in the x direction to the first part, and apply the position embedding in the y direction to the second part.
4. The chemical laboratory dynamic obstacle detection method according to any one of claims 1-3, characterized in that, In the optical flow matching of Step 3, calculate the matching cost matrix of the two groups of multi-scale feature maps through the optical flow of the detection model and the cross-attention mechanism of the stereo matching layer in the following way. Constrain the matching probability distribution through the entropy-regularized optimal transport method. Solve the optimal transport matrix that satisfies the marginal distribution consistency from the matching cost matrix using the Sinkhorn-Knopp algorithm. Calculate the preliminary disparity prediction from the optimal transport matrix using the winner-takes-all strategy to complete stereo matching, including: Step 311, obtain the bidirectional cross-attention features between two sets of multi-scale feature maps F1 and F2 obtained in Step 2 through two cross-attention Transformers deployed by the optical flow and stereo matching layer. This process is symmetrically performed on the two sets of multi-scale feature maps F1 and F2. The query of the cross-attention comes from one of the multi-scale feature maps, and the key and value come from the other multi-scale feature map, respectively obtaining the first feature and the second feature Among them, T cross is the cross-attention Transformer; Step 312, obtain the first feature using the dot product-based correlation calculation method for each pixel in and the feature similarity C of all pixels in the second feature where d F is the number of channels of the first feature and the second feature ; is a normalization coefficient to avoid an overly large value after the dot product operation; is a set of real numbers, H is the height of the original resolution of the two input frames of the scene image, W is the width of the original resolution of the two input frames of the scene image, and the superscript T represents the transpose of the matrix; Step 313, perform softmax differentiable processing through the following formula to obtain a weighted average matching distribution Step 314, calculate the correspondence with the 2D coordinates G of the pixel grid by weighted average matching distribution Calculate the correspondence with the 2D coordinates G of the pixel grid: The preliminary optical flow prediction representing the final sub-pixel accuracy with the difference between the front and rear coordinates is as follows: In the disparity matching of Step 3, calculate the matching cost matrix of the two groups of multi-scale feature maps through the optical flow and the cross-attention mechanism of the stereo matching layer in the following way. Constrain the matching probability distribution through the entropy-regularized optimal transport method. Solve the optimal transport matrix that satisfies the marginal distribution consistency from the matching cost matrix using the Sinkhorn-Knopp algorithm. Calculate the preliminary disparity prediction from the optimal transport matrix using the winner-takes-all strategy to complete stereo matching, including: Step 321, calculate the one-way cross-attention scores of two groups of multi-scale feature maps F1 and F2 through two cross-attention Transformers deployed by the optical flow and stereo matching layer Among them, T cross is a cross-attention Transformer; Step 322: Use the entropy-regularized optimal transport method to find the optimal transport matrix from the marginal distribution \(a\) of the first-frame image to the marginal distribution \(b\) of the second-frame image of two-frame scene images through the following optimal transport equation The optimal transport equation is as follows: The constraint condition of the optimal transport equation is: Among them, R + represents an I w ×I w matrix; I w is the width of the edge distribution a of the first frame image and the width of the edge distribution b of the second frame image; represents the matching probability between the i-th pixel in the first frame image of two frame scene images and the j-th pixel in the second frame image of two frame scene images; M ij represents the cost of matching the i-th pixel in the first frame image with the j-th pixel in the second frame image, taking the negative value of the cross-attention score calculated in step 321 ; γ is a regularization parameter used to control the regularization intensity; is an entropy regularization term that makes the value of the optimal transport matrix more uniform to increase the distribution smoothness, defined as The solution of the above optimal transport equation is obtained by iteratively solving it with the Sinkhorn-Knopp algorithm to obtain the optimal transport matrix Step 323, use the following improved winner-takes-all method to obtain the initial disparity prediction for each pixel from the optimal transport matrix Specifically, determine the pixel index position k with the highest matching probability for the i-th pixel in the first frame image in the second frame image, which is: Construct a window N3(k) with a width of 3×3 pixels centered at the index position k, and renormalize the probabilities within the window N3(k) to obtain new probabilities Among them, represents the matching probability of the i-th pixel in the first frame image and the matching probability of the l-th pixel in the second frame image; The new probability after re-normalization reflects the relative matching strength of each pixel within the window N3(k); Step 324, weighted calculation of the pixel positions within window N3(k) gives the preliminary disparity prediction 5. The chemical laboratory dynamic obstacle detection method according to claim 4, characterized in that In step 4, the high-quality prediction results are propagated from the matched pixels to the unmatched regions in the two sets of multi-scale feature maps through the propagation layer in the following manner, including: Propagate the high-quality prediction results from the matched pixels to the unmatched regions through the propagation layer using the self-attention mechanism to obtain a new high-quality optical flow prediction and disparity prediction as follows:
6. The dynamic obstacle detection method for a chemical laboratory for a robot's dynamic obstacle avoidance according to any one of claims 1-3, characterized in that In step 5, the high-resolution features of the two sets of multi-scale feature maps obtained in step 2 are processed through the iterative refinement layer of the detection model in the following manner. The key context information is extracted and integrated using the feature connection module of the iterative refinement layer. The cascaded GRU module of the iterative refinement layer generates residual update amounts round by round starting from the initial prediction, and the residual update amounts are superimposed on the refined matching result to obtain the refined disparity prediction, including: Step 51, the iterative refinement layer consists of multiple GRU modules cascaded by a feature connection module FCM. Taking the high-resolution features in the feature pyramid of the two sets of multi-scale feature maps obtained in Step 2 as inputs, the feature connection module FCM processes them to generate the context features required for each GRU module. Starting from the initial disparity prediction f0, the GRU module calculates the residual update Δf i to gradually refine the prediction result. In each iteration, these residual updates are added to the disparity prediction result f of the previous round i-1 to generate the refined disparity prediction f i . Subsequently, the updated refined disparity prediction f i is passed to the FCM-GRU module in the next cascade; Step 52, for the feature connection module FCM described in Step 51, it first calculates the high-resolution feature F of the first-frame multi-scale feature map in two groups of multi-scale feature maps left and the high-resolution feature F of the second-frame multi-scale feature map right of the bidirectional cross-attention score. According to the bidirectional cross-attention score, the feature similarity is extracted through a dot-product-based correlation calculation method, and the extracted feature similarity is concatenated with the disparity prediction result f refined in the previous iteration i to obtain the input feature x of the GRU module i .
7. The chemical laboratory dynamic obstacle detection method according to any one of claims 1-3, characterized in that In the upsampling restoration of step 6, the refined optical flow prediction and disparity prediction are upsampled by 8 times in the following manner to restore to the original resolution to obtain the optical flow prediction and disparity prediction at the original resolution, including: Both the iteratively refined optical flow prediction and disparity prediction are at 1 / 8 resolution. The restoration methods for the optical flow prediction and disparity prediction are the same. The full-resolution flow of each pixel is calculated as a weighted combination of the 3×3 grid in its coarse-resolution neighborhood. The weighted mask is predicted by a small 2-layer convolutional network with a size of H / 8×W / 8×8×8×9 for 8-fold upsampling; After the above upsampling process, the optical flow prediction and disparity prediction at the original resolution are restored to the original resolution.
8. The dynamic obstacle detection method for a chemical laboratory according to any one of claims 1-3, characterized in that In the training supervision of step 7, the multi-stage supervised optical flow weighted loss function L flow and the disparity weighted loss function L disp are respectively as follows: Among them, i represents the i-th stage of model iteration; N is the number of iteration stages of the model; γ is the decay factor, set to 0.9; is the optical flow prediction loss, is the optical flow prediction result at the i-th stage of model iteration, V gt is the optical flow ground truth; is the disparity prediction loss, is the disparity prediction result at the i-th stage of model iteration, d gt is the disparity ground truth; Optical flow prediction loss for each iteration stage It is expressed as the sum of the endpoint error of the optical flow and the smoothness loss, which is: where n is the total number of pixels; are the optical flow components predicted for pixel p; u gt,i (p), v gt,i (p) are the ground truth optical flow components corresponding to pixel p, respectively; represent the gradient in the horizontal direction and the gradient in the vertical direction, respectively; Select the l1 loss as the disparity prediction loss for each stage which is Among them, is the parallax predicted at the i-th round stage for point p, d gt,i (p) is the ground truth of the parallax corresponding to point p.
9. The dynamic obstacle detection method for a chemical laboratory according to any one of claims 1-3, characterized in that, In the dynamic obstacle detection of step 8, the dynamic obstacle regions are screened from the first frame image of the two-frame scene images through the optical flow prediction with optimized accuracy in the following manner, and the depth information calculated from the disparity prediction with optimized accuracy is combined to determine whether there is a collision risk in the dynamic obstacle regions, including: The regions with displacement differences from the background meeting a predetermined value are screened out using the obtained optical flow prediction with optimized accuracy, marked as dynamic obstacle regions, and the dynamic obstacle regions are extracted as the optical flow mask; Combined with the parallax prediction information, the area covered by the optical flow mask in the parallax map is selected as the region of interest. According to the following formula, the depth value D of the region of interest is calculated from the parallax prediction with optimized accuracy combined with the camera parameters as the depth value of the dynamic obstacle region, which is: where B is the baseline of the color camera; f is the focal length of the color camera; d is the upsampled parallax; If there are dynamic obstacle regions in the current frame and the depth value of the dynamic obstacle regions exceeds a predetermined threshold, it is confirmed that there is a potential collision risk with the dynamic obstacle regions.
10. A method for a robot in a chemical laboratory to dynamically avoid obstacles, characterized in that, The detection model formed by sequentially connecting the feature extraction network, the optical flow and stereo matching layer, the propagation layer, and the iterative refinement layer is deployed on the controller of the robot. The detection model performs dynamic obstacle detection in the chemical laboratory according to the detection method described in any one of claims 1-9. When it is confirmed that there is a potential collision risk with the detected dynamic obstacle regions, an emergency stop operation is immediately performed on the robot to achieve dynamic obstacle avoidance.
Citation Information
Cited By
AMR robot operation intelligent obstacle avoidance method and system based on multi-view vision
CN121433264A