Urine test tube automatic sorting method based on mechanical arm
Through deep learning technology based on robotic arms, automatic sorting of urine test tubes is achieved, which solves the problems of inefficient sorting efficiency and safety risks in the existing technology, improves the accuracy and safety of sorting, and provides technical support for intelligent medical systems.
Patent Information
- Application Number
- CN202510138451.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art test tube sorting task is inefficient in urine testing scenarios, with safety risks and high infection possibility.
The automatic sorting method of urine test tubes based on robotic arm is used to extract test tube image features through deep learning technologies (such as ResNet-101 and residual high-efficiency spatial pyramid module), and combine feature fusion modules and task subnets to predict the robotic arm grab parameters to realize the unmanned intervention sorting and classification of test tubes.
It realizes efficient and accurate sorting of test tubes, reduces the risk of secondary pollution, improves the safety of the medical environment, and provides technical support for the development of intelligent medical systems.
Smart Images

Figure SMS_1 
Figure SMS_4 
Figure SMS_8
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and robotic arm grasping control, and specifically relates to an automatic sorting method for urine test tubes based on a robotic arm. Background Art
[0002] Introducing artificial intelligence and robotic technology into the medical field can not only provide innovative support in theory and practice for doctors, researchers and patients, but also effectively improve the diagnosis and treatment effect, reduce medical costs, and contribute to the rational allocation of medical resources.
[0003] In recent years, the application of robotic technology has been increasing in many fields such as automated factories and daily life (such as hospitals and distribution stations). Currently, many domestic hospitals receive a large number of patients every day and face various complex urine test requirements. Urine test tubes need to be classified according to test items before being sent to different inspection departments. In many cases, urine test tubes are randomly placed on medical trays, especially in clinical wards and emergency departments. Sorting urine test tubes manually alone poses a risk of secondary contamination, and excessive shaking of the test tubes during the sorting process may affect the accuracy of subsequent tests. Therefore, in this context, it is particularly crucial to introduce an automated urine test tube sorting technology based on a robotic arm.
[0004] The automated urine test tube sorting technology can efficiently and accurately complete the sorting and classification of urine test tubes without manual intervention. At the same time, this technology can effectively prevent the test tubes from being secondarily contaminated during the processing, thus ensuring the cleanliness and safety of the medical environment and further promoting the progress and development of intelligent medical systems. Summary of the Invention
[0005] In view of the limitations of the prior art in the test tube sorting task in the urine test scenario, especially in the face of challenges such as low efficiency, safety risks and high infection possibilities, the present invention proposes an automatic sorting method for urine test tubes based on a robotic arm. This technology aims to achieve non-manual intervention sorting and classification of test tubes, ensure the safety of the test tubes and the reliability of the tests, and lay a technical foundation for building an efficient and intelligent medical system.
[0006] An automatic sorting method for urine test tubes based on a robotic arm includes the following steps:
[0007] (1) Collect test tube images;
[0008] (2) Use ResNet-101 to extract the preliminary extraction feature map of the test tube image;
[0009] (3) Use multiple parallel-structured residual efficient spatial pyramid modules with different dilation rates and groups to process the preliminary extraction feature map respectively to obtain the corresponding spliced feature map;
[0010] (4) Use the feature fusion module to perform multi-scale information fusion on the spliced feature map and the preliminarily extracted feature map to generate a large-scale feature map and a small-scale feature map;
[0011] (5) Respectively perform splicing operations on the obtained large-scale feature map, small-scale feature map, and preliminarily extracted feature map to obtain a grasping feature map;
[0012] (6) Input the obtained grasping feature map into the task sub-network, predict the robotic arm grasping parameters, and use the robotic arm to perform test tube sorting.
[0013] In step (1), the test tube image is a test tube RGB image; for image acquisition, we use a depth camera vision platform designed by us and implement RGB image acquisition and management using a Python program, which can automatically store images.
[0014] After adopting the RGB image, the following preprocessing can be carried out: Crop the RGB image obtained in step (1) to retain the test tube area, remove the irrelevant background, reduce the calculation amount and improve the recognition accuracy. At the same time, standardize the image size for subsequent processing.
[0015] Furthermore, when training the ResNet-101, the residual efficient spatial pyramid module, the feature fusion module, and the task sub-network, use the test tube RGB image containing a single test tube and the corresponding label for training. After training, it can be used for sorting in a single test tube scenario, or combined with existing multi-task learning methods for sorting in a multi-test tube scenario.
[0016] Based on the preprocessed RGB image in step (2), use ResNet-101 as the backbone network, extract the key features of the RGB image through residual mapping, and generate a preliminary feature map (preliminarily extracted feature map).
[0017] Furthermore, there are eight residual efficient spatial pyramid modules, including 4 different dilation rates and 2 groups. Based on the preliminary feature map obtained in step (2) of the present invention, it is processed by 8 parallel REASP blocks, and multi-scale features are extracted by convolutions with different dilation rates and groups, and a rich spliced feature map is generated through a splicing operation.
[0018] Furthermore, the spliced feature map includes a high dilation rate spliced feature map and a low dilation rate spliced feature map; the high dilation rate spliced feature map is obtained by extracting and splicing with a residual efficient spatial pyramid module with a high dilation rate, and the low dilation rate spliced feature map is obtained by extracting and splicing with a residual efficient spatial pyramid module with a low dilation rate. Taking different dilation rates (rate = {1, 6, 9, 12}) and groups (num = {1, 2}) as an example, the low dilation rate spliced feature map is obtained by extracting and splicing with rate = {1, 6}, num = {1, 2}; the high dilation rate spliced feature map is obtained by extracting and splicing with rate = {9, 12}, num = {1, 2}.
[0019] In the present invention, by introducing a feature fusion module, based on the preliminary feature map obtained in step (2) and the spliced feature map obtained in step (3), multi-scale information is fused through convolution and upsampling operations to generate large-scale and small-scale feature maps, improving the target recognition accuracy.
[0020] Furthermore, the feature fusion module combines features at different levels by using multiple convolution and upsampling operations to obtain large-scale feature maps and small-scale feature maps respectively.
[0021] In step (5), according to the large-scale feature map and the small-scale feature map obtained in step (4), a new grasping feature map is obtained through a splicing operation. Further, in step (5), the splicing operation is specifically as follows:
[0022] The small-scale feature map and the preliminary extraction feature map are spliced, and then spliced with the large-scale feature map to obtain the grasping feature map A;
[0023] The large-scale feature map and the small-scale feature map are spliced, and then spliced with the preliminary extraction feature map to obtain the grasping feature map B;
[0024] The grasping feature map is composed of the grasping feature map A and the grasping feature map B.
[0025] In step (6), based on the new grasping feature map in step (5), 4 task sub-networks are designed, and the SENet attention mechanism is introduced to predict the grasping quality, category, angle and width. Finally, the best grasping point and configuration information are output, including parameters such as (u, v) coordinates, grasping category, angle and width, improving the grasping performance.
[0026] Furthermore, the task sub-network includes four sub-networks, namely the quality sub-network, the grasping category sub-network, the grasping angle sub-network and the grasping width sub-network.
[0027] Furthermore, the input of the grasping quality sub-network and the grasping category sub-network is the grasping feature map A, and the input of the grasping angle sub-network and the grasping width sub-network is the grasping feature map B.
[0028] Furthermore, the four sub-networks have the same structure, all including a convolutional layer, a transposed convolutional layer, a concatenation layer, and an upsampling layer. At the same time, the attention mechanism SENet is introduced into each sub-network.
[0029] The (u, v) coordinates are converted into the coordinates of the robotic arm through coordinate transformation and finally transmitted to the robotic arm controller to perform grasping. Further, in step (6):
[0030] (6-1) First, use the four sub-networks to pre-test the grasping category parameters of each test tube in the test tube image: grasping quality, grasping category, grasping angle, and grasping width;
[0031] (6-2) Then, determine whether there is a test tube with a grasping category of collision-free grasping points. If so, proceed to the next step; if not, return to step (1) to re-acquire the test tube image;
[0032] (6-3) Convert the coordinates of the best grasping point with the highest grasping quality into the coordinates of the robotic arm, and at the same time adjust the grasping angle and grasping width of the robotic arm to perform test tube sorting.
[0033] For the multi-test tube scenario, in the case of collision-free grasping points, manual intervention on the test tube tray can be adopted. For example, the test tube tray can be shaken horizontally with a set force to further separate the test tubes, and then return to step (1) to achieve the sorting of each test tube one by one.
[0034] Compared with the prior art, by using the method of the present invention, the robotic arm can safely and stably perform the grasping action and complete the sorting task of urine test tubes. Description of the Drawings
[0035] Figure 1 is the overall flowchart of an automatic sorting method for urine test tubes based on a robotic arm.
[0036] Figure 2 is the structural diagram of the REASP block with a parallel structure of 8.
[0037] Figure 3 is the algorithm flowchart of the specific task sub-network.
[0038] Figure 4 is the overall flowchart of the network model.
[0039] Figure 5 is the specific physical diagram of this patent application.
[0040] Figure 6is an input image in the embodiment.
[0041] Figure 7 is the cropped image. Detailed implementation manners
[0042] To describe the present invention in more detail, the technical solutions of the present invention will be described in detail below in conjunction with specific embodiments. However, the scope to be protected by the present invention is not limited thereto.
[0043] Below is a method for automatically sorting urine test tubes based on a robotic arm. Before sorting, when training the ResNet-101, residual efficient spatial pyramid module, feature fusion module, and task sub-network, a test tube RGB image containing a single test tube and the corresponding label are used for training by using a conventional training method. The overall flowcharts of training and actual sorting are as Figure 5 shown.
[0044] ⑴ Image acquisition
[0045] The present invention adopts a carefully designed robot vision platform, in which the camera is fixedly installed at a position 45° directly above the left of the robotic arm and at a height of about 1 meter from the working plane. This specific installation position ensures that the camera can cover the entire operation area, ensuring comprehensive observation of the target object without affecting the normal operation of the robotic arm.
[0046] To achieve efficient acquisition and management of images, we developed an image acquisition program using the Python programming language. This program can not only automatically create directories to store the acquired images, but also continuously capture RGB images, providing data support for subsequent image processing and analysis.
[0047] Specifically, during the process of capturing RGB images, when the user presses the s key, the program saves the current image to the previously created directory. If the user presses the q or Esc key, the program exits. Figure 6 is an acquired image, which shows a situation where there is only one test tube in the test tube tray.
[0048] ⑵ Image preprocessing
[0049] We performed cropping processing on the images acquired in step (1), as Figure 7 shown. Specifically, during the cropping process, irrelevant background information in the image is removed, and the urine test tube in the central area is retained, so that each image focuses on the target object to be grasped. This operation not only reduces the amount of calculation, but also improves the accuracy of test tube recognition. In addition, by cropping, the sizes of all images are standardized, ensuring that the input images have a consistent format, which is convenient for subsequent processing.
[0050] ⑶ Extract key features
[0051] Select ResNet-101 as the backbone network, which is responsible for extracting key features from the input RGB image (the image obtained in step 2). The core of ResNet is to construct a deep network through residual mapping, and the form of the residual block is expressed as:
[0052] y = F(x,{W i}) + x (1)
[0053] where F(x,{W i}) is a residual mapping, x is the input feature map (i.e., the input RGB image), y is the output feature map, and {W i} are the weight parameters in the residual mapping. Through this step, we can obtain the preliminarily extracted feature map.
[0054] ⑷ Further extract key features
[0055] Use the preliminarily extracted feature map obtained in the above step (3) as the input and transmit it to the REASP block with 8 parallel structures (full name: Residual Efficient Atrous Spatial Pyramid, Chinese: Residual Efficient Spatial Pyramid, the structure diagram of the REASP block with 8 parallel structures is as Figure 2 shown, where + represents Addition and C represents Concatenation) to improve the utilization efficiency of information. Each pyramid is constructed using sub-attribute convolutions with different dilation rates (rate = {1, 6, 9, 12}) and groups (num = {1, 2}). This configuration allows for feature extraction at various scales, thus allowing for a more comprehensive capture of rich semantic information. The formula is as follows:
[0056] REASP(x) = Concat(Conv rate (x), Conv rate (x)) (2)
[0057] where Conv rate (x) represents the convolution operation with the specified dilation rate rate. Concat() represents the operation of concatenating feature maps. The low-dilation-rate concatenated feature map is obtained by extraction and concatenation with rate = {1, 6} and num = {1, 2}; the high-dilation-rate concatenated feature map is obtained by extraction and concatenation with rate = {9, 12} and num = {1, 2}. Through this step, we can obtain the low-dilation-rate concatenated feature map or the high-dilation-rate concatenated feature map.
[0058] ⑸ Add a feature fusion module
[0059] The feature fusion module fuses the features from the backbone network and the feature extraction module (8 parallel REASP blocks). This module allows features at different levels to be combined to obtain mixed large-scale and small-scale features, thus obtaining more comprehensive and accurate information during the segmentation process. The feature fusion module includes convolution and upsampling operations to ensure that the fused features have appropriate resolution and semantic information.
[0060] The specific fusion steps are as follows Figure 4 As shown, C represents the Concatenation operation. Specifically, it includes three layers, namely Layer I, Layer II, and Layer III. The outputs of Layer I and Layer II are concatenated to obtain a large-size feature map. The output of Layer III is concatenated with the preliminarily extracted feature map output by ResNet-101 to obtain a small-size feature map. Further, Layer I contains a first convolutional layer, an upsampling layer, and a second convolutional layer. The output of the low-dilation rate concatenated feature map obtained by the REASP module after being processed by the first convolutional layer and the upsampling layer is concatenated with the preliminarily extracted feature map output by sNet-101, and then input into the second convolutional layer for processing to obtain the output of this layer; in Layer II, it contains a convolutional layer and an upsampling layer. The high-dilation rate concatenated feature map obtained by the REASP module is processed by this convolutional layer and the upsampling layer to obtain the modified output; in Layer II, it includes a convolutional layer and an upsampling layer. The high-dilation rate concatenated feature map obtained by the REASP module is processed by this convolutional layer and the upsampling layer to obtain the modified output.
[0061] Through such a design, it is possible to more effectively extract key features from the input RGB image and fuse information at multiple scales, thereby improving the accuracy in identifying and locating targets. Through this step, we can obtain multi-scale feature maps: large-scale feature maps and small-scale feature maps.
[0062] ⑹ Prepare the grasping configuration
[0063] Based on the large-scale and small-scale feature maps obtained in step (5), a new grasping feature map is obtained through the concatenation operation.
[0064] The small-scale feature map and the preliminarily extracted feature map are concatenated, and then concatenated with the large-scale feature map to obtain the grasping feature map A.
[0065] The large-scale feature map and the small-scale feature map are concatenated, and then concatenated with the preliminarily extracted feature map to obtain the grasping feature map B.
[0066] Different concatenation orders result in different feature information contained in their feature maps. The grasping feature map A pays more attention to regional information and does not contain the edge information of large-scale objects. The grasping feature map B contains all the information of objects at different scales.
[0067] ⑺ Generate grasping configuration information
[0068] Based on the grasping feature map A and the grasping feature map B obtained in step (6), design 4 task sub-networks (the structure is as Figure 3 shown. The structures of the 4 task sub-networks are the same, only the input feature maps are different). The input of the grasping quality sub-network and the grasping category sub-network is the grasping feature map A, and the input of the grasping angle sub-network and the grasping width sub-network is the grasping feature map B.
[0069] In the entire network model, consider the grasping quality score detection as a binary classification task, the grasping width detection as a regression task, and the grasping category and grasping angle detections as multi-classification tasks. The following are the loss functions for each output task head:
[0070] ① Grasping quality score: First, use the sigmoid function to normalize the prediction result, and then use the binary cross-entropy function (BCE) to calculate the loss, which is defined as:
[0071]
[0072] In the formula, N is the size of the output feature map, is the predicted probability at pixel position n, is the corresponding true quality score.
[0073] ② Grasping angle: After normalizing the angle output using the sigmoid function, use the binary cross-entropy function to calculate the grasping angle loss, which is defined as:
[0074]
[0075] In the formula, represents the probability that the predicted grasping angle at pixel position n is within the range, is the corresponding label, and S defines a valid range or threshold for the grasping angle to determine whether the grasping angle predicted by the model is within a reasonable range.
[0076] ③ Grasping width: Use the binary cross-entropy function to calculate the loss of the grasping width branch. The formula is as shown in Equation 5:
[0077]
[0078] In the formula, is the predicted grasping width at pixel position n, is the actual grasping width.
[0079] ④ Grasping category: Select the multi-class cross-entropy function as the loss function for the category branch calculation, which is defined as
[0080]
[0081] In the formula, M is the number of categories, which is set to three categories. represents the predicted category at pixel position n, and represents its corresponding grasping category label.
[0082] At the same time, the SENet attention mechanism is introduced into each sub-network to improve the quality of feature representation (the SENet attention mechanism plays a key role in each sub-network. By dynamically adjusting the importance of feature channels, it enhances the features related to specific tasks, thereby improving the prediction performance of each task). The formula is as follows:
[0083]
[0084] where W 1 and W 2 are learnable weight matrices, σ is the activation function, and avgpool is the average pooling operation. x refers to the multi-scale fusion feature map.
[0085] An adaptive coefficient α is added before the σ function to enable the model to learn the importance changes of different channels during the training process:
[0086]
[0087] where α is a learnable scalar used to adjust the importance between channels.
[0088] Through this step, we can finally predict the grasping quality, grasping category, grasping angle, and grasping width, that is, our output. The output is as follows:
[0089] In the output of the grasping quality sub-network, the grasping quality score corresponding to each test tube is obtained. The score value ranges from 0 to 100. If it is less than the preset threshold (85), its quality score is set to 0, and the rest are set to 1. The grasping quality score reflects the success probability or stability of the robotic arm in grasping the test tube. The higher the score, the greater the stability or success probability of the grasping. The physical meaning of the grasping quality is the success probability or stability of the grasping task, ensuring that the object is grasped effectively and safely.
[0090] In the output of the grasping category sub-network, the grasping points of two types of test tubes are shown. Category 1 is the pixel points for collision-free grasping, and Category 2 is the grasping points that need to be adjusted. During sorting, it is necessary to determine whether there are suitable grasping points. Here, only Category 1 is considered because grasping is easier and more convenient. At the same time, considering the grasping quality score, we can finally obtain the pixel coordinates (u, v) of the corresponding grasping points. If there is no Category 1, in the way of manual intervention, perform a horizontal jitter on the test tube tray with a set force (which will neither damage the test tubes nor cause them to separate from each other), so that the test tubes are fully separated, and then return to step (1). Obtain the pixel coordinates (u, v) of the final corresponding grasping points from the grasping category sub-network.
[0091] In the output of the grasping angle sub-network, the grasping angle of each test tube is obtained, within the range of [0, 360°).
[0092] In the output of the grasping width sub-network, the grasping width value is obtained. Since the test tube types are the same, the width output values are almost the same.
[0093] By integrating the output information of the 4 sub-networks, we can obtain the grasping configuration information of a single test tube as follows:
[0094] Grasping point: (u = 100, v = 150)
[0095] Grasping quality score: 98
[0096] Grasping category: 1
[0097] Grasping angle: 45°
[0098] Grasping width: 50mm
[0099] In the multi-test tube scenario, a multi-task learning method is adopted, specifically referring to multi-task learning based on region proposals (Region Proposal Networks, RPN). The model simultaneously outputs the grasping configuration of each test tube. Specifically, the model will not only give the position information (coordinates) of each test tube in the image, but also output the corresponding grasping angle, width, quality score, etc. The grasping configuration of each test tube is treated as an independent task, but some weights and feature extraction parts of the model are shared.
[0100] For all test tubes, we need to select and judge their grasping points to select the test tubes to be grasped.
[0101] First, select the grasping points with a grasping category of 1.
[0102] Secondly, judge whether the grasping quality score corresponding to the grasping point exceeds the threshold. If it does not exceed the threshold, the grasping is not considered.
[0103] Then, it is determined whether the grasping angle and width of the grasping point are suitable for grasping.
[0104] Finally, we sort the qualified grasping points according to their grasping quality scores, and preferentially select those with high grasping quality scores. According to the above method, the test tubes to be grasped are determined, and so on.
[0105] ⑻ Execute grasping
[0106] According to the grasping category and quality score output in step (7), first select the test tubes with the grasping category of 1, then select the test tube with the highest quality score among them, and input the grasping angle and grasping width to the controller of the robotic arm; we give priority to the test tube with the highest quality score and convert its pixel coordinates (u, v) into three-dimensional coordinates (X c , Y c , Z c ) in the camera coordinate system. The formula is as follows:
[0107] X c =(u - c x )·Z c / f x (9)
[0108] Y c =(v - c y )·Z c / f y (10)
[0109] Z c = depth value (11)
[0110] Among them, the depth value is directly measured by the camera D435i, and c x and c y are the principal point coordinates of the camera D435i, and f x and f y are the focal lengths of the camera D435i.
[0111] The principal point coordinates and focal lengths are obtained through internal parameter calibration using the Zhang Youzheng algorithm and the camera_calibration function package on the ROS platform. The calibration results are as follows:
[0112]
[0113] Then, the point (X c , Y c , Z c ) in the camera coordinate system is converted into a point (X m , Y m , Z m ) in the robotic arm coordinate system. The formula is as follows:
[0114]
[0115] Among them, R is the rotation matrix obtained through hand-eye calibration, and T is the translation vector obtained through hand-eye calibration. The hand-eye calibration algorithm we used is the classic Tsai-Lenz method. The numerical values of R and T actually obtained finally are as follows:
[0116]
[0117] Finally, the grasping point (X m , Y m , Z m ) is transmitted to the robotic arm controller. The controller plans the motion path of the robotic arm according to the target pose and executes the grasping action.
[0118] Figure 5 For the specific flow chart of the robotic arm sorting test tubes in the actual application scenario, as shown in the figure, the processes represented from left to right in the picture are detecting test tubes, initial pose, action execution, and placing them into the tube slots.
[0119] In the case of multiple test tubes operating simultaneously, this invention patent can also be implemented. We screen and sort according to the mass fraction, give priority to the test tube with the highest mass fraction, and grasp them in sequence until all the test tubes in the picture are grasped.
[0120] In summary, the present invention is based on the test tube grasping technology that combines a robotic arm and a depth camera, and can accurately grasp urine test tubes. Compared with the traditional manual sorting, this method greatly improves the sorting efficiency, reduces the infection risk, and provides technical support for the intelligentization of medical equipment.
Claims
1. A method for automatically sorting urine test tubes based on a robotic arm, characterized in that: include: (1) Collecting test tube images; (2) Using ResNet-101 to extract the preliminary feature map of the test tube image; (3) using multiple parallel-structured residual efficient spatial pyramid modules with different expansion rates and group numbers to process the preliminary extracted feature maps respectively and obtain the corresponding spliced feature maps; (4) Using the feature fusion module to perform multi-scale information fusion on the spliced feature map and the preliminary extracted feature map to generate a large-scale feature map and a small-scale feature map; (5) The obtained large-scale feature map, small-scale feature map and preliminary extracted feature map are respectively concatenated to obtain a captured feature map; (6) The obtained grasping feature map is input into the task sub-network, the grasping parameters of the robotic arm are predicted, and the robotic arm is used to sort the test tubes.
2. The method for automatic urine test tube sorting based on a robotic arm according to claim 1, characterized in that: The test tube image is a test tube RGB image; when training the ResNet-101, the residual efficient spatial pyramid module, the feature fusion module, and the task subnetwork, the test tube RGB image containing a single test tube and the corresponding labels are used for training.
3. The method for automatic urine test tube sorting based on a robotic arm according to claim 1, characterized in that: There are eight residual efficient spatial pyramid modules, including four different expansion rates and two group numbers.
4. The method for automatically sorting urine test tubes based on a robotic arm according to claim 3, characterized in that: The splicing feature map includes a high dilation rate splicing feature map and a low dilation rate splicing feature map; the high dilation rate splicing feature map is extracted and spliced by a residual efficient spatial pyramid module with a high dilation rate, and the low dilation rate splicing feature map is extracted and spliced by a residual efficient spatial pyramid module with a low dilation rate.
5. The method for automatic urine test tube sorting based on a robotic arm according to claim 1, characterized in that: The feature fusion module utilizes multiple convolution and upsampling operations to combine features at different levels to obtain a large-scale feature map and a small-scale feature map, respectively.
6. The method for automatic urine test tube sorting based on a robotic arm according to claim 1, characterized in that: In step (5), the splicing operation is specifically as follows: The small-scale feature map and the preliminary extracted feature map are spliced together, and then spliced together with the large-scale feature map to obtain the captured feature map A; The large-scale feature map and the small-scale feature map are spliced, and then spliced with the preliminary extracted feature map to obtain the captured feature map B; The grasping feature map is composed of the grasping feature map A and the grasping feature map B.
7. The method for automatically sorting urine test tubes based on a robotic arm according to claim 6, characterized in that: The task subnetwork includes four subnetworks, namely, a quality subnetwork, a grasping category subnetwork, a grasping angle subnetwork and a grasping width subnetwork.
8. The method for automatic urine test tube sorting based on a robotic arm according to claim 7, characterized in that: The input of the grasping quality subnetwork and the grasping category subnetwork is the grasping feature map A, and the input of the grasping angle subnetwork and the grasping width subnetwork is the grasping feature map B.
9. The method for automatically sorting urine test tubes based on a robotic arm according to claim 7, characterized in that: The four sub-networks have the same structure, including convolutional layers, deconvolutional layers, concatenation layers, and upsampling layers. At the same time, the attention mechanism SENet is introduced in each sub-network.
10. The method for automatic urine test tube sorting based on a robotic arm according to claim 7, characterized in that: In step (6): (6-1) First, four sub-networks are used to predict the grasping category parameters of each test tube in the test tube image: grasping quality, grasping category, grasping angle, and grasping width; (6-2) Then determine whether there is a test tube with a non-collision grasping point in the grasping category, if yes, proceed to the next step; if not, return to step (1) to re-capture the test tube image; (6-3) The coordinates of the best grasping point with the highest grasping quality are converted into the coordinates of the robot arm, and the grasping angle and grasping width of the robot arm are adjusted to perform test tube sorting.