Automatic injection robot vision data processing method, system and injection robot
The use of deep convolutional neural networks for precise injection site detection on pigs addresses inefficiencies and labor challenges in large-scale vaccination systems, enhancing automation and safety.
Patent Information
- Application Number
- CN202111098333.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-18
AI Technical Summary
The prior art problems in the injection efficiency of liquid agents (vaccines) in large-scale farms, high labor intensity, difficult virus prevention and control, and poor positioning accuracy in injection areas.
The deep convolutional neural network algorithm is used to detect the injectable areas of the pig side and hips. Combined with RFID identity recognition, the deep convolutional neural network model of the side and injectable areas is trained, and high-precision injection is performed by reaching the designated position by the end of the robot arm, and the normal vector is calculated to adjust the posture of the robot arm to ensure that the syringe is perpendicular to the injection position.
Efficient and automated vaccine injections are achieved, which improves injection efficiency, reduces labor intensity, avoids viral infections, and ensures injection accuracy.
Smart Images

Figure CN113971756B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual data processing, and in particular to a visual data processing system and method for an automatic injection robot. Background Art
[0002] At present, there are mainly three types of equipment for injecting liquid medicines, vaccines, etc. into large livestock (including live pigs, cattle, horses, etc.). Taking live pigs as an example, they are as follows:
[0003] 1. Based on a handheld traditional syringe: Manually use a syringe with a needle to inject vaccines into live pigs. Needle injection causes excessive stress reactions in live pigs, there is a risk of cross-infection, the injection efficiency is low, the needle is likely to break in the body of the live pig, generally more than three people are required for injection, and there is also a problem of pollution by medical waste.
[0004] 2. Based on a handheld needleless syringe: Manually use a needleless syringe to inject vaccines into live pigs. A needleless syringe requires auxiliary equipment such as a high-pressure gas cylinder and a pneumatic amplifier, the labor intensity of the injection personnel is high, and manual injection is difficult for virus prevention and control.
[0005] 3. Based on various fixed auxiliary devices: Use a fixed device to fix the live pig, and then manually inject vaccines into the live pig. It is not easy to fix larger live pigs with the auxiliary device, and multiple personnel are required, resulting in low injection efficiency.
[0006] In view of the problems of low efficiency, high labor intensity, and difficulty in virus prevention and control in manual injection of liquid medicines (vaccines) in large-scale farms, it is necessary to use computer vision algorithms to continuously detect the injectable area of livestock (buttocks) at a specified position to control the robotic arm of the robot to inject vaccines into the livestock.
[0007] Currently, the prior art has disclosed a live pig vaccine injection robot including an injection robotic arm (patent application number: 202110907776.1) and a robotic arm motion module (CN211073593U) that can be used for automatic injection of pig vaccines, but it still has the problem of poor positioning accuracy of the injection area. Summary of the Invention
[0008] The purpose of the present invention is to disclose a visual data processing system and method for an automatic injection robot, which is highly efficient, highly accurate, and automated, and solves the problems of low efficiency, high labor intensity, and difficulty in virus prevention and control in the injection of liquid medicines (vaccines) in large-scale farms and similar scenarios.
[0009] Based on the deep convolutional neural network algorithm, the present invention can quickly and accurately detect the injectable areas on the sides and hips of live pigs at fixed positions (such as drinking places), so as to shorten the time for single vaccine injection; the accurately located injection sites and the normal vectors of their local planes can effectively cooperate with the vaccine injection robot to guide the robotic arm to automatically inject vaccines into live pigs, thereby achieving the effects of liberating human labor, improving efficiency, and avoiding virus infection.
[0010] In view of the deficiencies of the prior art, the present invention proposes an automatic injection robot vision data processing method, which includes the following steps:
[0011] Step 1: For the side and hip images of each live pig, manually annotate the back and side images of the pig, use a rectangular box to mark the side position and the tail position, and record the pixel positions of the rectangular box to construct a side detection data set;
[0012] Manually annotate the hip image, and use a rectangular box to mark the injectable area on the hip to construct an injectable area detection data set;
[0013] Step 2: Train a deep convolutional neural network model for detecting the injectable areas in the top view and rear view of the sides and hips of live pigs, including the following steps:
[0014] Use the side detection data set to train the side detection model end-to-end to locate the tail position of the live pig, which is used to guide the end of the robotic arm of the robot to reach the rear of the hip of the live pig. Use the injectable area data set to train the injectable area detection model end-to-end to locate the injection position;
[0015] Step 3: After determining the injection position, fit a plane to a small area centered on the injection point, and calculate the normal vector of this plane to adjust the posture of the robotic arm so that the robot syringe is perpendicular to the injection position for injection.
[0016] In the automatic injection robot vision data processing method, step 1 further includes receiving the ear tag information on the ear of the live pig through an RFID identity recognition receiver, and judging whether the live pig needs to be vaccinated according to the received ear tag information.
[0017] In the automatic injection robot vision data processing method, the deep convolutional neural network model in step 2 can adopt an object detection model of the CenterNet framework algorithm.
[0018] In the automatic injection robot vision data processing method, in the image data set, each image is associated with a 3D coordinate space to obtain the spatial coordinates of the pixel points in the camera coordinate system.
[0019] The described method for processing visual data of an automatic injection robot, where the camera uses the Brown-Conrady distortion model and it is necessary to calculate the spatial coordinates after distortion correction to obtain the spatial coordinates of the pixel points in the camera coordinate system.
[0020] The described method for processing visual data of an automatic injection robot further includes the following steps:
[0021] The camera coordinate system coincides with the syringe tip coordinate system through a rotation transformation R and a translation transformation T. Determine the orientation of the live pig based on the relative positions of the side detection frame and the tail detection frame, and use this to guide the end of the robotic arm to move to a specified position behind the pig's buttocks.
[0022] The camera at the end of the robot's robotic arm captures the buttocks image and the corresponding depth information, and inputs the buttocks image into the detection model to output the detection frame of the injectable area and its corresponding detection result confidence: If the detection algorithm only detects one rectangular frame for the given input image, use the center point of this rectangular frame as the injection point.
[0023] If two rectangular frames are detected and the difference in confidence is greater than 0.1, use the center point of the rectangular frame with a higher confidence as the injection point.
[0024] If two rectangular frames are detected and the difference in confidence is less than 0.1, use the center point of the rectangular frame with a larger area within the rectangular frames as the injection point.
[0025] Calculate the spatial coordinates of this point in the camera coordinate system based on the pixel coordinates and the corresponding depth value at the injection point in the image, and convert this spatial coordinate to the syringe tip coordinate system to obtain the injection position.
[0026] The described method for processing visual data of an automatic injection robot, where step 3 further includes using the singular value decomposition method or the eigenvector method to obtain the normal vector of the fitting plane.
[0027] The present invention also proposes a system for processing visual data of an automatic injection robot, which includes:
[0028] A module for establishing a detection image data set, which is used to perform manual annotation on the side and buttocks images of each live pig, mark the side position and the tail position with a rectangular frame, and record the pixel positions of the rectangular frame to construct a side detection data set; it is used to perform manual annotation on the buttocks images, and mark the injectable area of the buttocks with a rectangular frame to construct an injectable area detection data set.
[0029] Train the deep convolutional neural network model module for detecting the depth of the injectable area in the top view and rear view of the side and hip of live pigs, use the side detection dataset to train the side detection model end-to-end to locate the position of the pig's tail, guide the end of the robotic arm of robot 15 to reach behind the hip of the live pig, and use the injectable area dataset to train the injectable area detection model end-to-end to locate the injection position.
[0030] The normal vector calculation module is used to determine a small area centered on the injection point to fit a plane and calculate the normal vector of the plane after determining the injectable position, so as to adjust the posture of the robotic arm to make the robot syringe perpendicular to the injection position for injection.
[0031] The automatic injection robot vision data processing system described above further includes: a data output module for outputting the normal vector to the automatic injection robot.
[0032] The present invention also proposes an automatic injection robot, which includes the automatic injection robot vision data processing system. Description of the Drawings
[0033] Figure 1 Schematic diagram of the working environment of the vaccine injection robot;
[0034] Figure 2 Top view of the back, side and tail regions of the live pig (the rectangle is the annotation box);
[0035] Figure 3 Schematic diagram of the injectable area of the hip (the rectangle is the marked injectable area);
[0036] Figure 4 Schematic diagram of the detection model CenterNet network. Detailed Embodiments
[0037] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and are described in detail in conjunction with the accompanying drawings of the specification as follows.
[0038] The present invention will be described in detail below (taking live pigs as an example) with reference to the accompanying drawings.
[0039] The present invention discloses a robot vision data processing system and method, that is, based on a deep convolutional neural network for detecting the injectable areas of the side and hip of live pigs, including the following steps:
[0040] 1. Install the hardware in the pigsty and establish a detection image dataset.
[0041] The working environment of the present invention is as Figure 1As shown in the schematic diagram, the pigsty 11 is designed with two (or multiple) rows of pig pens at the same interval, with an aisle 12 in the middle. A designated drinking position 13 is designed near the aisle 12. The robot 15 moves in the aisle 12. When the live pig 14 drinks water at the designated drinking position 13, it is injected through the cooperation of the RGB image acquisition device of the robot 15 and the robotic arm motion injection robot 15. At the designated drinking position 13, an RFID identity recognition receiver is installed, and each pig wears an electronic ear tag, so the identity of each live pig entering the drinking area can be effectively recognized. The main function of the RFID identity recognition receiver is to receive the ear tag information on the pig's ear and determine whether the live pig needs to be vaccinated based on the received ear tag information. The RGB image acquisition device is mainly used to collect pig images.
[0042] To guide the robot to perform effective injection, it is necessary to know the position of the live pig and determine the injection area. The present invention adopts a visual positioning strategy. To train the detection model, a detection data set is first established. In the real environment of the pigsty 11, the RGB image acquisition device is held to collect the side and hip images of each live pig. To obtain more robust detection results, the images of the live pig standing can be collected from multiple different angles, such as top view, side view, and rear view. To obtain richer samples and make the trained network more robust, the angles when collecting the top view and rear view are not fixed. Here, it is not necessary to be strictly the top view and rear view. Since the live pig stays relatively long during the drinking activity, it is only necessary to collect images at the drinking area. Due to the existence of the fence at the aisle and the height of the robot body being about 70 - 80 cm, a strictly side view cannot be collected. However, as mentioned before, the top view and other perspectives are not very strict, and the top view and side view angles can be combined to a certain extent. The image acquisition device on the robot 15 collects the side image of the live pig through the top view + test perspective for rough positioning of the tail.
[0043] Manually annotate the back and side images of the pig, use the rectangular box 24 to mark the side position and the tail position, and record the pixel position of the rectangular box 24 to construct the side detection data set; see Figure 2 , which is the top view of the back, side, and tail areas of the live pig (the rectangle is the annotation box), where 21 represents the live pig, 22 represents the pig's head, 23 represents the tail, 24 is the annotation rectangle of the live pig, and 25 is the annotation rectangle of the tail. Manually annotate the hip image, use the rectangular box to mark the injectable area of the hip to construct the injectable area detection data set. See Figure 3 , 31 represents the hip of the live pig; 32 and 33 represent the marked injectable areas.
[0044] 2. Train and construct a deep convolutional neural network model for detecting the injectable areas in the top view and rear view of the side and hip of the live pig.
[0045] The present invention detects an object in an image and identifies it with a rectangular box, realizing the side detection task of the top view of live pigs. It can be achieved by using an object detection model, such as the CenterNet framework algorithm:
[0046] The target is described as a center point, and other characteristics of the target, such as size and orientation, are directly regressed in the feature map. CenterNet mainly includes three parts. First, the input image is scaled to a size of 512×512, that is, the long side is scaled to 512 and the short side is padded with 0. Then the image is input into the backbone network for feature extraction. Finally, the features output by the backbone network are predicted using a prediction module, which includes 3 branches, specifically including the center point heatmap branch, the target size branch, and the center point offset branch. The heatmap branch contains C channels, and each channel contains a category. The local maximum value in the heatmap represents the center point position of the target. The target size branch is used to predict the deviation values of the width w and height h of the target rectangular box, and the center point offset branch is used to compensate for the pixel error caused by mapping the points on the pooled low heatmap to the original image. This method has a simple principle, strong compatibility, and does not require complex post-processing, achieving true end-to-end.
[0047] Use the side detection dataset to train the side detection model end-to-end to locate the position of the pig's tail, and use it to guide the end of the robotic arm of the robot 15 to reach behind the pig's hip. Use the injectable area dataset to train the injectable area detection model end-to-end to locate the injection position.
[0048] Among them, it includes the training of the side detection model and the training of the injectable area detection model. Use the side detection dataset to train the side detection model Ms end-to-end, that is, input the side image and the corresponding annotation information, and the side detection model Ms outputs the detection boxes 34 of the pig's side and tail. For the tail detection box 35, take the pixel coordinates at the center point and the average value of the depth values of each point after removing the background as the depth information (in practical applications, the RGB-D depth camera of the robot 15 can be used to obtain the RGB image and the corresponding depth information at the same time), and calculate the spatial coordinate p1 of this point. The calculation process is as follows:
[0049] For an image captured by a camera, in terms of pixels, the coordinate [0, 0] refers to the center of the top-left pixel of the image, and [w - 1, h - 1] refers to the center of the bottom-right pixel in an image that contains exactly w columns and h rows (where w and h represent the width and height of the image respectively). That is, from the perspective of the camera, the x-axis points to the right, the y-axis points downwards, and the coordinates in this space are called "pixel coordinates", which are used to index the image to find the content of a specific pixel. Each image is also associated with a 3D coordinate space (camera coordinate system), in meters, and the coordinate [0, 0, 0] refers to the center of the RGB-D depth camera. In this space, the positive x-axis points to the right, the positive y-axis points downwards, and the positive z-axis points forward. The coordinates in this space are called "points", which are used to describe the positions visible in a specific image in 3D space. For a point in the image, such as (x c1 , y c1 ), if its depth value depth and the internal parameters of the camera (focal lengths (f x , f y ), optical center (o x , o y ), and 5 distortion parameters (d1, d2, d3, d4, d5)) are known, then the spatial coordinates of the pixel point (x c1 , y c1 ) in the camera coordinate system can be calculated through the following formula:
[0050] x = (x c1 - o x ) / f x , y = (y c1 - o y ) / f y
[0051] If the camera is in the Brown-Conrady distortion model, the spatial coordinates after distortion correction need to be calculated. The distortion correction process is as follows:
[0052] r2 = x 2 + y 2
[0053] f = 1 + d1 * r2 + d2 * r2 * r2 + d5 * r2 * r2 * r2
[0054] x = x * f + 2 * d3 * x * y + d4 * (r2 + 2x 2 )
[0055] y = y * f + 2 * d4 * x * y + d3 * (r2 + 2y 2 )
[0056] Finally, the pixel point (x c1 , y c1)The spatial coordinate p1(x, y, depth) in the camera coordinate system.
[0057] Since the positions of the camera and the end of the syringe are relatively fixed, the camera coordinate system can coincide with the end coordinate system of the syringe through a rotation transformation R and a translation transformation T. That is, there is a one-to-one correspondence between the camera coordinate system and the end coordinate system of the syringe. Then, p1 can be converted to P1 through P1 = p1R + T.
[0058] Finally, determine the orientation of the live pig according to the relative positions of the side detection frame and the tail detection frame, and use this to guide the end of the robotic arm to move to about 30 cm behind the buttocks of the live pig.
[0059] Similarly, use the object detection model to implement the training of the injectable area detection model Mb. In practical applications, after the robotic arm moves to the specified position behind the buttocks of the live pig, the RGB-D depth camera at the end of the robotic arm of the robot 15 collects the buttocks image and the corresponding depth information, and inputs the buttocks image into the detection model Mb to output the detection frame of the injectable area and the corresponding detection result confidence. The confidence represents the reliability of the detection result, which can be given by the deep neural network, and the value range is (0-1). The higher the value, the more reliable it is.
[0060] If the detection algorithm only detects one rectangular frame 42 for the given input image, use the center point of this rectangular frame as the injection point;
[0061] If two rectangular frames 42 and 43 are detected and the difference in confidence is greater than 0.1, use the center point of the rectangular frame with a higher confidence as the injection point;
[0062] If two rectangular frames 42 and 43 are detected and the difference in confidence is less than 0.1, use the center point of the rectangular frame with a larger area inside the rectangular frame as the injection point.
[0063] Calculate according to the pixel coordinates and the corresponding depth value at the injection point in the image to obtain the spatial coordinate p2 of this point in the camera coordinate system. Similar to the calculation process of the spatial coordinate p1, calculate the spatial coordinate p2 in the camera coordinate system according to the pixel coordinates, the corresponding depth value and the internal parameters of the camera at the injection point in the image, and convert p2 to the end coordinate system of the syringe to obtain the injection position P2. The conversion of p2 to P2 is the same as the conversion of p1 to P1. Since the positions of the camera and the end of the syringe are relatively fixed, the camera coordinate system can coincide with the end coordinate system of the syringe through a rotation transformation R and a translation transformation T. That is, there is a one-to-one correspondence between the camera coordinate system and the end coordinate system of the syringe. Then, p2 can be converted to P2 through P2 = p2R + T.
[0064] Another embodiment of the present invention:
[0065] In the embodiment of the object detection network of the present invention, an improved CenterNet network (such as Figure 4 ) is used to detect the side body and injectable area of live pigs. In the figure, S41 is the input image, S42 - S46 are Conv1, Conv2_x, Conv3_x, Conv4_x, Conv5_x in Resnet50 respectively, S47 is the MaxPool pooling operation, S48 - S411 are deconvolution operations, S412 represents the prediction of the target center point, S413 represents the prediction of the target size (width and height), and S414 represents the prediction of the target center point offset.
[0066] The CenterNet network scales the input image S41 to a fixed size (512×512 pixels). The backbone network for feature extraction uses Resnet50 for feature extraction. Resnet50 represents a network structure of the Resnet network with 50 operation layers. Resnet proposed network structures of 18, 34, 50, 101, and 152 layers, which are respectively represented as Resnet18, Resnet34, Resnet50, Resnet101, and Resnet152. The network structure of Resnet50 is shown in Table 1, mainly including 5 modules Conv1, Conv2_x, Conv3_x, Conv4_x, and Conv5_x corresponding to S42 - S46 respectively. After each module, a MaxPool pooling operation S47 is used to halve the feature size. Among them, 1×1 and 3×3 are the sizes of the convolution kernels, 64 and 256 are the numbers of feature channels, and the subsequent ×3 means {…} is executed 3 times. After the feature passes through three deconvolutions (S48, S49, and S410), the feature size changes from 16×16 to 128×128. After the feature passes through the deconvolution, it is sent to three branches for prediction respectively. Here, the three branches respectively convolve the features extracted by the network to the corresponding sizes for predicting the center point of the target (i.e., heatmap prediction), the width and height of the target, and the offset of the target center point. For the heatmap S412 (center point) prediction, the size of the side body detection model is 128×128×2 (two categories: side body and tail). The size of the injectable area detection model is 128×128×1 (only one category: injectable area). S413 represents the prediction of the width and height of the detected target. Since the final size of the feature is inconsistent with the size of the input image, there is an offset when the center point is mapped to the image pixels. S414 represents the prediction of the center point offset. The role of the loss function is to measure the quality of the model prediction. During model training, the loss is reduced by the gradient descent method to achieve the purpose of optimizing the model. The loss function is used as an evaluation of the prediction result. The smaller its value, the more accurate the predicted target position and the more accurate the injection position. The loss function of the CenterNet network can be expressed as L det = L k+λ size L size +λ off L off (λ size = 0.1, λ off = 1), where L represents loss, and L det represents the loss of the model, and L det is divided into three parts, corresponding to the three branches of the network respectively: L k represents the loss of predicting the center point of the target, and L size represents the loss of predicting the width and height of the target, and L off represents the loss of the offset of the center point of the predicted target, λ size and λ off are hyperparameters used to control the influence of the losses of the corresponding branches.
[0067] In the CenterNet network, due to the inconsistency between the input image and the network output feature size, it is necessary to predict the offset of the center point of the target, which not only increases the network complexity but also reduces the accuracy of network detection. For the above problems, the present invention improves the CenterNet network. First, the input image S41 is scaled to 256×256 size, and the last MaxPool pooling operation of Resnet50 is removed. Then the feature size after passing through Resnet50 is 16×16. After three deconvolutions (S48, S49, and S410), one more deconvolution (S411) is performed. At this time, the feature size is 256×256, which is the same as the size of the input image S41. The predicted center point position is the same as the corresponding pixel position of the input image, and the prediction of the center point offset is no longer required. The improved CenterNet network can reduce the network parameters and make the network easier to fit.
[0068]
[0069]
[0070] Table 1 Resnet50 network structure
[0071] In the present invention, other Resnets and their variants, DLA, etc. can also be used as the backbone network to obtain other possible embodiments.
[0072] The residual module of the Resnet network effectively solves the degradation problem in deep learning. Its variants add attention mechanisms, etc. on the basis of the Resnet network to further improve the network performance; the DLA network combines different layers in the network richly, effectively enhancing the feature expression ability.
[0073] In the present invention, for the training of the detection model, other mature object detection frameworks can also be used, such as the YOLO series, the R-CNN series, etc.
[0074] The YOLO algorithm treats the object detection problem as a regression problem, and a convolutional neural network structure can directly predict the bounding box and class probability from the input image. RCNN first uses the selective search method to extract about 2,000 candidate regions on the image, scales the candidate regions to a fixed size and then inputs them into the CNN network, and finally inputs the output of the CNN into the SVM for class determination.
[0075] 3. After determining the position to be injected, fit a plane to a small area centered on the injection point p2, calculate the normal vector of this plane to adjust the posture of the robotic arm, so that the robotic syringe injects perpendicularly to the injection position. The calculation process is as follows:
[0076] After detecting the injection position p2 in the image, a small area centered on the injection point (a square p(x c +i,y c +j), where (x c ,y c ) is the coordinate of the injection point, -10 ≤ i < 10, -10 ≤ j < 10, and obtain the spatial coordinates P(x k ,y k ,z k ) of all pixel points in this area according to the depth information of all points in this small area, where 0 ≤ k < 400, fit a plane according to the spatial coordinates of these points and calculate the normal vector of this plane. The matrix formed by the spatial coordinates of these points is:
[0077]
[0078] The present invention uses two common matrix decomposition methods in linear algebra to find the normal vector of the fitted plane:
[0079] The singular value decomposition method uses formula 1 to calculate the singular value decomposition of the centroid-removed spatial coordinates, and the rightmost column of the left singular vector u(3×3) is the normal vector of the fitted plane.
[0080] In the eigenvector method, for all points in the region, there is equation 2. For multiple points (more than 3), the left pseudo-inverse matrix (formula 3) is used, and the normal vector of the fitted plane is calculated according to formula 4. Adjust the posture of the robotic arm according to this normal vector so that the syringe is perpendicular to the plane fitted to the local area around the injection point.
[0081] In other words, in the camera coordinate system, the normal vector of the plane fitting the area around the injection point (temporarily set as n, for example) is obtained, and then according to N = nR + T, n is transformed into the normal vector N in the coordinate system of the syringe tip. N guides the manipulator to adjust its posture so that the syringe is perpendicular to the plane fitting the local area around the injection point.
[0082] u, s, vh = SVD(P T - mean(P T )) (1)
[0083]
[0084] A + =(A T A) -1 A T (3)
[0085]
[0086] The vaccine injection robot can replace manual labor in the vaccine injection task with only simple modification of the pigsty, greatly reducing the labor cost. Using the auxiliary vision system of the present invention can be divided into two stages: live pig body detection and detection of the area to be injected on the buttocks. By using a deep convolutional neural network to train the detection model, the injection position can be quickly and accurately located, and it is ensured that the injection is perpendicular to the injection position, improving the injection efficiency and enhancing the safety of live pigs at the same time. If detection data sets of different categories of large livestock and small animals can be collected, the present invention can be conveniently extended to the injection tasks of other livestock vaccines and liquid medicines.
[0087] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0088] The present invention also proposes an automatic injection robot vision data processing system, which includes:
[0089] A module for establishing a detection image data set, which is used to manually annotate the side and buttocks images of each live pig, mark the side position and tail position with a rectangular box, and record the pixel positions of the rectangular box to construct a side detection data set; and is used to manually annotate the buttocks images and mark the injectable area on the buttocks with a rectangular box to construct an injectable area detection data set.
[0090] A deep convolutional neural network model module for training the depth of the injectable area in the top view and rear view of the side and hip of live pigs, which is used to train the side detection model end-to-end using the side detection dataset, locate the position of the pig's tail, and guide the end of the robotic arm of the robot 15 to reach the rear of the pig's hip. The injectable area detection model to be trained end-to-end using the injectable area dataset to locate the injection position.
[0091] A normal vector calculation module, after determining the position to be injected, fits a plane to a small area centered on the injection point, calculates the normal vector of the plane, and adjusts the posture of the robotic arm so that the robot syringe is perpendicular to the injection position for injection.
[0092] The automatic injection robot vision data processing system described above further includes: a data output module for outputting the normal vector to the automatic injection robot.
[0093] The present invention also proposes an automatic injection robot, which includes the automatic injection robot vision data processing system.
Claims
1. An automatic injection robot vision data processing system, characterized in that, Including: A dataset construction module, which is used to label the pig images with the side position and the tail position, and construct a side detection dataset; Based on the pig images labeled with the injectable area of the hip, construct an injection area detection dataset; A model training module, which is used to end-to-end train a side detection model with the side detection dataset to locate the tail position of the pig, so as to guide the end of the robotic arm of the injection robot to reach the rear of the pig's hip, and use the injectable area dataset to end-to-end train the to-be-injected area detection model to locate the injection position; wherein, in this dataset, each image is associated with a 3D coordinate space to obtain the spatial coordinates of the pixel points in the camera coordinate system; the camera coordinate system coincides with the syringe end coordinate system through a rotation transformation R and a translation transformation T, and the orientation of the pig is determined according to the relative positions of the side detection frame and the tail detection frame, so as to guide the end of the robotic arm to move to a specified position behind the pig's hip; the camera at the end of the robotic arm of the robot collects the hip image and the corresponding depth information, inputs the hip image into the detection model, and outputs the detection frame of the injectable area and the corresponding detection result confidence: if the detection algorithm only detects one rectangular frame for the given input image, the center point of this rectangular frame is used as the injection point; if two rectangular frames are detected and the confidence difference is greater than 0.1, the center point of the rectangular frame with a greater confidence is used as the injection point; if two rectangular frames are detected and the confidence difference is less than 0.1, the center point of the rectangular frame with a larger area inside the rectangular frame is used as the injection point; the spatial coordinates of this point in the camera coordinate system are calculated according to the pixel coordinates and the corresponding depth value at the injection point in the image, and the spatial coordinates are converted to the syringe end coordinate system to obtain the injection position; A robotic arm adjustment module, which is used to fit the injection position area into a plane, calculate the normal vector of this plane, and adjust the posture of the robotic arm according to this normal vector, so that the syringe of the injection robot is perpendicular to the injection position for injection.
2. The visual data processing system of the automatic injection robot according to claim 1, wherein This robotic arm adjustment module includes: a data output module, which is used to output the normal vector to the automatic injection robot.
3. An injection robot including the automatic injection robot vision data processing system as described in claim 1 or 2.
Citation Information
Patent Citations
Live pig vaccine injection robot and vaccine injection method
CN113599013A
Mechanical arm action module capable of being used for automatically injecting pig vaccines
CN211073593U
Surgical robot system for radioactive seed interventional therapy of tumors and operating method
CN105616004A
Automatic micro-injection device and control method thereof
CN111558110A