Pet size intelligent measurement system and method based on multi-task segmentation network
Through the combination of binocular stereo vision and multi-task segmentation network, the problems of large measurement errors, stress response and inaccurate edge processing of traditional pet intelligent hair combing equipment are solved, and high-precision and fast pet size measurement is achieved, providing key technical support for pet intelligent hair combing.
Patent Information
- Application Number
- CN202510551358.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional pet intelligent combing equipment has problems such as large measurement errors, contact measurement causes pet stress response, and segmentation algorithms are inaccurate in handling dynamic edges.
The pet size intelligent measurement system based on a multi-task segmentation network is adopted, and depth information is obtained using binocular stereoscopic vision technology, lens distortion is corrected through binocular stereoscopic calibration module, binocular stereoscopic matching module calculates parallax and restores depth, and the object detection and segmentation module performs high-precision pet positioning and segmentation, and finally three-dimensional measurement is performed.
It realizes contactless and high-precision pet size measurement, with the measurement accuracy being improved to <1%, the segmentation accuracy reaches 99%, and the processing speed is fast. It is suitable for pet intelligent hair combing equipment.
Smart Images

Figure CN120495376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intersection of computer vision, deep learning and automatic control, and specifically to three-dimensional visual perception based on binocular stereo vision, target detection driven by multi-task neural networks, instance segmentation and three-dimensional measurement technology, and more particularly to an intelligent pet size measurement system and method based on a multi-task segmentation network and three-dimensional vision. Background Art
[0002] As people's living standards improve, more and more people keep pets. Pets can not only make people's lives healthier and more fulfilling, but also bring them happiness and cultivate people's sense of responsibility. Therefore, they have become an important part of life.
[0003] When people raise pets, they generally use pet smart care devices, among which pet smart grooming devices are more common.
[0004] Existing smart pet grooming devices generally have the following technical defects:
[0005] 1. Traditional monocular vision systems cannot obtain depth information, resulting in large errors in pet size measurement;
[0006] 2. The measurement method based on contact sensors can easily cause stress reactions in pets;
[0007] 3. Existing segmentation algorithms do not accurately handle the edges of dynamic targets (especially the edges of long-haired pet hair). Summary of the Invention
[0008] In view of this, in order to solve the above-mentioned problems in the prior art, the present invention proposes a pet size intelligent measurement system and method based on a multi-task segmentation network, which uses non-contact to accurately measure the pet's body shape and accurately segment the hair distribution area, providing key technology for subsequent pet intelligent grooming.
[0009] The present invention solves the above problems through the following technical means:
[0010] In one aspect, the present invention provides a pet size intelligent measurement system based on a multi-task segmentation network, comprising:
[0011] The binocular stereo calibration module is used to determine the geometric relationship between the two cameras. This includes: calculating the camera intrinsic parameters, calculating the inherent parameters of a single camera, and describing the camera's imaging characteristics; calculating the camera extrinsic parameters, calculating the relative position and orientation between the two cameras and defining their spatial relationship; eliminating lens distortion, correcting radial and tangential distortion of the lens; establishing epipolar constraints, and obtaining the fundamental matrix or essential matrix through calibration, which compresses the stereo matching search range from the two-dimensional image to the one-dimensional epipolar line;
[0012] The binocular stereo matching module is used to calculate the disparity between pixels and obtain depth information by analyzing image pairs of the same scene captured by left and right cameras. This includes: generating a disparity map by matching corresponding pixels in the left and right images, calculating the horizontal displacement of each pixel, and outputting a disparity map in which the brightness value represents the magnitude of the disparity; recovering depth information by converting the disparity map into a depth map based on the principle of triangulation and combined with calibrated camera parameters, which directly reflects the distance between each point in the scene; processing occluded areas, detecting areas that are not visible in the left and right perspectives, and filling in reasonable disparity values through interpolation or global optimization algorithms; and suppressing noise to eliminate mismatches during the matching process and improve the robustness of the depth map.
[0013] The object detection and segmentation module uses a multi-task segmentation network to perform object detection, object classification, and pixel-level instance segmentation on pets, improving the high-precision positioning and recognition of pet objects. This module includes: object detection, which locates the position of the pet in the image or video and marks it with a bounding box; object classification, which determines whether the detected pet belongs to a predefined category to avoid false detection; and instance segmentation, which performs pixel-level segmentation on the pet's outline to obtain the pet's outline information.
[0014] The 3D measurement module is used to measure the pet's size based on the depth information obtained by the binocular stereo matching module and the pet's outline information obtained by the target detection and segmentation module.
[0015] Preferably, the binocular stereo calibration module includes:
[0016] A data acquisition unit is used to use a binocular camera to shoot the same checkerboard from different angles, ensuring that the checkerboard is completely visible in both the left and right images and covers the edge and center areas of the image;
[0017] The single-target positioning unit is used to detect corner points. A corner detection algorithm is used to extract checkerboard corner points for each image. A sub-pixel optimization algorithm is used to improve the corner point accuracy to a certain pixel level. The unit also calculates parameters. Based on the Zhang Zhengyou calibration method, a camera model is established to solve the intrinsic parameter matrix and distortion coefficients, and then the reprojection error is minimized to optimize the parameters.
[0018] The dual-target positioning unit is used for stereo corner matching, matching the corner points of the left and right images to ensure the correspondence between the same physical point in the two images. It also calculates the extrinsic parameters, solves the essential matrix through epipolar geometry, decomposes it into the rotation matrix and translation vector, and uses the bundle adjustment method to jointly optimize all parameters to minimize the 3D point reprojection error.
[0019] The stereo correction unit is used to calculate the correction transformation matrix based on the calibration results, align the epipolar lines of the left and right images horizontally, generate a correction mapping table, and correct the new image in real time.
[0020] Preferably, the binocular stereo matching module includes:
[0021] The binocular data acquisition unit is used to collect video images from the left and right cameras at the same time. If there is a time difference, it will cause errors in the subsequent matching and thus result errors;
[0022] The disparity map generation unit is used to correct the input images of the left and right cameras based on the mapping table obtained by the binocular stereo calibration module so that the images of the left and right cameras are mapped to the same horizontal plane. The SGBM algorithm is used with the parameters set to a disparity range of 64 pixels and a penalty coefficient of P1 = 8 / P2 = 32, and the output is a 16-bit disparity map.
[0023] A disparity map filtering unit is used to filter the disparity map using the WLS algorithm to minimize disparity noise while protecting image edges;
[0024] The depth map conversion unit is used to convert the disparity map filtered by the left camera into 3D space using the reprojection matrix Q to obtain the world coordinate system of the depth map and each pixel of the image.
[0025] Preferably, the target detection and segmentation module includes:
[0026] The model training unit is used to collect a large amount of pet data and design a multi-task network. The network output includes object detection, instance segmentation, and object recognition;
[0027] The model quantization unit is used to quantize the trained model into a binary file that can be run on the embedded system through the quantization tool;
[0028] The model inference unit is used to obtain images from the left or right camera and normalize the size to the model input size; the model output targets are the pet position coordinate information, pet classification information, credibility, and pet mask in the image.
[0029] Preferably, the three-dimensional measurement module includes:
[0030] The contour acquisition unit is used to obtain the contour information of the pet in the image based on the mask obtained by model inference;
[0031] A contour calculation unit, used to calculate the centroid point of the contour;
[0032] The shortest distance point pair calculation unit calculates the intersection point pairs (P1, P2) of multiple straight lines passing through the contour centroid and the contour;
[0033] The valid point pair retention unit is used to search in the 3D image based on the shortest distance point pair, exclude points with invalid or out-of-range depth values, and retain valid point pairs P1(x1,y1,z1), P2(x2,y2,z2), ...;
[0034] Pet size measurement unit for use according to the formula Calculate the distances of multiple point pairs and calculate the average value avg_dist of these point pairs to obtain the average cross-sectional width of the pet.
[0035] On the other hand, the present invention provides a pet size intelligent measurement method based on a multi-task segmentation network, comprising the following steps:
[0036] Binocular stereo calibration determines the geometric relationship between the two cameras. This includes: calculating camera intrinsic parameters, calculating the inherent parameters of a single camera, and describing the camera's imaging characteristics; calculating camera extrinsic parameters, calculating the relative position and orientation between the two cameras and defining their spatial relationship; eliminating lens distortion, correcting radial and tangential distortion of the lens; establishing epipolar constraints, and obtaining the fundamental matrix or essential matrix through calibration, which compresses the stereo matching search range from the two-dimensional image to the one-dimensional epipolar line.
[0037] Binocular stereo matching analyzes image pairs of the same scene captured by left and right cameras, calculates the disparity between pixels, and thus obtains depth information. This includes: generating a disparity map by matching corresponding pixels in the left and right images, calculating the horizontal displacement of each pixel, and outputting a disparity map where the brightness value represents the magnitude of the disparity; restoring depth information by converting the disparity map into a depth map based on the principle of triangulation and combined with calibrated camera parameters, which directly reflects the distance between each point in the scene; processing occluded areas, detecting areas that are not visible in the left and right perspectives, and filling in reasonable disparity values through interpolation or global optimization algorithms; and suppressing noise to eliminate mismatches during the matching process and improve the robustness of the depth map.
[0038] Object detection and segmentation uses a multi-task segmentation network to perform object detection, object classification, and pixel-level instance segmentation on pets, improving the high-precision positioning and recognition of pet objects. This includes: object detection, locating the position of the pet in the image or video and marking it with a bounding box; object classification, determining whether the detected pet belongs to a predefined category to avoid false detection; instance segmentation, performing pixel-level segmentation on the pet's outline to obtain the pet's outline information;
[0039] Three-dimensional measurement is based on the depth information obtained from binocular stereo matching and the pet outline information obtained from target detection and segmentation, and the two are combined to measure the pet's size.
[0040] Preferably, the binocular stereo calibration includes the following steps:
[0041] Data collection: Use a binocular camera to capture the same checkerboard from different angles, ensuring that the checkerboard is completely visible in both the left and right images, and covers the image edges and center areas;
[0042] Single-target localization: Detect corners. A corner detection algorithm is used to extract checkerboard corners from each image. A sub-pixel optimization algorithm is used to improve corner accuracy to a certain pixel level. Parameters are calculated. Based on the Zhang Zhengyou calibration method, a camera model is established to solve for the intrinsic parameter matrix and distortion coefficients. The reprojection error is then minimized to optimize the parameters.
[0043] Dual-target positioning: Stereo matching corner points, matching the corner points of the left and right images to ensure the correspondence between the same physical point in the two images; calculating extrinsic parameters, solving the essential matrix through epipolar geometry, decomposing it to obtain the rotation matrix and translation vector, and using the bundle adjustment method to jointly optimize all parameters to minimize the 3D point reprojection error;
[0044] Stereo correction: Based on the calibration results, the correction transformation matrix is calculated to align the epipolar lines of the left and right images horizontally, and a correction mapping table is generated to correct the new image in real time.
[0045] Preferably, the binocular stereo matching includes:
[0046] Binocular data acquisition: The video images of the left and right cameras are collected at the same time. If there is a time difference, it will cause errors in the subsequent matching and thus lead to error in the results;
[0047] Disparity map generation: Based on the mapping table obtained from the binocular stereo calibration module, the input images of the left and right cameras are corrected so that the images of the left and right cameras are mapped to the same horizontal plane. The SGBM algorithm is used with the parameters set to a disparity range of 64 pixels and a penalty coefficient of P1 = 8 / P2 = 32, and a 16-bit disparity map is output.
[0048] Disparity map filtering: used to filter the disparity map using the WLS algorithm to minimize disparity noise while protecting image edges;
[0049] Depth map conversion: Use the reprojection matrix Q to convert the disparity map filtered by the left camera into 3D space to obtain the world coordinate system of the depth map and each pixel of the image.
[0050] Preferably, the target detection and segmentation includes:
[0051] Model training: Collect a large amount of pet data and design a multi-task network. The network output includes object detection, instance segmentation, and object recognition.
[0052] Model quantization: The trained model is quantized into a binary file that can be run on an embedded system using a quantization tool.
[0053] Model inference: Obtain images from the left or right camera and normalize the size to the model input size; the model output targets are the pet location coordinate information, pet classification information, credibility, and pet mask in the image.
[0054] Preferably, the three-dimensional measurement includes:
[0055] Contour acquisition: Obtain the contour information of the pet in the image based on the mask obtained by model inference;
[0056] Contour calculation: used to calculate the centroid of the contour;
[0057] Shortest distance point pair calculation: Calculate the intersection point pairs (P1, P2) of multiple straight lines passing through the contour centroid and the contour;
[0058] Retain valid point pairs: Search in the 3D graph based on the shortest distance point pairs, exclude points with invalid or out-of-range depth values, and retain valid point pairs P1(x1,y1,z1), P2(x2,y2,z2), ...;
[0059] Pet size measurement: According to the formula Calculate the distances of multiple point pairs and calculate the average value avg_dist of these point pairs to obtain the average cross-sectional width of the pet.
[0060] Compared with the prior art, the beneficial effects of the present invention include at least:
[0061] The present invention uses non-contact methods to accurately measure the pet's body shape and accurately segment the hair distribution area, improving measurement accuracy with a relative error of <1%. The segmentation accuracy is high, with mAP@0.5 reaching 99%. The processing speed is fast, reaching 100ms / frame, providing a key technology for subsequent intelligent pet grooming. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0063] Figure 1 is a structural diagram of the pet size intelligent measurement system based on the multi-task segmentation network of the present invention;
[0064] Figure 2 It is a structural diagram of the binocular stereo calibration module of the present invention;
[0065] Figure 3 It is a structural diagram of the binocular stereo matching module of the present invention;
[0066] Figure 4 It is a structural diagram of the target detection and segmentation module of the present invention;
[0067] Figure 5 is a structural diagram of the three-dimensional measurement module of the present invention;
[0068] Figure 6 is a flow chart of the pet size intelligent measurement method based on the multi-task segmentation network of the present invention;
[0069] Figure 7 Schematic diagram of the principle of binocular stereo calibration of the present invention;
[0070] Figure 8 This is a flow chart of binocular stereo matching of the present invention;
[0071] Figure 9 It is a flow chart of target detection and segmentation of the present invention;
[0072] Figure 10 It is a structural block diagram of the electronic equipment of the present invention. DETAILED DESCRIPTION
[0073] To make the above-mentioned objectives, features, and advantages of the present invention more clearly understood, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are also within the scope of protection of the present invention.
[0074] Example 1
[0075] like Figure 1 As shown, the present invention provides a pet size intelligent measurement system based on a multi-task segmentation network, including a binocular stereo calibration module, a binocular stereo matching module, a target detection and segmentation module and a three-dimensional measurement module.
[0076] Binocular stereo calibration module: Binocular stereo calibration is a key step in computer vision. It is mainly used to determine the geometric relationship between two cameras (binocular system), providing a basis for tasks such as 3D reconstruction, depth perception, and dimension measurement. Its functions are as follows:
[0077] 1. Calculate the internal and external parameters of the camera:
[0078] Camera intrinsic parameters: The intrinsic parameters of a single camera (such as focal length, principal point coordinates, and distortion coefficients) that describe the imaging characteristics of the camera.
[0079] Camera extrinsics: The relative position and orientation (rotation matrix and translation vector) between the two cameras, defining their spatial relationship.
[0080] 2. Eliminate lens distortion: Correct the radial distortion (such as barrel distortion and pincushion distortion) and tangential distortion of the lens to ensure that the image is closer to the ideal projection.
[0081] 3. Establish epipolar constraints: Obtain the fundamental matrix or essential matrix through calibration, compress the stereo matching search range from the two-dimensional image to the one-dimensional epipolar line, and improve matching efficiency.
[0082] Binocular stereo matching module: Binocular stereo matching is the core module of the binocular vision system. Its function is to calculate the disparity between pixels by analyzing the image pairs of the same scene captured by the left and right cameras, and then obtain depth information. It prepares for target size measurement and has the following functions:
[0083] 1. Disparity map generation: By matching corresponding pixels in the left and right images, the horizontal displacement (disparity) of each pixel is calculated and the output is a disparity map, where the brightness value represents the disparity size (inversely proportional to the depth).
[0084] 2. Depth information recovery: Based on the principle of triangulation and combined with the calibrated camera parameters (baseline distance, focal length), the disparity map is converted into a depth map, which directly reflects the distance of each point in the scene.
[0085] 3. Occlusion area processing: Detect the invisible areas (occluded parts) in the left and right perspectives and fill them with reasonable disparity values through interpolation or global optimization algorithms.
[0086] 4. Noise suppression: Eliminate false matches during the matching process (such as texture repetition areas and weak texture areas) to improve the robustness of the depth map.
[0087] Object Detection and Segmentation Module: This module uses a multi-task segmentation network to perform object detection, object classification, and pixel-level instance segmentation on pets, improving the high-precision positioning and recognition of pet objects. Its functions are as follows:
[0088] 1. Object detection: Locate the pet in the image or video and mark it with a bounding box.
[0089] 2. Target classification: Determine whether the detected pet belongs to a predefined category to avoid false detection (such as dogs, other objects).
[0090] 3. Instance segmentation: Perform pixel-level segmentation on the pet’s outline to obtain the pet’s outline information.
[0091] 3D measurement module: Based on the depth information obtained by the binocular stereo matching module and the pet outline information obtained by the target detection and segmentation module, the two are combined to measure the pet's size.
[0092] like Figure 2As shown, the binocular stereo calibration module includes a data acquisition unit, a monocular calibration unit, a binocular calibration unit and a stereo correction unit;
[0093] The data acquisition unit is used to use a binocular camera to shoot the same checkerboard from different angles (at least 10 groups), ensuring that the checkerboard is completely visible in both the left and right images and covers the edge and center areas of the image;
[0094] The single-target calibration unit (intrinsic parameter + distortion calibration) is used to detect corner points. A corner detection algorithm is used to extract checkerboard corner points for each image. A sub-pixel optimization algorithm is used to improve the corner point accuracy to 0.1 pixel level. Parameters are calculated. Based on the Zhang Zhengyou calibration method, a camera model is established to solve the intrinsic parameter matrix K and distortion coefficients (k1, k2, p1, p2, k3). The Levenberg-Marquardt algorithm is then used to minimize the reprojection error and optimize the parameters.
[0095] The dual-target calibration unit (external parameter calibration) is used for stereo matching corner points, matching the corner points of the left and right images to ensure the correspondence between the same physical point in the two images; calculating the extrinsic parameters, solving the essential matrix E through epipolar geometry, decomposing the rotation matrix R and the translation vector T, and jointly optimizing all parameters using the bundle adjustment method to minimize the 3D point reprojection error;
[0096] The stereo correction unit is used to calculate the correction transformation matrix (R1, R2, P1, P2) based on the calibration results, align the epipolar lines of the left and right images horizontally, generate a correction mapping table, and correct the new image in real time.
[0097] like Figure 3 As shown, the binocular stereo matching module includes a binocular data acquisition unit, a disparity map generation unit, a disparity map filtering unit and a depth map conversion unit;
[0098] The binocular data acquisition unit is used to collect video images from the left and right cameras at the same time. If there is a time difference, it will cause errors in the subsequent matching and thus result errors;
[0099] The disparity map generation unit is used to correct the input images of the left and right cameras according to the mapping table obtained by the binocular stereo calibration module so that the images of the left and right cameras are mapped to the same horizontal plane; the SGBM algorithm is used with the parameters set as the disparity range of 64 pixels and the penalty coefficient P1=8 / P2=32, and the 16-bit disparity map is output;
[0100] The disparity map filtering unit is used to filter the disparity map using a WLS algorithm to minimize disparity noise while protecting image edges;
[0101] The depth map conversion unit is used to convert the disparity map filtered by the left camera into a 3D space using a reprojection matrix Q to obtain a depth map and a world coordinate system for each pixel of the image.
[0102] like Figure 4 As shown, the target detection and segmentation module includes a model training unit, a model quantization unit and a model inference unit;
[0103] The model training unit is used to collect a large amount of pet data and design a multi-task network. The network output includes target detection, instance segmentation, and target recognition. Taking pet cats as an example, the types of targets are divided into two categories (long-haired cats and short-haired cats). In order to achieve real-time performance on an embedded platform, the present invention adopts a lightweight network structure such as yolov8n-seg, which meets the three required functions at the same time.
[0104] The model quantization unit is used to quantize the trained model into a binary file that can be run on the embedded system through a quantization tool;
[0105] The model inference unit is used to obtain the image of the left or right camera and normalize the size to the input size of the model; the model output target is the pet position coordinate information, pet classification information, credibility and pet mask in the image.
[0106] like Figure 5 As shown, the three-dimensional measurement module includes a contour acquisition unit, a contour calculation unit, a shortest distance point pair calculation unit, a valid point pair retention unit and a pet size measurement unit;
[0107] The contour acquisition unit is used to obtain the contour information of the pet in the image based on the mask obtained by model inference;
[0108] A contour calculation unit, used to calculate the centroid point of the contour;
[0109] The shortest distance point pair calculation unit calculates the intersection point pairs (P1, P2) of multiple straight lines passing through the contour centroid and the contour;
[0110] The valid point pair retention unit is used to search in the 3D image based on the shortest distance point pair, exclude points with invalid or out-of-range depth values, and retain valid point pairs P1(x1,y1,z1), P2(x2,y2,z2), ...;
[0111] Pet size measurement unit for use according to the formula Calculate the distances of multiple point pairs and calculate the average value avg_dist of these point pairs to obtain the average cross-sectional width of the pet.
[0112] Example 2
[0113] like Figure 6As shown, the present invention provides a pet size intelligent measurement method based on a multi-task segmentation network, comprising the following steps:
[0114] S1. Binocular stereo calibration: Determine the geometric relationship between the two cameras; including: calculating camera intrinsic parameters, calculating the inherent parameters of a single camera, describing the camera's imaging characteristics; calculating camera extrinsic parameters, calculating the relative position and orientation between the two cameras, and defining their spatial relationship; eliminating lens distortion, correcting radial and tangential distortion of the lens; establishing epipolar constraints, obtaining the fundamental matrix or essential matrix through calibration, and compressing the stereo matching search range from the two-dimensional image to the one-dimensional epipolar line;
[0115] S2, binocular stereo matching, by analyzing the image pairs of the same scene captured by the left and right cameras, calculating the disparity between pixels, and then obtaining depth information; including: generating a disparity map by matching the corresponding pixels in the left and right images, calculating the horizontal displacement of each pixel, and outputting a disparity map, where the brightness value represents the magnitude of the disparity; restoring depth information, based on the principle of triangulation and combined with the calibrated camera parameters, converting the disparity map into a depth map that directly reflects the distance between each point in the scene; processing occluded areas, detecting areas that are not visible in the left and right perspectives, and filling in reasonable disparity values through interpolation or global optimization algorithms; suppressing noise, eliminating false matches during the matching process, and improving the robustness of the depth map;
[0116] S3, Object Detection and Segmentation, uses a multi-task segmentation network to perform object detection, object classification, and pixel-level instance segmentation on pets, improving the high-precision positioning and recognition of pet objects. This includes: object detection, locating the position of the pet in the image or video and marking it with a bounding box; object classification, determining whether the detected pet belongs to a predefined category to avoid false detection; instance segmentation, performing pixel-level segmentation on the pet's outline to obtain the pet's outline information;
[0117] S4. Three-dimensional measurement: Based on the depth information obtained from binocular stereo matching and the pet contour information obtained from target detection and segmentation, the two are combined to measure the pet size.
[0118] like Figure 7 As shown in Figure 2, binocular stereo calibration includes the following steps:
[0119] Step 1: Data Collection
[0120] Use a binocular camera to capture the same checkerboard from different angles (at least 10 groups), ensuring that the checkerboard is completely visible in both the left and right images and covers the edges and center areas of the images.
[0121] Step 2: Monocular camera calibration (intrinsic reference + distortion calibration)
[0122] 1) Corner detection: A corner detection algorithm is used to extract checkerboard corners for each image, and a sub-pixel optimization algorithm is used to improve the corner accuracy to 0.1 pixel level.
[0123] 2) Parameter calculation: Based on Zhang Zhengyou calibration method, establish a camera model, solve the intrinsic parameter matrix K and distortion coefficients (k1, k2, p1, p2, k3), and then use the Levenberg-Marquardt algorithm to minimize the reprojection error and optimize the parameters.
[0124] Step 3: Dual-target calibration (external reference calibration)
[0125] 1) Stereo matching corner points: Match the corner points of the left and right images to ensure the correspondence between the same physical point in the two images.
[0126] 2) Calculate extrinsic parameters: Solve the essential matrix E through epipolar geometry and decompose it into the rotation matrix R and translation vector T. Use the bundle adjustment method to jointly optimize all parameters to minimize the 3D point reprojection error.
[0127] Step 4: Stereo Correction
[0128] 1) According to the calibration results, calculate the correction transformation matrix (R1, R2, P1, P2) to align the left and right images horizontally.
[0129] 2) Generate a correction mapping table and correct the new image in real time.
[0130] like Figure 8 As shown in Figure 2, binocular stereo matching includes the following steps:
[0131] Step 1: Binocular data collection
[0132] The video images of the left and right cameras are collected at the same time. If there is a time difference, it will cause errors in the subsequent matching and thus result errors.
[0133] Step 2: Disparity map generation
[0134] 1) According to the mapping table obtained by the dual-target positioning module, the input images of the left and right cameras are corrected so that the images of the left and right cameras are mapped to the same horizontal plane.
[0135] 2) Use the SGBM algorithm with the parameters set to a disparity range of 64 pixels, a penalty coefficient P1 = 8 / P2 = 32, and output a 16-bit disparity map.
[0136] Step 3: Disparity map filtering
[0137] The WLS algorithm is used to filter the disparity map to minimize the disparity noise while protecting the image edges.
[0138] Step 4: Depth Map Conversion
[0139] The reprojection matrix Q is used to transform the filtered disparity map of the left camera into 3D space to obtain the depth map and the world coordinate system of each pixel of the image.
[0140] like Figure 9 As shown in Figure 2, target detection and segmentation include the following steps:
[0141] Step 1: Model training
[0142] Taking pet cats as an example, we collected a large amount of cat data and designed a multi-task network. The network outputs include object detection, instance segmentation, and object recognition. The objects are divided into two types (long-haired cats and short-haired cats). To achieve real-time performance on an embedded platform, we used the lightweight network structure yolov8n-seg, which simultaneously meets all three required functions.
[0143] Step 2: Model Quantization
[0144] Quantize the trained model into a binary file that can be run on an embedded system using a quantization tool;
[0145] Step 3: Model Inference
[0146] 1) Obtain the image from the left (right) camera and normalize the size to the input size of the model;
[0147] 2) The model output targets are the pet location coordinate information, pet classification information, credibility, and pet mask in the image.
[0148] Specifically, three-dimensional measurement includes the following steps:
[0149] 1) Obtain the outline information of the cat in the image based on the mask obtained by model reasoning;
[0150] 2) Calculate the centroid of the contour;
[0151] 3) Calculate the intersection points (P1, P2) of multiple straight lines passing through the contour centroid and the contour;
[0152] 4) Based on these point pairs, search in the 3D graph, exclude points with invalid or out-of-range depth values, and retain valid point pairs P1(x1,y1,z1), P2(x2,y2,z2), ...;
[0153] 5) According to the formula Calculate the distances of multiple point pairs and calculate the average value avg_dist of these point pairs to obtain the average cross-sectional width of the pet.
[0154] Example 3
[0155] Based on the same concept, the present invention also provides a schematic diagram of a physical structure, such as Figure 10 As shown, the server may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the steps of the intelligent pet size measurement method based on the multi-task segmentation network.
[0156] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0157] Example 4
[0158] Based on the same concept, the present invention also provides a non-transitory computer-readable storage medium, which stores a computer program. The computer program includes at least one code segment, which can be executed by a main control device to control the main control device to implement the steps of the pet size intelligent measurement method based on the multi-task segmentation network.
[0159] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in this application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0160] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0161] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A pet size intelligent measurement system based on a multi-task segmentation network, characterized in that: include: The binocular stereo calibration module is used to determine the geometric relationship between the two cameras. This includes: calculating the camera intrinsic parameters, calculating the inherent parameters of a single camera, and describing the camera's imaging characteristics; calculating the camera extrinsic parameters, calculating the relative position and orientation between the two cameras and defining their spatial relationship; eliminating lens distortion, correcting radial and tangential distortion of the lens; establishing epipolar constraints, and obtaining the fundamental matrix or essential matrix through calibration, which compresses the stereo matching search range from the two-dimensional image to the one-dimensional epipolar line; The binocular stereo matching module is used to calculate the disparity between pixels and obtain depth information by analyzing image pairs of the same scene captured by left and right cameras. This includes: generating a disparity map by matching corresponding pixels in the left and right images, calculating the horizontal displacement of each pixel, and outputting a disparity map in which the brightness value represents the magnitude of the disparity; recovering depth information by converting the disparity map into a depth map based on the principle of triangulation and combined with calibrated camera parameters, which directly reflects the distance between each point in the scene; processing occluded areas, detecting areas that are not visible in the left and right perspectives, and filling in reasonable disparity values through interpolation or global optimization algorithms; and suppressing noise to eliminate mismatches during the matching process and improve the robustness of the depth map. The object detection and segmentation module uses a multi-task segmentation network to perform object detection, object classification, and pixel-level instance segmentation on pets, improving the high-precision positioning and recognition of pet objects. This module includes: object detection, which locates the position of the pet in the image or video and marks it with a bounding box; object classification, which determines whether the detected pet belongs to a predefined category to avoid false detection; and instance segmentation, which performs pixel-level segmentation on the pet's outline to obtain the pet's outline information. The 3D measurement module is used to measure the pet's size based on the depth information obtained by the binocular stereo matching module and the pet's outline information obtained by the target detection and segmentation module.
2. The pet size intelligent measurement system based on multi-task segmentation network according to claim 1 is characterized in that: The binocular stereo calibration module includes: A data acquisition unit is used to use a binocular camera to shoot the same checkerboard from different angles, ensuring that the checkerboard is completely visible in both the left and right images and covers the edge and center areas of the image; The single-target positioning unit is used to detect corner points. A corner detection algorithm is used to extract checkerboard corner points for each image. A sub-pixel optimization algorithm is used to improve the corner point accuracy to a certain pixel level. The unit also calculates parameters. Based on the Zhang Zhengyou calibration method, a camera model is established to solve the intrinsic parameter matrix and distortion coefficients, and then the reprojection error is minimized to optimize the parameters. The dual-target positioning unit is used for stereo corner matching, matching the corner points of the left and right images to ensure the correspondence between the same physical point in the two images. It also calculates the extrinsic parameters, solves the essential matrix through epipolar geometry, decomposes it into the rotation matrix and translation vector, and uses the bundle adjustment method to jointly optimize all parameters to minimize the 3D point reprojection error. The stereo correction unit is used to calculate the correction transformation matrix based on the calibration results, align the epipolar lines of the left and right images horizontally, generate a correction mapping table, and correct the new image in real time.
3. The pet size intelligent measurement system based on multi-task segmentation network according to claim 1 is characterized in that: The binocular stereo matching module includes: The binocular data acquisition unit is used to collect video images from the left and right cameras at the same time. If there is a time difference, it will cause errors in the subsequent matching and thus result errors; The disparity map generation unit is used to correct the input images of the left and right cameras based on the mapping table obtained by the binocular stereo calibration module so that the images of the left and right cameras are mapped to the same horizontal plane. The SGBM algorithm is used with the parameters set to a disparity range of 64 pixels and a penalty coefficient of P1 = 8 / P2 = 32, and the output is a 16-bit disparity map. A disparity map filtering unit is used to filter the disparity map using the WLS algorithm to minimize disparity noise while protecting image edges; The depth map conversion unit is used to convert the disparity map filtered by the left camera into 3D space using the reprojection matrix Q to obtain the world coordinate system of the depth map and each pixel of the image.
4. The pet size intelligent measurement system based on multi-task segmentation network according to claim 1 is characterized in that: The target detection and segmentation module includes: The model training unit is used to collect a large amount of pet data and design a multi-task network. The network output includes object detection, instance segmentation, and object recognition; The model quantization unit is used to quantize the trained model into a binary file that can be run on the embedded system through the quantization tool; The model inference unit is used to obtain images from the left or right camera and normalize the size to the model input size; the model output targets are the pet position coordinate information, pet classification information, credibility, and pet mask in the image.
5. The pet size intelligent measurement system based on multi-task segmentation network according to claim 1 is characterized in that: The three-dimensional measurement module includes: The contour acquisition unit is used to obtain the contour information of the pet in the image based on the mask obtained by model inference; A contour calculation unit, used to calculate the centroid point of the contour; The shortest distance point pair calculation unit calculates the intersection point pairs (P1, P2) of multiple straight lines passing through the contour centroid and the contour; The valid point pair retention unit is used to search in the 3D image based on the shortest distance point pair, exclude points with invalid or out-of-range depth values, and retain valid point pairs P1(x1,y1,z1), P2(x2,y2,z2), ...; Pet size measurement unit for use according to the formula Calculate the distances of multiple point pairs and calculate the average value avg_dist of these point pairs to obtain the average cross-sectional width of the pet.
6. A pet size intelligent measurement method based on a multi-task segmentation network, characterized in that: The steps include: Binocular stereo calibration determines the geometric relationship between the two cameras. This includes: calculating camera intrinsic parameters, calculating the inherent parameters of a single camera, and describing the camera's imaging characteristics; calculating camera extrinsic parameters, calculating the relative position and orientation between the two cameras and defining their spatial relationship; eliminating lens distortion, correcting radial and tangential distortion of the lens; establishing epipolar constraints, and obtaining the fundamental matrix or essential matrix through calibration, which compresses the stereo matching search range from the two-dimensional image to the one-dimensional epipolar line. Binocular stereo matching analyzes image pairs of the same scene captured by left and right cameras, calculates the disparity between pixels, and thus obtains depth information. This includes: generating a disparity map by matching corresponding pixels in the left and right images, calculating the horizontal displacement of each pixel, and outputting a disparity map where the brightness value represents the magnitude of the disparity; restoring depth information by converting the disparity map into a depth map based on the principle of triangulation and combined with calibrated camera parameters, which directly reflects the distance between each point in the scene; processing occluded areas, detecting areas that are not visible in the left and right perspectives, and filling in reasonable disparity values through interpolation or global optimization algorithms; and suppressing noise to eliminate mismatches during the matching process and improve the robustness of the depth map. Object detection and segmentation uses a multi-task segmentation network to perform object detection, object classification, and pixel-level instance segmentation on pets, improving the high-precision positioning and recognition of pet objects. This includes: object detection, locating the position of the pet in the image or video and marking it with a bounding box; object classification, determining whether the detected pet belongs to a predefined category to avoid false detection; instance segmentation, performing pixel-level segmentation on the pet's outline to obtain the pet's outline information; Three-dimensional measurement is based on the depth information obtained from binocular stereo matching and the pet outline information obtained from target detection and segmentation, and the two are combined to measure the pet's size.
7. The pet size intelligent measurement method based on multi-task segmentation network according to claim 6 is characterized in that: The binocular stereo calibration comprises the following steps: Data collection: Use a binocular camera to capture the same checkerboard from different angles, ensuring that the checkerboard is completely visible in both the left and right images, and covers the image edges and center areas; Single-target localization: Detect corners. A corner detection algorithm is used to extract checkerboard corners from each image. A sub-pixel optimization algorithm is used to improve corner accuracy to a certain pixel level. Parameters are calculated. Based on the Zhang Zhengyou calibration method, a camera model is established to solve for the intrinsic parameter matrix and distortion coefficients. The reprojection error is then minimized to optimize the parameters. Dual-target positioning: Stereo matching corner points, matching the corner points of the left and right images to ensure the correspondence between the same physical point in the two images; Calculate the extrinsic parameters, solve the essential matrix through epipolar geometry, decompose it to obtain the rotation matrix and translation vector, and use the bundle adjustment method to jointly optimize all parameters to minimize the 3D point reprojection error; Stereo correction: Based on the calibration results, the correction transformation matrix is calculated to align the epipolar lines of the left and right images horizontally, and a correction mapping table is generated to correct the new image in real time.
8. The pet size intelligent measurement method based on multi-task segmentation network according to claim 6 is characterized in that: The binocular stereo matching includes: Binocular data acquisition: The video images of the left and right cameras are collected at the same time. If there is a time difference, it will cause errors in the subsequent matching and thus lead to error in the results; Disparity map generation: Based on the mapping table obtained from the binocular stereo calibration module, the input images of the left and right cameras are corrected so that the images of the left and right cameras are mapped to the same horizontal plane. The SGBM algorithm is used with the parameters set to a disparity range of 64 pixels and a penalty coefficient of P1 = 8 / P2 = 32, and a 16-bit disparity map is output. Disparity map filtering: used to filter the disparity map using the WLS algorithm to minimize disparity noise while protecting image edges; Depth map conversion: Use the reprojection matrix Q to convert the disparity map filtered by the left camera into 3D space to obtain the world coordinate system of the depth map and each pixel of the image.
9. The pet size intelligent measurement method based on multi-task segmentation network according to claim 6, characterized in that: The target detection and segmentation include: Model training: Collect a large amount of pet data and design a multi-task network. The network output includes object detection, instance segmentation, and object recognition. Model quantization: The trained model is quantized into a binary file that can be run on an embedded system using a quantization tool. Model inference: Obtain images from the left or right camera and normalize the size to the model input size; the model output targets are the pet location coordinate information, pet classification information, credibility, and pet mask in the image.
10. The pet size intelligent measurement method based on multi-task segmentation network according to claim 6, characterized in that: The three-dimensional measurement includes: Contour acquisition: Obtain the contour information of the pet in the image based on the mask obtained by model inference; Contour calculation: used to calculate the centroid of the contour; Shortest distance point pair calculation: Calculate the intersection point pairs (P1, P2) of multiple straight lines passing through the contour centroid and the contour; Retain valid point pairs: Search in the 3D graph based on the shortest distance point pairs, exclude points with invalid or out-of-range depth values, and retain valid point pairs P1(x1,y1,z1), P2(x2,y2,z2), ...; Pet size measurement: According to the formula Calculate the distances of multiple point pairs and calculate the average value avg_dist of these point pairs to obtain the average cross-sectional width of the pet.
Citation Information
Cited By
Semantic segmentation-based logging cable outer diameter anomaly detection method
CN121258973A