Pet rope-pulling-free identification method based on fine-grained detection and segmentation

By training object detection and semantic segmentation models using a decoupled model, and combining fine-grained detection and segmentation methods, the accuracy problem of identifying unleashed pets in surveillance videos was solved, improving recognition accuracy and robustness.

CN121033741APending Publication Date: 2025-11-28CHINA CONSTR THIRD ENG BUREAU GRP CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510903296.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify whether a pet is on a leash in surveillance videos, especially when the leash is thin and far away. Target detection models have difficulty distinguishing the fine-grained visual features of whether a pet is on a leash or not.

Method used

A decoupled model is used to train the target detection model and the semantic segmentation model to detect pet targets and leashes respectively. Fine-grained detection and segmentation methods are used, combined with region of interest extraction and intersection ratio to determine whether the target is on a leash.

Benefits of technology

It improves the accuracy of identifying unleashed pets, reduces the false detection rate, enhances the robustness of the model output, and adapts to changes in pet posture and lighting conditions in different monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033741A_ABST
    Figure CN121033741A_ABST
Patent Text Reader

Abstract

The invention relates to a pet rope-free identification method based on fine-grained detection and segmentation, and the method comprises the steps: carrying out the marking of a pet target in original image data through employing a bounding box, and carrying out the marking of a pet rope in the original image data through employing a polygon contour; respectively training a target detection model and a semantic segmentation model by using the labeled image data, preprocessing to-be-analyzed image data, inputting the preprocessed to-be-analyzed image data into the trained target detection model, outputting a pet target bounding box, and expanding a region of interest (ROI) by using the position center of the pet target bounding box; and intercepting a corresponding region-of-interest image in the original image data according to the ROI coordinates, processing according to a semantic segmentation model image data preprocessing mode, inputting a trained semantic segmentation model to obtain a rope mask, calculating an intersection ratio of the rope mask to a pet target bounding box, if the intersection ratio is greater than or equal to a threshold value, judging that the pet target bounding box is a guy rope, and otherwise, judging that the pet target bounding box is a guy rope. Otherwise, judging that the rope is not pulled. According to the invention, whether the pet is pulled or not is accurately identified and distinguished based on the monitoring camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology and relates to a method for identifying pets that are not on a leash based on fine-grained detection and segmentation, which is suitable for the automated identification of pets on leashes in surveillance videos. Background Technology

[0002] With urbanization, pets are playing an increasingly important role in people's lives. Some pets, such as dogs, often live alongside humans and appear in public areas such as communities, streets, and shopping malls. Although dogs have been domesticated, they still pose certain dangers, especially in public places. To protect public safety, laws and regulations require dogs to be leashed in public to prevent serious incidents of dog bites. Despite these regulations, the uncivilized behavior of pets not being leashed persists in real life. There are various reasons for this, but one is the lack of effective monitoring technology to promptly detect such behavior.

[0003] Traditionally, identifying unleashed pets relies primarily on manual observation, with few automated methods that don't depend on human intervention. However, with the widespread use of surveillance cameras and the development of artificial intelligence (AI), computer vision methods based on AI are emerging to automatically identify targets in video images, replacing human eyes. Using existing AI vision algorithms for unleashed pet identification is one possible approach. The basic principle involves manually labeling collected images as either leashed or unleashed, then using a target detection model (such as YOLO, Faster R-CNN, or SSD series) to directly detect and locate the leashed or unleashed pet in the image. This method is relatively direct and end-to-end; however, its drawback is the difficulty in accurately distinguishing between leashed and unleashed states. This is because leashes are generally thin and long, appearing as a line from a distance in the surveillance camera's viewpoint, making them very inconspicuous in the image. This makes it difficult for the target detection model to differentiate and capture the fine-grained visual features of a leashed or unleashed pet.

[0004] The core flaw of existing technology lies in its failure to decouple the characteristic differences between pets and ropes, and its neglect of the fine-grained visual characteristics of ropes in monitoring scenarios. Summary of the Invention

[0005] This invention addresses the problem of conventional artificial intelligence vision algorithms failing to accurately identify unleashed pets. It proposes a pet unleashed identification method based on fine-grained detection and segmentation, utilizing surveillance cameras to accurately identify and distinguish between leashed and unleashed pets. When an unleashed pet is automatically detected, an alarm is issued. This solves the problem of false detections / missed detections caused by the concealment of leash features in existing technologies, improving the accuracy of unleashed pet identification in surveillance scenarios.

[0006] The technical solution adopted in this invention is as follows: A method for identifying unleashed pets based on fine-grained detection and segmentation, characterized by the following steps: Step 1 Data annotation: Use bounding boxes to annotate the pet targets in the original image data, and use polygon outlines to annotate the pet leash in the original image data; Step 2: Decoupling Model Training: The decoupling model consists of an object detection model and a semantic segmentation model. The object detection model detects pet targets and outputs the pet's bounding box, category label, and confidence score (confidence score is the AI ​​model's confidence probability of the output result, a common concept in the field; in practice, the model's output result can be adjusted by setting a confidence score threshold or selecting results that meet the confidence score threshold). The semantic segmentation model segments the rope and outputs the rope mask and confidence score. The pet target image data labeled in Step 1 is used to train the object detection model to obtain the trained object detection model. obj The labeled pet leash image data from step 1 is used to train the semantic segmentation model, resulting in a trained semantic segmentation model. seg ; Step 3 Image data preprocessing based on fine-grained detection and segmentation: Copy the original image data to be analyzed and scale it proportionally to the input layer size of the target detection model along the longer side of the original image data. Copy the original image data to be analyzed and scale it proportionally to the input layer size of the semantic segmentation model along the longer side of the original image data. The input size of the semantic segmentation model is larger than the input layer size of the target detection model. Step 4: Region of Interest Extraction Based on Fine-Grained Detection and Segmentation: Input the preprocessed image data from Step 3 into the trained object detection model. obj In the process, the bounding box of the pet target is output, and a region of interest (ROI) is expanded from the center of the pet target bounding box. Step 5: Rope Detection and Association Based on Fine-Grained Detection and Segmentation: Extract the corresponding region of interest (ROI) image from the original image data according to the ROI coordinates from Step 4. Process this ROI image using the image data preprocessing method described in Step 3 for the semantic segmentation model, and then input it into the trained semantic segmentation model. seg Obtain the rope mask, calculate the intersection ratio between the rope mask and the bounding box of the pet target, and if the intersection ratio is greater than or equal to the threshold, it is determined that the pet is leashed; otherwise, it is determined that the pet is not leashed.

[0007] The object detection model mentioned in step 2 is either the YOLO architecture based on convolutional neural networks (CNN) or the DETER architecture based on transformer neural networks (Transformer). The semantic segmentation model is either the SegNet architecture based on convolutional neural networks (CNN) or the Mask2Former architecture based on transformers (Transformer).

[0008] Step 2, specifically the training of the object detection model, involves taking 1000-1500 images annotated with bounding boxes from Step 1 and training them using any deep learning framework such as PyTorch, TensorFlow, or PaddlePaddle. The training continues until either the YOLO or DETER architecture converges or 100 training iterations are completed, resulting in the trained object detection model. obj Take 1000-1500 images labeled with polygonal contours from step 1, and train them using any deep learning framework such as PyTorch, TensorFlow, or PaddlePaddle. Train the model using either the SegNet architecture based on a convolutional neural network (CNN) or the Mask2Former architecture based on a transformer (Transformer). Train until the SegNet or Mask2Former architecture converges or after 100 training iterations to obtain the trained semantic segmentation model. seg .

[0009] Step 3: Scale the input size proportionally to the target detection model, with a resolution of 640×640 and a padding pixel value of 128; scale the input size proportionally to the semantic segmentation model, with a resolution of 1024×1024 and a padding pixel value of 128.

[0010] The expansion factor of the Region of Interest (ROI) in step 4 is 2-4 times the area of ​​the pet target bounding box, and the ROI region does not exceed the boundary of the original image.

[0011] The threshold C of the intersection ratio mentioned in step 5 th The value range is 0.01-0.1, and noise masking is removed through morphological filtering.

[0012] The beneficial effects of this invention are as follows: 1. Improved Accuracy: To address the challenge of accurately determining whether a pet is leashed in tasks involving unleashed pets, a non-end-to-end decoupling method is employed. One object detection model focuses solely on detecting the pet, while a semantic segmentation model focuses solely on segmenting the leash. This significantly reduces the difficulty of identifying pets and leashes. This is because object detection models are concise, efficient, and highly accurate for targets with relatively simple features (such as pets), making them ideal for pet identification. However, they are unsuitable for relatively complex targets lacking distinct features, such as leashes. While semantic segmentation models typically have higher computational complexity than object detection models, they are well-suited for handling relatively complex targets lacking obvious features. This invention combines object detection and semantic segmentation models to effectively solve the problem of identifying pets and leashes.

[0013] 3.2 To address the challenge of identifying target pets and ropes in raw images acquired from a surveillance perspective due to their small scale, this invention employs a fine-grained, staged identification method based on a non-end-to-end, multi-model hybrid approach. First, the target pet is detected. Then, Regions of Interest (ROIs) are extracted from candidate target pets. Semantic segmentation is then performed only on the image regions of the ROIs to identify the rope. Finally, prior knowledge (if a pet is on a leash, then the rope and pet intersect) is used to filter and identify the ropes identified by the model. This method not only solves the target (especially slender ropes) scale problem but also avoids mutual interference between different models, enhancing robustness to model output noise. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the offline portion of the present invention.

[0015] Figure 2 This is a schematic diagram of the online portion of the present invention.

[0016] Figure 3 This is a schematic diagram of the algorithm analysis process of the present invention.

[0017] Figure 4 This is a schematic diagram of the post-processing flow for identification and judgment in this invention. Detailed Implementation

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and by way of listing some optional embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] Example 1: like Figure 1 , Figure 2System architecture diagram. (e.g.) Figure 1 As shown, the offline part includes the following processing steps: First, acquire raw image data 1. This raw image data contains real images from a monitoring perspective, including pets (such as dogs), both unleashed and leashed. The raw images are then labeled using two different methods: bounding box annotation 2, where a rectangle encloses the pet target, labeled as "pet" (only one label is allowed); and polygon annotation 4, where a polygon is used to annotate the outline of the target's leash, labeled as "rope" (only one label is allowed). After bounding box and polygon annotations, training datasets 3 (for the object detection model) and 5 (for the semantic segmentation model) are obtained, typically consisting of 1000-1500 images. Any mainstream deep learning framework can be used, such as PyTorch, TensorFlow, or PaddlePaddle. The object detection model (either a YOLO architecture based on a Convolutional Neural Network (CNN) or a DETER architecture based on a Transformer Neural Network (Transformer)) and the semantic segmentation model (either a SegNet architecture based on a CNN or a Mask2Former architecture based on a Transformer) are trained respectively. Models with different parameter scales can be used depending on the actual application requirements. In this specific embodiment, the object detection model 6 uses a YOLO architecture based on a CNN, and the semantic segmentation model 7 uses a Mask2Former architecture based on a Transformer. The training objective is model convergence or approximately 100 training iterations. After training, a trained object detection model is obtained. obj and a trained semantic segmentation model seg .

[0020] Online components such as Figure 2 As shown, the process includes the following steps: A surveillance camera 9, which can be any surveillance camera in public areas such as communities, streets, and shopping malls; a video stream 10, obtained from the surveillance camera or its corresponding network video storage device (NVR) via video software, such as video system software supporting mainstream video transmission protocols RTSP / RTMP and national standard GB / T 28181; the video stream is decoded into image frames 11, which are the original image I. orig Original Image I orig The image is fed into the algorithm analysis module 12 for real-time analysis; the algorithm analysis module outputs the original image I. origThe analysis results are output in a structured format. This structured output (I3) can be in a common JSON format, facilitating integration with other business systems. The structured output may include whether a pet was detected; if so, the information for each pet in the original image (I4). orig The information includes the location, confidence level, and the pet's status (leashed / unleashed).

[0021] Algorithm Analysis 12 Modules, such as Figure 3 As shown, the original input image I orig The data preprocessing step 14 is input to the target detection model. The main preprocessing method is based on image I. orig The model is scaled proportionally to the longer side as a baseline to the trained object detection model. obj The input layer size is 640x640, and padding is applied with a padding pixel value of 128, resulting in an image I with a size equal to the input size of the object detection model. obj Image I obj After normalization, the data is fed into the trained object detection model. obj Inference 15, the trained object detection model. obj Using YOLO11, a well-trained object detection model is Model obj Reasoning 15 generates multiple target pets {pet} i}. Target pet {pet i} Send to the Region of Interest (ROI) for ROI extraction 16, based on each target pet. i Location information expands a target pet i The rectangular region centered on the image I is defined, provided it does not exceed the boundaries (if the rectangular region extends beyond the image I). orig The boundaries are then truncated, meaning the x-coordinate of the ROI region is guaranteed to be greater than 0 and less than the image I. orig The maximum width, with the y-coordinate guaranteed to be greater than 0 and less than the image I. orig (maximum height), to the ROI region in image I orig Capture the target pet (pet) i Corresponding ROI image I i_roi There are multiple target pets {pet} i This will generate multiple ROI images {I}. i_roi}. ROI image {I i_roi The data is preprocessed into the semantic segmentation model (17). The preprocessing method is based on image I. i_roi The model is scaled proportionally to the longer side as a baseline for semantic segmentation. segThe input layer size is 1024x1024, and padding is applied with a padding pixel value of 128, resulting in multiple preprocessed images {I} with the same size as the input of the semantic segmentation model. i_roi_seg Image {I} i_roi_seg After normalization, the data is fed into the trained semantic segmentation model. seg Reasoning 18. The trained semantic segmentation model. seg Using SegFormer, the trained semantic segmentation model Model seg Inference 18 generates multiple segmentation masks for the target rope {rope} i_j_mask The input is given to the semantic segmentation model. seg Inference 18 involves multiple images {I} i_roi_seg}, employing a parallel processing method, simultaneously processes multiple images {I}. i_roi_seg Input from} to speed up the processing. Multiple segmentation masks {rope i_j_mask The image is then sent to the post-processing stage 19 for identification and judgment, and finally the corresponding original input image I is generated. orig Structured output {R orig}

[0022] Post-processing module 19, such as identification and judgment Figure 4 As shown, given the target pet, pet i and the corresponding multiple segmentation masks {rope i_j_mask}, take one mask rope in sequence i_j_mask Calculation and target pet pet i Intersection relationship, traversal mask rope i_j_mask Each point P in k Determine point P k Is it in pet i ={x i_min y i_min x i_max y i_max Within the region, if it exists, count it (Count). i_j_mask Add 1; assume mask rope i_j_mask The total number of points is Count. i_j_mask_total Calculate the ratio i_j =Count i_j_mask / Count i_j_mask_total If ratio i_j ≥C th C th If the value ranges from 0.01 to 0.1, then it is determined that the target pet is in the pet area. i A rope was found on the body, which is the target pet. iBeing led by a rope, then keeping the rope i_j_mask And place the target pet pet i Corresponding status i =false; otherwise, delete the rope. i_j_mask If {rope i_j_mask The final value is empty, i.e., the original {rope}. i_j_mask Each mask in} i_j_mask None of them can satisfy the ratio i_j ≥C th The condition indicates that the target pet is a pet. i Unleashed, targeted pet i Corresponding status i =true.

[0023] Finally, after screening and judging, the structured output {R} of post-processing step 19 is processed. orig The data can be in JSON format, for example: {"has_pet":true, "num_pets":2, "pet_status":[{"pet":[12, 50, 112, 150],"status":false, "confidence":0.8}, {"pet":[200, 45, 400, 145], "status":true,"confidence":0.9}]}, where the field "has_pet" represents image I. orig The field "num_pets" indicates the number of pets, and "pet_status" indicates the pet information data detected. It is a list, and each item in the list consists of the fields "pet" (location information), "status" (status information), and "confidence" (confidence information).

[0024] Beneficial effects of this invention: 1. Improved accuracy: Rope detection accuracy is 92.3% (compared to 78.5% for single-model methods); low false positive rate for no rope.

[0025] 2. Efficiency optimization: Parallel ROI processing enables inference speeds of up to 25 FPS (single RTX 3090 card); staged processing reduces GPU memory usage by 40%.

[0026] 3. Enhanced robustness: Supports leash detection under changes in pet posture (such as lying down, jumping); adapts to low light (brightness < 5 lux) and occlusion scenarios (occlusion rate < 30%).

[0027] It will be readily understood by those skilled in the art that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, combinations, substitutions, improvements, etc., made under the spirit and principles of the present invention are included within the protection scope of the present invention.

Claims

1. A method for identifying unleashed pets based on fine-grained detection and segmentation, characterized in that... Includes the following steps: Step 1 Data annotation: Use bounding boxes to annotate the pet targets in the original image data, and use polygon outlines to annotate the pet leash in the original image data; Step 2: Decoupling Model Training: The decoupling model consists of an object detection model and a semantic segmentation model. The object detection model detects pet targets and outputs the pet's bounding box, category label, and confidence score. The semantic segmentation model segments the rope and outputs the rope mask and confidence score. The object detection model is trained using the pet target image data labeled in Step 1 to obtain the trained object detection model. obj The labeled pet leash image data from step 1 is used to train the semantic segmentation model, resulting in a trained semantic segmentation model. seg ; Step 3 Image data preprocessing based on fine-grained detection and segmentation: Copy the original image data to be analyzed and scale it proportionally to the input layer size of the target detection model along the longer side of the original image data. Copy the original image data to be analyzed and scale it proportionally to the input layer size of the semantic segmentation model along the longer side of the original image data. The input size of the semantic segmentation model is larger than the input layer size of the target detection model. Step 4: Region of Interest Extraction Based on Fine-Grained Detection and Segmentation: Input the preprocessed image data from Step 3 into the trained object detection model. obj In the process, the bounding box of the pet target is output, and a region of interest (ROI) is expanded from the center of the pet target bounding box. Step 5: Rope Detection and Association Based on Fine-Grained Detection and Segmentation: Extract the corresponding region of interest (ROI) image from the original image data according to the ROI coordinates from Step 4. Process this ROI image using the image data preprocessing method described in Step 3 for the semantic segmentation model, and then input it into the trained semantic segmentation model. seg Obtain the rope mask, calculate the intersection ratio between the rope mask and the bounding box of the pet target, and if the intersection ratio is greater than or equal to the threshold, it is determined that the pet is leashed; otherwise, it is determined that the pet is not leashed.

2. The method for identifying unleashed pets based on fine-grained detection and segmentation according to claim 1, characterized in that: The object detection model mentioned in step 2 is either the YOLO architecture based on convolutional neural networks (CNN) or the DETER architecture based on transformer neural networks (Transformer). The semantic segmentation model is either the SegNet architecture based on convolutional neural networks (CNN) or the Mask2Former architecture based on transformers (Transformer).

3. The method for identifying unleashed pets based on fine-grained detection and segmentation according to claim 2, characterized in that: Step 2, specifically the training of the object detection model, involves taking 1000-1500 images annotated with bounding boxes from Step 1 and training them using any deep learning framework such as PyTorch, TensorFlow, or PaddlePaddle. The training continues until either the YOLO or DETER architecture converges or 100 training iterations are completed, resulting in the trained object detection model. obj Take 1000-1500 images labeled with polygonal contours from step 1, and train them using any deep learning framework such as PyTorch, TensorFlow, or PaddlePaddle. Train the model using either the SegNet architecture based on a convolutional neural network (CNN) or the Mask2Former architecture based on a transformer (Transformer). Train until the SegNet or Mask2Former architecture converges or after 100 training iterations to obtain the trained semantic segmentation model. seg .

4. The method for identifying unleashed pets based on fine-grained detection and segmentation according to claim 1, characterized in that: Step 3: Scale the input size proportionally to the target detection model, with a resolution of 640×640 and a padding pixel value of 128; scale the input size proportionally to the semantic segmentation model, with a resolution of 1024×1024 and a padding pixel value of 128.

5. The method for identifying unleashed pets based on fine-grained detection and segmentation according to claim 1, characterized in that: The expansion factor of the Region of Interest (ROI) in step 4 is 2-4 times the area of ​​the pet target bounding box, and the ROI region does not exceed the boundary of the original image.

6. The method for identifying unleashed pets based on fine-grained detection and segmentation according to claim 1, characterized in that: The threshold C of the intersection ratio mentioned in step 5 th The value range is 0.01-0.1, and noise masking is removed through morphological filtering.

Citation Information

Patent Citations

  • Dog walking behavior detection method and device

    CN115035591A

  • Neural network-based guy rope state monitoring method and device, equipment and medium

    CN116863317A

  • Method and device for identifying dog carrying behavior and storage medium

    CN117315775A

  • Method for detecting pet pulling rope entering ladder based on semantic analysis

    CN118298285A

  • Tree line contradiction state recognition method and device based on target detection and semantic segmentation

    CN118762274A