Dynamic pet urination and defecation identification method based on hybrid cascade

By using a hybrid cascaded approach that combines target detection, multi-target tracking, and ROI extraction, the problems of data collection difficulties, multi-target interference, and poor recognition of small targets in pet defecation behavior recognition are solved, achieving accurate recognition and improved robustness in complex scenarios.

CN121033740APending Publication Date: 2025-11-28CHINA CONSTRUCTION THIRD ENGINEERING BUREAU WUCHUANG YUNWEI TECHNOLOGY (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510903249.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies face challenges in identifying pet defecation behavior, including difficulties in data collection, interference from multiple targets, and poor recognition of small targets. In particular, they struggle to accurately identify the actions of multiple pets under long-distance monitoring, and are also subject to severe background dynamic interference.

Method used

By employing a hybrid cascaded approach, which decouples target detection and action recognition and combines multi-target tracking and ROI extraction, accurate identification of pet defecation behavior can be achieved.

Benefits of technology

It effectively solves the problem of misjudgment in multi-pet scenarios, improves the accuracy of action recognition, reduces deployment costs, and enhances the robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033740A_ABST
    Figure CN121033740A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic pet urination and defecation identification method based on hybrid cascading, and the method comprises the steps: collecting a video stream through deploying a monitoring camera, analyzing the video stream in real time, and outputting the alarm information of pet urination and defecation. Through hybrid cascade design, target detection and action recognition are decoupled, the dependence of an end-to-end model on a large amount of customized data is avoided, the method can be completed through a general model, and the deployment cost is reduced; multi-target tracking and ROI extraction: through track screening and ROI interception, background interference is eliminated, continuous actions of a single pet are focused, and the problem of misjudgment in a multi-pet scene is solved; background noise in a small target scene is reduced through ROI extraction, and the action recognition accuracy is improved; dynamic windows and parameters are adjustable, parameters such as sampling frequency and window time can be configured, different monitoring scenes are adapted, and robustness is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and intelligent monitoring technology, specifically to a dynamic pet excrement recognition method based on hybrid cascading, applicable to pet behavior monitoring in public areas such as communities, streets, and shopping malls. Background Technology

[0002] Traditionally, pet waste identification relies primarily on manual observation, with few automated methods that don't depend on human intervention. However, with the widespread use of surveillance cameras and the development of artificial intelligence (AI), computer vision methods based on AI are emerging to automatically identify behavioral events in video images, replacing human eyes. Using existing AI vision algorithms for pet waste identification is one possible approach. The basic principle involves manually labeling the actions of pets defecating in a large number of collected video clips, and then using an action recognition network model (such as the SlowFast series or TSN series) to directly identify whether a pet is defecating in the video clip. This method is relatively direct and end-to-end; its drawbacks include the need to collect a large number of video clips of pet defecation actions, especially those from long-distance monitoring perspectives, which is time-consuming (finding pet defecation in surveillance videos is a low-frequency event). Furthermore, this method struggles to handle situations where multiple pets are present in the video, i.e., processing the actions of multiple pets simultaneously without interference and correctly distinguishing which pet's actions are being performed. Furthermore, existing motion recognition algorithms struggle to accurately identify the actions of small targets (those occupying only a small portion of the frame) because other dynamic changes in the background can cause significant interference. However, such complex scenarios are very common from the perspective of surveillance cameras in public areas.

[0003] The main defects are as follows: 1. Data collection difficulties: Pet excrement is a low-frequency event, requiring a large amount of labeled video footage, which is especially difficult to obtain under long-distance monitoring. 2. Multi-target interference: It cannot effectively distinguish the actions of multiple pets, making it easy to misjudge; 3. Poor recognition of small targets: Pets are small in proportion at long distances, and dynamic background interference leads to low recognition accuracy.

[0004] This invention proposes a hybrid cascade method that solves the above problems by decoupling target detection and action recognition, multi-target tracking, and ROI extraction. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a dynamic pet defecation recognition method based on hybrid cascading, which can achieve accurate and robust recognition of pet defecation behavior in complex scenarios.

[0006] The technical solution adopted in this invention is as follows: A dynamic pet excrement recognition method based on hybrid cascade, characterized by the following steps: Step 1: Downsample the input continuous video image frames at a sampling frequency of S. det, And 5≤S det ≤15fps, continuous sampling time T det And 5≤T det ≤10s, obtain the total number of frames F sa =S det ×T det Store the sampled frames into a first-in-first-out (FIFO) image queue {I} orig_i}, where i represents the i-th image in the image queue, and when the queue length reaches F. sa When the time comes, proceed to step 2, and simultaneously update the queue using the sliding window method, popping up the percentage R. out The historical frames, and 1 / 4≤Rout≤1 / 2; Step 2: For each frame of image I in the queue orig_i Preprocessing is performed by scaling the image proportionally to the input size of the object detection model along its longer side, padding with 128 pixels, normalizing the image, and then inputting it into the object detection model. det Output N bounding boxes for the pet targets {pet} i_j}, pet i_j For image queue {I orig_i The i-th image frame I orig_i The j-th pet was detected, where N is the total number of pets, and j ≤ N; Step 3: Transfer the image queue {I} orig_i} and the detected pet target bounding box {pet i_j Input the multi-target tracker TRacker. The multi-target tracker TRacker tracks multiple targets simultaneously and generates historical tracking trajectories {Tra j}, j≤N, each trajectory Tra j Indicates the target pet 1_j Coordinates of consecutive bounding boxes in the queue; Step 4: For each trajectory Tra j Perform filtering, with a length equal to F. sa Then retain it and calculate a rectangular ROI region. j Make the rectangular ROI j The region can contain historical tracking tracks. j For each bounding box in the range, ensure that the rectangular ROI is defined. j The region does not exceed the original image frame I orig_i Under the boundary conditions; sequentially in the image frame queue {Iorig_i In each frame of the image, a rectangular region of interest (ROI) of size} is extracted. j Sub-images of the region are used to construct a pet-targeted image. 1_j The length is F sa Continuous image queue { I pet_j_i After drawing the trajectory bounding box, proceed according to the sampling parameter S. act, For a continuous image queue { I pet_j_i} The continuous image queue {I} input to the action behavior recognition model is obtained by downsampling. act_j_k}, and 1 / 5 ≤ S act ≤1 / 2; Step 5: For {I act_j_k The image queue { I} is preprocessed to the input size of the action behavior recognition model by scaling and padding the images proportionally to their longer sides. act_obj_j_k}, preprocessed image queue { I act_obj_j_k After normalization, the data is fed into the action / behavior recognition model. act Action and behavior recognition model act The inference outputs a structured result, which includes the location information, confidence level information, and a judgment on whether a pet has defecated or urinated.

[0007] This invention processes multi-pet scenarios in parallel and outputs structured results in JSON format. The structured results output includes an analysis window {I}. orig_i The system includes the location information, confidence level, and determination of whether each identified pet is defecating or urinating.

[0008] The target detection model in step 2 is a general open set detection model, and the detection label is limited to pet type. No customized training is required, and the pet type is selected as dog.

[0009] The multi-target tracker TRacker in step 3 is based on the track-by-detection method. It does not require training a dedicated detection model. It supports dynamic tracking of multiple targets by associating the detection results with the trajectory.

[0010] ROI in step 4 j The extraction rule is: based on the trajectory Tra j The maximum left, right, top, and bottom coordinates of all bounding boxes in the ROI determine the rectangular region; if the ROI j If the image exceeds the original image boundary, it will be truncated and corrected according to the boundary.

[0011] In step 4, the action behavior recognition model input image queue {I act_j_k The sampling parameters S of} actUsed to reduce the frame rate, such as retaining 1 / 3 of the frame when Sact=1 / 3, with an input size of 896×896, a multiple of 224, to adapt to the pre-trained backbone network.

[0012] The beneficial effects of this invention are as follows: 1. Hybrid cascaded design: Object detection and action recognition are decoupled, avoiding the dependence of end-to-end models on a large amount of customized data. General models (such as GroundingDINO, Qwen2.5-VL) can be used to complete the task, reducing deployment costs.

[0013] 2. Multi-target tracking and ROI extraction: By filtering the trajectory and truncation the ROI, background interference is eliminated, and the continuous movements of a single pet are focused on, solving the problem of misjudgment in multi-pet scenarios; ROI extraction reduces background noise for small targets and improves the accuracy of action recognition.

[0014] 3. Dynamic window and adjustable parameters: Sampling frequency (S det S act ), window time (T) det Parameters such as ) can be configured to adapt to different monitoring scenarios (such as long distance, different pet densities) and enhance robustness. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the system structure of the present invention.

[0016] Figure 2 This is a schematic diagram of the algorithm analysis process of the present invention.

[0017] Figure 3 This is a schematic diagram of the multi-target ROI extraction process of the present invention. Detailed Implementation

[0018] The technical solutions (including preferred technical solutions) of the present invention will be further described in detail below with reference to the accompanying drawings and by way of listing some optional embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] A dynamic pet excrement recognition method based on hybrid cascade, characterized by the following steps: Step 1: Downsample the input continuous video image frames at a sampling frequency of S. det(The frame rate of a typical video is about 30fps. The downsampling frequency should not be too low or too high. If the downsampling is too low, almost every frame of the video needs to be processed, which is unnecessary for recognizing pet defecation behavior, as this behavior does not change as rapidly as running. If the downsampling is too high, the behavior may not be captured. The downsampling frequency S of this invention...) det As a variable parameter, the setting range is 5~15fps, and the continuous sampling time T det Sampling time T det Based on the identified actions and behaviors, this invention, specifically for pet excrement recognition applications, sets 5≤T. det ≤10s, obtain the total number of frames F sa =S det ×T det Store the sampled frames into a first-in-first-out (FIFO) image queue {I} orig_i}, where i represents the i-th image in the image queue, and when the queue length reaches F. sa When {I}, proceed to step 2, and simultaneously update the queue using the sliding window method. orig_i} Length reaches F sa Then, the percentage R pops up. out The historical frames, and 1 / 4≤Rout≤1 / 2; Step 2: When the image queue {I} orig_i The length of} is equal to F sa At that time, for each frame image I in the queue orig_i Preprocessing is performed by scaling the image proportionally to the input size of the target detection model along its longer side. In this embodiment, the image size is 640x640. After padding with 128 pixels and normalizing, the image is input into the target detection model. The target detection model is a general open-set detection model, and the detection label is limited to pet type, requiring no customized training. The pet type is selected as dog. The output consists of N bounding boxes for pet targets {pet}. i_j}, pet i_j For image queue {I orig_i The i-th image frame I orig_i The j-th pet was detected, where N is the total number of pets, and j ≤ N; Step 3: Transfer the image queue {I} orig_i} and the detected pet target bounding box {pet i_j Input a multi-target tracker (Tracker), which tracks multiple targets simultaneously and generates historical tracking trajectories. j}, j≤N, each trajectory Tra j Indicates the target pet 1_jThe coordinates of continuous bounding boxes in the queue; the multi-object tracker is based on the track-by-detection method, which does not require training a dedicated detection model. It supports dynamic tracking of multiple objects by associating trajectories with detection results. Step 4: For each trajectory Tra j Perform filtering, with a length equal to F. sa Then retain it and calculate a rectangular ROI region. j Make the rectangular ROI j The region can contain historical tracking tracks. j For each bounding box in the range, ensure that the rectangular ROI is defined. j The region does not exceed the original image frame I orig_i Under the boundary conditions; sequentially in the image frame queue {I orig_i In each frame of the image, a rectangular region of interest (ROI) of size} is extracted. j Sub-images of the region are used to construct a pet-targeted image. 1_j The length is F sa Continuous image queue { I pet_j_i After drawing the trajectory bounding box, proceed according to the sampling parameter S. act, For a continuous image queue { I pet_j_i} The continuous image queue {I} input to the action behavior recognition model is obtained by downsampling. act_j_k}, and 1 / 5 ≤ S act ≤1 / 2; ROI j The extraction rule is: based on the trajectory Tra j The maximum left, right, top, and bottom coordinates of all bounding boxes in the ROI determine the rectangular region; if the ROI j If the image exceeds the original image boundary, it will be truncated and corrected according to the boundary. Step 5: For {I act_j_k The image queue { I} is preprocessed to the input size of the action behavior recognition model by scaling and padding the images proportionally to their longer sides. act_obj_j_k}, preprocessed image queue { I act_obj_j_k After normalization, the data is fed into the action / behavior recognition model. act Action and behavior recognition model act The inference outputs a structured result, which includes location information, confidence level, and a judgment on whether a pet has defecated or urinated; the action behavior recognition model inputs an image queue {I}. act_j_k The sampling parameters S of} act To reduce the frame rate, such as retaining 1 / 3 of the frame when Sact=1 / 3, with an input size of 896×896, this invention generally uses a multiple of 224x224 as the final scaling size.

[0020] This invention processes multi-pet scenarios in parallel and outputs structured results in JSON format. The structured results output includes an analysis window {I}. orig_i The system includes the location information, confidence level, and determination of whether each identified pet is defecating or urinating.

[0021] The system structure used in this invention is as follows: Figure 1 As shown, 1 represents a surveillance camera, which can be any surveillance camera in public scenes such as communities, streets, and shopping malls; 2 represents the video stream, obtained from the surveillance camera or its corresponding network video storage device (NVR) through video software, such as video system software that supports mainstream video transmission protocols RTSP / RTMP and national standard GB / T 28181; video stream 2 is decoded into image queue 3. The image queue is obtained as follows: the decoded continuous video image frame data is downsampled at a sampling frequency of S. det downsampling frequency S det As a variable parameter, the setting range is 5~15fps; continuous sampling time T det If the range is set to 5-10 seconds, then the total number of frames sampled is F. sa =S det x T det For continuously input video image frames, each sampled frame is placed into a first-in-first-out (FIFO) image queue {I}. orig_i When the number of sampled frames reaches the total number of frames F sa Then, that is, the image queue {I orig_i} Length equals F sa , will the image queue {I orig_i The image queue is then processed in the next step of algorithm analysis 4. A sliding window method is used to update the image queue {I}. orig_i}, that is, when {I orig_i} Length reaches F sa Then, the percentage R pops up. out Historical image frames, R out This is a variable parameter, with a range of 1 / 4 to 1 / 2. Algorithm Analysis Module 4 accepts image queue 3 data {I orig_iAfter analyzing and processing this batch of data, the analysis results are obtained and output in a structured form. The structured output uses the common JSON format for easy integration with other business systems. An example of the structured output is: {"has_pet":true, "num_pets":2, "pet_status":[{"pet":[12,50, 112, 150], "status":false, "confidence":0.8}, {"pet":[200, 45, 400, 145],"status":true, "confidence":0.9}]}, where the field "has_pet" represents image I. orig The system checks whether pets exist. The field "num_pets" represents the number of pets, and "pet_status" represents the pet information data detected. It is a list, and each item in the list consists of the fields "pet" (location information), "status" (status information, true indicates that the pet is defecating, false indicates that the pet is not defecating), and "confidence" (confidence information).

[0022] Algorithm Analysis 4 Modules Figure 2 As shown, the original image queue data {I orig_i First, the data undergoes preprocessing using the object detection model 6. The preprocessing method involves dividing the queue data {I} into... orig_i Each frame I in} orig_i The image is scaled proportionally to the longer side to the object detection model. det The network input layer size is 640x640, and padding is applied with a padding pixel value of 128, resulting in an image I with a size equal to the input size of the object detection model network. obj_i After normalization, it becomes a new image sequence {I}. obj_i Object detection model obj 7. Inference received image sequence {I} obj_i}, employing parallel inference, outputs the detected multiple targets, assuming there are N targets, {pet i_j}, pet i_j For the original image frame I orig_i The j-th pet is detected, and its bounding box, pet, is represented by the top-left and bottom-right points. i_j ={x i_j_min y i_j_min x i_j_max y i_j_max}, j≤N, where coordinates {x i_j_min yi_j_min x i_j_max y i_j_max} is from image I obj_i Coordinate system restored to the original image I orig_i Coordinates in a coordinate system; Object detection model obj The GroundingDINO open-world object detection model is used, with the target label set to pet type, such as "dog". Based on the detected multiple targets {pet}... i_j}, employing a multi-target tracking method on the original image queue data {I orig_i Continuous tracking is performed on the target. The multi-target tracking method is based on track-by-detection and uses the ByteTrack algorithm to implement a multi-target tracker, TRacker. The multi-target tracker TRacker tracks multiple targets simultaneously and records multiple historical tracking trajectories {Tra}. j}, j≤N, where Tra j It represents pets. 1_j The corresponding continuous tracking trajectory. When initializing the tracker, the selected initial tracking target is the image frame queue {I}. orig_i The first frame I in} orig_1 All target pets 1_j Historical tracking trajectory {Tra j Each trajectory Tra in} j Is the target pet 1_j In the continuous image frame queue {I orig_i The location information of} is represented using bounding boxes, i.e., Tra j ={x j_min_i y j_min_i x j_max_i y j_max_i}, where (x j_min_i y j_min_i x j_max_i y j_max_i The bounding box coordinates of the j-th trajectory in the i-th frame's image coordinate system are represented by ). Historical tracking trajectories {Tra j} Input multi-target ROI extraction 9, output multiple consecutive image queues { I act_j_k}. Target image queue { I act_j_k The data is preprocessed by the action recognition model 10, and each image I... act_j_k Image preprocessing is performed by scaling the image proportionally to its longer side and padding it with 128 pixels, resulting in a square image suitable for the action recognition model. actThe size of the network input layer is set to 896x896, and then normalized to obtain the image queue { I act_obj_j_k Image queue { I act_obj_j_k} Input into the action recognition model act 11. Reasoning. Action Recognition Model act 11. It adopts the general-purpose Vision-Language-Model (VLM) Qwen2.5-VL-3B-Instruct, requiring no data collection or annotation, and no fine-tuning training. (Action recognition model) act After 11 inferences, the final output is a structured result {R}. orig}

[0023] Multi-target ROI extraction process 9 as follows Figure 3 As shown, the input raw image queue {I orig_i} and historical tracking trajectory {Tra j}, take a historical tracking trajectory Tra j Determine whether the length of this trajectory is equal to F. sa If not equal to F sa Then delete the tracking trajectory Tra j That is, abandoning the corresponding target pet. 1_j If equal to F sa Then keep this target pet. 1_j Find a rectangular region of interest (ROI). j Make the rectangular ROI j The region can contain historical tracking tracks. j Each target box (i.e., the target pet) in the [database / database] 1_j In the image frame queue {I orig_i (Continuous tracking boxes in}). Validation rectangle ROI. j Does the region exceed the original image frame {I orig_i Under the boundary conditions, if there is an area exceeding the boundary, then the ROI will be... j Maximum and minimum coordinates according to I orig_i Boundary values ​​are truncated to complete the correction. Based on ROI... j In sequence in the image frame queue {I orig_i In each frame of the image, a rectangular region of interest (ROI) of size} is extracted. j Sub-images of the region are used to construct a target pet. 1_j The length is F sa Continuous image queue { I pet_j_i Then, the continuous trajectory Tra... j The tracking bounding box, in red, is drawn in the continuous image queue { I pet_j_iOn, that is, the target pet 1_j The red bounding box in the consecutive image queue { I pet_j_i The data was marked. A controllable sampling parameter S was used. act Further analysis of the continuous image queue { I pet_j_i} Perform downsampling, S act The value range is 1 / 5-1 / 2, resulting in a continuous image queue { I act_j_k Repeat the above process until the traversal of {Tra} is complete. j}

[0024] It will be readily understood by those skilled in the art that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, combinations, substitutions, improvements, etc., made under the spirit and principles of the present invention are included within the protection scope of the present invention.

Claims

1. A dynamic pet excretion identification method based on hybrid cascade, characterized in that Comprising the following steps: Step 1: Downsample the input continuous video image frames at a sampling frequency of S. det, And 5≤S det ≤15fps, continuous sampling time T det And 5≤T det ≤10s, obtain the total number of frames F sa =S det ×T det ; The sampling frame is stored in a first-in first-out (FIFO) image queue {I orig_i}, i represents the i-th in the image queue, when the queue length reaches F sa , step 2 is entered, and the queue is updated by using a sliding window method, the historical frames occupying R out are popped out, and 1 / 4≤R out ≤1 / 2. Step 2: For each frame of image I in the queue orig_i Preprocessing is performed by scaling the image proportionally to the input size of the object detection model along its longer side, padding with 128 pixels, normalizing the image, and then inputting it into the object detection model. det Output N bounding boxes for the pet targets {pet} i_j }, pet i_j For image queue {I orig_i The i-th image frame I orig_i The j-th pet was detected, where N is the total number of pets, and j ≤ N; Step 3: input the image queue {I orig_i} and the detected pet target bounding box {pet i_j} into the multi-target tracker TRacker, which simultaneously tracks multiple targets to generate historical tracking trajectories {Tra j}, j≤N, each trajectory Tra j represents a pet target pet 1_j the continuous bounding box coordinates in the queue; Step 4: For each trajectory Tra j perform a filter, length equals F sa then keep, find a rectangular ROI region ROI j such that the rectangular ROI j region can contain each bounding box in the historical tracking trajectory Tra j and ensure that the rectangular ROI j region does not exceed the boundary conditions of the original image frame I orig_i . Each frame image in the image frame queue {I orig_i} is intercepted in sequence to obtain a sub-image with a rectangular ROI j region, and a continuous image queue {I 1_j} with a length of F sa for the pet target pet is constructed. pet_j_i After the trajectory frame is drawn, the continuous image queue {I act,} is down-sampled according to a sampling parameter S pet_j_i to obtain a continuous image queue {I act_j_k} for input of the action behavior recognition model, and 1 / 5≤S act ≤1 / 2. Step 5: Scale and fill the image queue {I act_j_k} along the long edge to the input size of the action behavior recognition model by preprocessing act_obj_j_k}, and the preprocessed image queue {I act_obj_j_k} is sent to the action behavior recognition model Model act after normalization. act The inference outputs structured results, which include the position information, confidence information and whether the pet target is in the structured results.

2. The dynamic pet excretion identification method based on hybrid cascade according to claim 1, characterized in that: The target detection model in step 2 is a general open-set detection model, and the detection label is limited to pet type, without the need for customized training, and the pet type selected is dog.

3. The dynamic pet excretion identification method based on hybrid cascade according to claim 1, characterized in that: The multi-target tracker TRacker in step 3 is based on the track-by-detection method, without the need for training a special detection model, and through the association of detection results, it supports multi-target dynamic tracking.

4. The dynamic pet excretion identification method based on hybrid cascade according to claim 1, characterized in that: ROI in step 4 j The extraction rule is: based on the trajectory Tra j The maximum left, right, top, and bottom coordinates of all bounding boxes in the ROI determine the rectangular region; if the ROI j If the image exceeds the original image boundary, it will be truncated and corrected according to the boundary.

5. The dynamic pet waste identification method based on hybrid cascade according to claim 1, characterized in that: The action behavior recognition model in step 4 inputs the sampling parameters S of the image queue {I act_j_k} in step 3 act For reducing frame rate, the input size is 896x896.

6. The dynamic pet waste identification method based on hybrid cascade according to claim 1, characterized in that: Parallel processing multi-pet scene, output structured results in JSON format, structured results output includes location information, confidence information and whether in the result of the judgment of each recognized pet target in the analysis window {I orig_i}