Detection Method, Detection Device and Electronic Device for Pet Leashes

By identifying the position information of pets and people in video frames, calculating pixel distances and judging the accompanying relationship, the accuracy of pet rope detection is solved, and efficient and accurate pet rope recognition is achieved.

CN114283364BActive Publication Date: 2025-07-29ANHUI IFLYTEK INTELLIGENT SYST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111589028.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-07-29
Estimated Expiration
2041-12-23

AI Technical Summary

Technical Problem

In the prior art, the detection effect of pet ropes is poor, and it is difficult to accurately identify the rope state between pets and people.

Method used

By identifying the position information of the pet and the person from the target video frame, calculating their pixel distance, and determining whether there is a concomitant relationship in successive frames, image recognition technology is used to determine whether there is a bolt.

Benefits of technology

It improves the accuracy and adaptability of pet rope detection, and reduces the detection inaccuracy caused by three-dimensional distance calculation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283364B_ABST
    Figure CN114283364B_ABST
Patent Text Reader

Abstract

The present invention provides a detection method, a detection device and an electronic device for a pet leash. The detection method for the pet leash includes: identifying a target object from a target video frame of a target video, where the target object includes a first target object corresponding to a pet and a second target object corresponding to a person; determining a first pixel distance between the first target object and the second target object in the target video frame based on the position information of the first target object and the second target object in the target video frame; when the first pixel distances corresponding to a plurality of consecutive target video frames in the target video are all less than a first target threshold, determining the first target object and the second target object as a target pair; and detecting whether the first target object is leashed based on the target pair. The detection method for the pet leash of the present invention can effectively improve the accuracy of the detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a pet leash detection method, a detection device, and an electronic device. Background Art

[0002] As living standards improve, more and more people are keeping pets. At the same time, to ensure the normal work and lives of others, various regions have begun to strengthen the detection of whether pets are leashed. Related technologies mainly use cameras for detection, but due to the small size of the leash, this detection method is difficult and has poor detection results. Summary of the invention

[0003] The present invention provides a pet leash detection method, a detection device, and an electronic device, which are used to solve the defect of poor pet leash detection effect in the prior art and achieve efficient and accurate pet leash detection.

[0004] The present invention provides a pet leash detection method, comprising:

[0005] identifying target objects from a target video frame of a target video, the target objects including a first target object corresponding to a pet and a second target object corresponding to a person;

[0006] determining a first pixel distance between the first target object and the second target object in the target video frame based on position information of the first target object and the second target object in the target video frame;

[0007] When first pixel distances corresponding to a plurality of consecutive target video frames in the target video are all less than a first target threshold, determining the first target object and the second target object as a target pair;

[0008] Based on the target pair, it is detected whether the first target object is tethered.

[0009] According to a pet leash detection method provided by the present invention, detecting whether the first target object is leashed based on the target pair includes:

[0010] Based on the target pair, performing image extraction on the target video frame to generate a target image block, where the target image block includes at least a portion of a pixel area between the first target object and the second target object;

[0011] Image recognition is performed on the target image block to detect whether the first target object is tied with a rope.

[0012] A method for detecting a pet leash according to the present invention, extracting an image from the target video frame based on the target pair to generate a target image block, including:

[0013] Based on the target pair, determining a target matte region corresponding to the target video frame;

[0014] Extracting an image from the target video frame based on the target matte region to generate the target image block.

[0015] A method for detecting a pet leash according to the present invention, the target video includes consecutive first video frames and second video frames, identifying a target object from the target video frames of the target video, including:

[0016] Performing feature extraction on the first video frame and the second video frame respectively to generate a first candidate object feature corresponding to the first video frame and a second candidate object feature corresponding to the second video frame;

[0017] Generating a second pixel distance based on the position information of the first candidate object feature in the first video frame and the position information of the second candidate object feature in the second video frame;

[0018] When the similarity between the first candidate object feature and the second candidate object feature is greater than a second target threshold and the second pixel distance is less than a third target threshold, determining the first candidate object feature and the second candidate object feature as the same target object.

[0019] A method for detecting a pet leash according to the present invention, identifying a target object from the target video frames of the target video, including:

[0020] Labeling the target object to generate labeling information, the labeling information including at least one of: an identifier of the target object, a feature of the target object in the target video frame, and position information of the target object in the target video frame.

[0021] A method for detecting a pet leash according to the present invention, after detecting whether the first target object is leashed based on the target pair, the method further includes:

[0022] When it is determined that the first target object is not leashed, outputting an alarm message.

[0023] The present invention also provides a pet leash detection device, including:

[0024] A first recognition module, configured to recognize a target object from a target video frame of a target video, where the target object includes a first target object corresponding to a pet and a second target object corresponding to a person;

[0025] A first determination module, configured to determine a first pixel distance between the first target object and the second target object in the target video frame based on position information of the first target object and the second target object in the target video frame;

[0026] A second determination module, configured to determine the first target object and the second target object as a target pair when first pixel distances corresponding to a plurality of consecutive target video frames in the target video are all less than a first target threshold;

[0027] A first detection module, configured to detect whether the first target object is on a leash based on the target pair.

[0028] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for detecting whether a pet is on a leash as described in any one of the above are implemented.

[0029] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for detecting whether a pet is on a leash as described in any one of the above are implemented.

[0030] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method for detecting whether a pet is on a leash as described in any one of the above are implemented.

[0031] The method, device, and electronic device for detecting whether a pet is on a leash provided by the present invention determine whether there is an accompanying relationship between the first target object and the second target object through the position trajectory between the first target object and the second target object. When it is determined that there is an accompanying relationship between the two, based on the target pair with the accompanying relationship, it is recognized whether the first target object in the target video frame is on a leash, with high accuracy and strong adaptability. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1It is one of the flow schematic diagrams of the detection method for pet leashes provided by the present invention;

[0034] Figure 2 It is the second of the flow schematic diagrams of the detection method for pet leashes provided by the present invention;

[0035] Figure 3 It is the structural schematic diagram of the detection device for pet leashes provided by the present invention;

[0036] Figure 4 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0037] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0038] Below in conjunction with Figures 1 to 2 Describe the detection method for pet leashes of the present invention.

[0039] The execution subject of this detection method for pet leashes can be a detection device for pet leashes, or it can be a server, or it can also be a staff member's terminal. Among them, the terminal includes but is not limited to: mobile phones, computers, tablet computers, watches, and other mobile or non-mobile terminals.

[0040] As Figure 1 shown, this detection method for pet leashes includes: step 110, step 120, step 130, and step 140.

[0041] Step 110: Identify target objects from the target video frames of the target video. The target objects include a first target object corresponding to a pet and a second target object corresponding to a person;

[0042] In this step, the target video is video data for detection.

[0043] The target video can be real-time video data collected by an image sensor; or it can be a video recorded and stored in a database; or it can be video data pulled from the network; or it can also be data obtained through other means, which is not limited in the present invention.

[0044] The target video frame can be any video frame in the target video, and both the first target object and the second target object are included in the target video frame.

[0045] Among them, the first target object is an object corresponding to a pet, for example, it can be a cat, a dog or other pets.

[0046] The second target object is an object corresponding to a person. It can be understood that the second target object described in the present invention can be an object having an accompanying relationship with the first target object. For example, the second target object can be the owner of the first target object; or the second target object can also be an object having no accompanying relationship with the first target object, such as the second target object being other passers-by.

[0047] In the actual execution process, a separately trained detection network can be used to identify the target object.

[0048] The input of the detection network is a single-frame target video frame, and the output is the recognition result.

[0049] Among them, the recognition result is used to represent whether the target video frame includes the first target object or the second target object.

[0050] It can be understood that the detection network can be a pre-trained neural network, where the neural network is trained with sample image frames as sample information and the corresponding recognition results as sample labels.

[0051] For example, in the actual execution process, any single-frame target video frame in the target video can be input into the detection network. For example, the first video frame in the target video is input into the detection network, and the detection network outputs the recognition result.

[0052] Among them, the recognition result can include: the target video frame includes a pet and does not include a person, or the target video frame includes a person and does not include a pet, or the target video frame includes a person and a pet, or the target video frame includes neither a person nor a pet.

[0053] In the embodiment of the present invention, only when the first target object corresponding to the pet and the second target object corresponding to the person are simultaneously recognized from the same target video frame of the target video, step 120 is executed.

[0054] In some embodiments, if only the first target object corresponding to the pet is recognized from the video frames of consecutive frames of the target video and the second target object corresponding to the person is not recognized, an alarm message is output.

[0055] The alarm message prompts the user that the pet is not leashed.

[0056] Among them, the number of frames of the video frames of consecutive frames can be user-defined. For example, it is set that when there is only a pet and no person in consecutive 3 frames or 5 frames of video frames, an alarm message of not leashed is output.

[0057] In some other embodiments, if only the second target object corresponding to a person is recognized in any video frame of the target video and the first target object corresponding to a pet is not recognized, the subsequent execution steps are not triggered.

[0058] In some embodiments, step 110 may further include: annotating the target object to generate annotation information, where the annotation information includes: the identifier of the target object, the features of the target object in the target video frame, and the position information of the target object in the target video frame.

[0059] In this embodiment, the annotation information is used to characterize the position information of the target object in the target video frame and the category of the target object.

[0060] The annotation information can be cached in a local or cloud database and retrieved when needed.

[0061] Wherein, the annotation information may include: a feature vector for characterizing the feature information of the target object (for example, a 1024-dimensional feature vector), the position information (x, y) of the pixel center point of the target object, and the size (w, h) of the detection region corresponding to the target object.

[0062] The size of the detection region corresponding to the target object is the region where the target object is located, and the detection region corresponding to the target object should include the pixel region corresponding to the target object.

[0063] It can be understood that each target object in each target video frame corresponds to annotation information.

[0064] There may be multiple target objects in the same target video frame, so the same target video frame may correspond to multiple annotation information.

[0065] The same target object may exist in multiple target video frames, so the same target object may correspond to annotation information in multiple target video frames.

[0066] In this embodiment, by recognizing the target video frame to generate a recognition result, in the case where the recognition result is that the target video frame includes a person and a pet, the recognition network can further output the size of the detection boxes corresponding to the person and the pet in the target video frame and the position information of the person and the pet in the target video frame.

[0067] That is to say, the recognition result may further include: whether the target video frame includes the first target object or the second target object, the position information of the target object in the target video frame, and the size information of the detection box corresponding to the target object, etc.

[0068] Among them, the detection box corresponding to the target object is the detection area corresponding to the target object. The detection box corresponding to the target object is used to label the target object, and the shape of the detection box can be user-defined.

[0069] For example, the shape of the detection box can be set to a rectangular box, an oval box, or any other shape, which is not limited in the present invention.

[0070] Taking the detection box as a rectangular box as an example, the rectangular box can be the minimum bounding rectangle of the target object. Specifically, when labeling the first target object, the detection box corresponding to the first target object is the minimum bounding rectangle of the first target object; when labeling the second target object, the detection box corresponding to the second target object is the minimum bounding rectangle of the second target object.

[0071] Based on the rectangular box, annotation information can be generated, which includes: the center point position information (x, y) of the target object, the length and width information (w, h) of the target rectangular box, and the category c of the target object.

[0072] In the actual execution process, any deep learning detection method can be used to execute this step. Taking yolov5 as an example, a convolutional network can be used to extract features from the target video frame first, generate multiple features, and then perform regression of position information, prediction of whether there is a target object, and prediction of the target category on the extracted features. The finally output annotation information includes: how many target objects are there in the target video frame, the position information of each target object, and the size and category of the target object.

[0073] It should be noted that in this step, the target objects in multiple consecutive target video frames can be labeled to generate annotation information corresponding to multiple consecutive target video frames.

[0074] For example, detect people and dogs in 15 consecutive video frames, and save the feature information of the detected target objects, the midpoint position information of the target objects, and the size information of the detection areas corresponding to the target objects to generate annotation information of people and dogs in 15 consecutive video frames. The annotation information includes the position sequence information of people and dogs in 15 consecutive video frames.

[0075] In some embodiments, there can be multiple second target objects.

[0076] For example, in the case of identifying the first target object and P second target objects from the target video frames of the target video, label the first target object and each second target object respectively to generate the position information of the first target object and each second target object in the same target video frame.

[0077] For example, detect a dog and multiple people in 15 consecutive video frames, and save the feature information of the detected target objects, the midpoint position information of the target objects, and the size information of the detection regions corresponding to the target objects, so as to generate annotation information of the dog and multiple people in 15 consecutive video frames. The annotation information includes the position sequence information of the dog and each person in 15 consecutive video frames.

[0078] In some embodiments, the target video includes consecutive first and second video frames. Step 110 may include:

[0079] Extract features from the first video frame and the second video frame respectively to generate first candidate object features corresponding to the first video frame and second candidate object features corresponding to the second video frame;

[0080] Generate a second pixel distance based on the position information of the first candidate object features in the first video frame and the position information of the second candidate object features in the second video frame;

[0081] When the similarity between the first candidate object features and the second candidate object features is greater than a second target threshold and the second pixel distance is less than a third target threshold, determine the first candidate object features and the second candidate object features as the same target object.

[0082] In this embodiment, the first candidate object features may be the features of any object in the first video frame. For example, they may be the object features corresponding to a pet in the first video frame, or they may also be the object features corresponding to a person in the first video frame.

[0083] The second candidate object features may be the features of any object in the second video frame. For example, they may be the object features corresponding to a pet in the second video frame, or they may also be the object features corresponding to a person in the second video frame.

[0084] It can be understood that one or more first candidate object features can be extracted from the first video frame; one or more second candidate object features can be extracted from the second video frame.

[0085] The second pixel distance is the horizontal pixel distance between the position of the first candidate object features in the first video frame and the position of the second candidate object features in the second video frame.

[0086] It should be noted that the first candidate object features and the second candidate object features may be the features corresponding to the same object in the target video, or may be the features corresponding to different objects in the target video.

[0087] After the first candidate object feature and the second candidate object feature are recognized, the position information of the first candidate object feature in the first video frame is respectively marked to generate third position information; the position information of the second candidate object feature in the second video frame is marked to generate fourth position information.

[0088] Then, calculate the horizontal pixel difference between the third position information and the fourth position information to obtain the second pixel distance Dis.

[0089] At the same time, compare the similarity Sim between the first candidate object feature and the second candidate object feature.

[0090] Among them, the second target threshold is used to judge the similarity degree between the first candidate object feature and the second candidate object feature.

[0091] The second target threshold can be user-defined, and different numerical second target thresholds can be defined in different scenarios.

[0092] It can be understood that when the similarity Sim between the first candidate object feature and the second candidate object feature is greater than the second target threshold, it can be approximately considered that the first candidate object feature and the second candidate object feature are the features corresponding to the same object in the target video; when the similarity Sim is not greater than the second target threshold, it is considered that the first candidate object feature and the second candidate object feature are the features corresponding to different objects in the target video.

[0093] The third target threshold is used to judge the distance between the first candidate object feature and the second candidate object feature in two consecutive video frames.

[0094] The third target threshold can be user-defined.

[0095] For example, the third target threshold can be set to half of the diagonal length of the rectangular box corresponding to the candidate object.

[0096] It can be understood that in consecutive video frames, the position change of the same object will not be too large. Then, when the second pixel distance Dis is less than the third target threshold, it can be approximately considered that the first candidate object feature and the second candidate object feature are the features corresponding to the same object in the target video; when the second pixel distance Dis is not less than the third target threshold, it is considered that the first candidate object feature and the second candidate object feature are the features corresponding to different objects in the target video.

[0097] In the actual execution process, by recognizing two consecutive frames of the target video, the first candidate object feature is recognized from the first video frame, and the first candidate object feature is marked to obtain the annotation information corresponding to the first candidate object feature.

[0098] The annotation information corresponding to the first candidate object feature includes, but is not limited to: the ID of the first candidate object feature, the feature of the first candidate object feature in the first video frame, and the position information of the first candidate object feature in the first video frame.

[0099] Cache the annotation information corresponding to the first candidate object feature.

[0100] Then, identify a second candidate object feature from the second video frame, and annotate the second candidate object feature to obtain the annotation information corresponding to the second candidate object feature.

[0101] Among them, when performing annotation, the similarity Sim between the second candidate object feature and the first candidate object feature can be compared with a second target threshold first, and the second pixel distance Dis between the second candidate object feature and the first candidate object feature can be compared with a third target threshold.

[0102] If the similarity Sim between the second candidate object feature and the first candidate object feature is greater than the second target threshold, and the second pixel distance Dis is less than the third target threshold, it is determined that the second candidate object feature and the first candidate object feature match successfully.

[0103] In the case of a successful match, keep the ID of the second candidate object feature unchanged, that is, keep the ID of the second candidate object feature the same as the ID of the first candidate object feature, and replace the feature and position information of the second candidate object feature in the cache with the feature and position information of the second candidate object feature in the second video frame, so as to generate the annotation information corresponding to the second candidate object feature.

[0104] If the similarity Sim between the second candidate object feature and the first candidate object feature is not greater than the second target threshold; or if the second pixel distance Dis is not less than the third target threshold; or if the similarity Sim between the second candidate object feature and the first candidate object feature is not greater than the second target threshold and the second pixel distance Dis is not less than the third target threshold, it is determined that the second candidate object feature and the first candidate object feature match fails.

[0105] In the case of a failed match, list the second candidate object feature as an observation object, and use the target video frame as the second frame, and match it with the candidate object features in subsequent consecutive video frames.

[0106] For the second frame, keep the position information of the observation object unchanged. For the third frame and subsequent frames, make the position of the observation object in the corresponding video frame be the position of the previous frame plus the position difference between the previous two frames. For example, it can be through the formula:

[0107]

[0108] Determine the position information of the candidate object feature in the current target video frame, where (X, Y) is the position information of the candidate object feature in the current target video frame, (X', Y') is the position information of the candidate object feature in the previous target video frame, and (X'', Y'') is the position information of the candidate object feature in the target video frame two frames before.

[0109] In the case where the observed object remains the observed object in multiple consecutive frames, it is determined that it disappears from the screen and is deleted from the cache.

[0110] In the case where the observed object matches other candidate object features in multiple consecutive frames, it is determined that there is another target object in the target video frame, and the observed object is saved to the cache and a new ID is assigned.

[0111] According to the pet leash detection method provided by the embodiments of the present invention, by judging the similarity and the second pixel distance between the first candidate object feature and the second candidate object feature in consecutive target video frames, it is determined whether the first candidate object feature and the second candidate object feature are the same object in the target video, which can avoid the judgment error that occurs under a single judgment condition, significantly improve the accuracy of the judgment result, and thus help improve the accuracy of the subsequent detection result.

[0112] Step 120: Based on the position information of the first target object and the second target object in the target video frame, determine the first pixel distance between the first target object and the second target object in the target video frame;

[0113] In this step, the first pixel distance is used to represent the horizontal pixel distance between the pixel regions of the first target object and the second target object in the same target video frame.

[0114] It can be understood that in the actual execution process, the three-dimensional distance cannot be obtained from the two-dimensional image, so the horizontal pixel distance is used to represent the distance between the first target object and the second target object.

[0115] For example, in the actual execution process, 15 consecutive target video frames can be used for judgment. Then, 15 pieces of annotation information corresponding to the 15 target video frames can be obtained from step 110. Based on the 15 pieces of annotation information, the first pixel distance corresponding to each target video frame can be generated.

[0116] It should be noted that in some embodiments, in the case where there are multiple second target objects, the first pixel distance between the first target object and each second target object in the same target video frame is generated based on the position information of the first target object and each second target object in the same target video frame.

[0117] For example, when the number of second target objects is P, for the same frame of target video frame, P first pixel distances can be determined, where each first pixel distance corresponds to a second target object respectively.

[0118] Step 130: When the first pixel distances corresponding to consecutive multiple target video frames in the target video are all less than the first target threshold, determine the first target object and the second target object as a target pair;

[0119] Such as Figure 2 As shown, in this step, the number of consecutive multiple target video frames can be user-defined, for example, set to 15 consecutive frames or 20 consecutive frames, etc., which is not limited in the present invention.

[0120] Each frame of the consecutive multiple target video frames needs to include the first target object and the second target object.

[0121] It should be noted that for the nodes in the position sequence information, when its length is less than the number of consecutive multiple target video frames set, the two ends of the position sequence information are filled with the outermost position information.

[0122] The first target threshold is used to determine whether the first target object and the second target object are in an accompanying state.

[0123] The first target threshold can be user-defined.

[0124] It should be noted that if the first target threshold represents the maximum distance between a pet and the pet's owner, the first target threshold can be determined based on the length of the pet leash.

[0125] It can be understood that for pet leashes, the length of the leashes on the market is generally between 2 - 4 meters, then the distance between the pet and the pet owner is generally less than 2 meters.

[0126] The width of a person is approximately between 20 cm - 50 cm. Using 35 cm as the standard, the length of the leash is approximately between 4 - 10 times the width of a person.

[0127] Converting the three-dimensional distance to two-dimensional pixels, the first pixel distance between the first target object and the second target object should be between 4 - 10 times the width of the detection frame corresponding to the second target object.

[0128] If 35 cm is used as the standard for the width of a person, the first target threshold can be set to 6 times the width of the detection frame corresponding to the second target object. Of course, in other embodiments, the first target threshold can also be set to other values, which is not limited in the present invention.

[0129] The target pair is used to characterize the coexistence relationship between the first target object and the second target object, that is, the second target object may be the owner of the first target object.

[0130] By comparing the first pixel distance corresponding to each target video frame in a series of consecutive target video frames with a first target threshold, if the first pixel distance corresponding to each target video frame in the series of consecutive target video frames is less than the first target threshold, then the first target object and the second target object corresponding to the first pixel distance are determined as the target pair.

[0131] During the actual execution process, the formula:

[0132]

[0133] is used to obtain the maximum first pixel distance between the first target object and the second target object in multiple target video frames, where represents the maximum first pixel distance, is the position of the first target object in the i-th target video frame, is the position of the second target object in the i-th target video frame, and S is used to represent belonging to the same target video frame.

[0134] In when it is less than the first target threshold, it is determined that the first pixel distances in the series of consecutive target video frames are all less than the first target threshold, and then the first target object and the second target object corresponding to the first pixel distance are determined as the target pair.

[0135] It can be understood that when there are multiple second target objects, by the above method, the maximum first pixel distance between each second target object and the first target object in multiple target video frames can be calculated respectively, and then by comparing the relationship between the maximum first pixel distance and the first target threshold, it is determined whether the second target object is associated with the first target object.

[0136] For example, when there are P second target objects, the distances between the position sequence information of the first target object and the P second target objects are respectively judged to obtain the corresponding to the P second target objects. By comparing with the first target threshold, it is determined that the corresponding to Q second target objects is less than the first target threshold, and then each of the Q second target objects is matched with the first target object to obtain Q target pairs, where P≥Q>0.

[0137] In this embodiment, after determining the target pair, step 140 is executed.

[0138] In the R & D process, the inventor found that the pet leash is a flexible object with a small volume. In the related technology, direct detection through video images has relatively low accuracy for detecting the state of the leash.

[0139] The inventor also found in the R & D process that an image is a two-dimensional plane, while the actual scene is a three-dimensional scene. In the related technology, estimating the three-dimensional distance between a person and a pet in a video image through the video image has relatively low accuracy, resulting in inaccurate detection results.

[0140] In the present invention, the first pixel distance between a person and a pet is determined based on the position information of the person and the pet in the same target video frame, and whether the person and the dog have a following relationship is judged through the trajectory feature formed by the first pixel distances corresponding to consecutive multiple target video frames. In the case of judging that there is a following relationship, it is further identified whether the leash is attached, which can effectively solve the above two problems from the side. It can not only avoid inaccurate detection results caused by errors in three-dimensional distance measurement, but also improve the accuracy and precision of pet leash recognition, thereby improving the accuracy of the recognition result.

[0141] According to this step, the first pixel distance between the first target object and the second target object in the same target video frame is compared with the first target threshold to judge whether there is a following relationship between the two. In the case of determining that there is a following relationship between the two, the first target object is determined as the pet of the second target object, and the two are determined as a target pair, which helps to improve the accuracy of the recognition result in the subsequent recognition process.

[0142] In some other embodiments, in the case where no target pair is recognized through consecutive multiple target video frames, that is, in the case where the first pixel distance between the second target object and the first target object in at least one target video frame among consecutive multiple target video frames is greater than the first target threshold, it is determined that there is no following relationship between the first target object and the second target object, and an alarm message is output to prompt that the pet is not on a leash.

[0143] Step 140: Based on the target pair, detect whether the first target object is on a leash.

[0144] In this step, after generating the target pair, the target pair can be extracted from the target video frame, and image recognition is performed on the extracted target pair to judge whether the leash is attached.

[0145] It can be understood that since there is a following state between the first target object and the second target object in the same target pair, it can be approximately considered that the first target object is the pet of the second target object.

[0146] In the actual execution process, based on the target pairs with a following relationship, image recognition is performed on the pixel region where the target pair is located to further determine whether the pixel region where the target pair is located includes a leash, so as to determine whether the first target object is leashed.

[0147] When the target pair is a single pair, based on this target pair, it is determined whether the pixel region between the first target object and the second target object in the target pair includes a leash.

[0148] When there are multiple target pairs, based on each pair of target pairs respectively, it is determined whether the pixel region between the first target object and the second target object in each target pair includes a leash.

[0149] For example, when two first target objects, a cat and a dog, are recognized from the target video, and three people, A, B, and C, are recognized from the target video, and three target pairs, "A - dog", "B - dog", and "C - cat", are determined through step 130, then based on the "A - dog" target pair, it is detected whether the pixel region between A and the dog includes a leash to detect whether A leashes the dog; based on the "B - dog" target pair, it is detected whether the pixel region between B and the dog includes a leash to detect whether B leashes the dog; based on the "C - cat" target pair, it is detected whether the pixel region between C and the cat includes a leash to detect whether C leashes the cat.

[0150] In some embodiments, step 140 may include:

[0151] Based on the target pair, image extraction is performed on the target video frame to generate a target image block, and the target image block includes at least part of the pixel region between the first target object and the second target object;

[0152] Image recognition is performed on the target image block to detect whether the first target object is leashed.

[0153] In this embodiment, the target image block is extracted from the target video frame and includes the image features of at least part of the pixel region of the first target object, at least part of the pixel region of the second target object, and at least part of the pixel region between the first target object and the second target object.

[0154] It can be understood that the target image block has a corresponding relationship with the second target object.

[0155] That is, when there are multiple target pairs in the same target video frame, based on the second target object in each target pair, a target image block corresponding to the second target object can be extracted from the target video frame.

[0156] For example, when there are two target pairs in the target video frame, namely the "A - pet dog" target pair and the "B - pet dog" target pair, the target video frame is extracted based on the "A - pet dog" target pair to generate a target image block corresponding to A; the target video frame is extracted based on the "B - pet dog" target pair to generate a target image block corresponding to B.

[0157] In the actual execution process, based on the target pair, image extraction is performed on multiple target video frames in the target video, and target image blocks corresponding to the multiple target video frames can be obtained.

[0158] In some embodiments, based on the target pair, image extraction is performed on the target video frame to generate a target image block, including:

[0159] Based on the target pair, determine the target matte region corresponding to the target video frame;

[0160] Based on the target matte region, perform image extraction on the target video frame to generate a target image block.

[0161] In this embodiment, the target matte region is the region for generating the target image block.

[0162] The range of the target matte region is determined based on the position information of the first target object and the second target object in the target pair.

[0163] The shape of the target matte region can be user - defined, for example, set as a rectangular frame, an oval frame or other irregular - shaped regions, which is not limited in the present invention.

[0164] For example, when the target matte region is set as a rectangular frame, the diagonal point positions of the target matte region can be set respectively as: the x - coordinate of the first diagonal point is set as the horizontal mid - point position of the rectangular frame corresponding to the second target object, and the y - coordinate of the first diagonal point is set as the position at 1 / 3 from top to bottom longitudinally of the rectangular frame corresponding to the second target object; the x - coordinate of the second diagonal point is set as the horizontal mid - point position of the rectangular frame corresponding to the first target object, and the y - coordinate of the second diagonal point is set as the longitudinal mid - point position of the rectangular frame corresponding to the second target object.

[0165] It can be understood that the target matte region includes the pixel region of the pet's lower body, the pixel region of the person's lower body, and the intermediate pixel region between the person and the pet. In the case where the pet is on a leash, the target matte region will include the leash feature.

[0166] In the actual execution process, a pre - trained neural network model can be used to execute this step. Among them, the neural network model is trained with sample video frames as samples and the corresponding sample image blocks as sample labels.

[0167] After obtaining the target image block, image recognition can be performed on the target image block to detect whether the first target object in the target image block is tethered.

[0168] Similarly, in the actual execution process, a pre-trained neural network model can be used to execute this step. Among them, the neural network model is trained with sample image blocks as samples and the corresponding sample recognition results of the sample target image blocks as sample labels.

[0169] Among them, the sample image block can be the image block extracted from the sample video frame in the previous step.

[0170] According to the pet tether detection method provided by the embodiment of the present invention, it is judged whether there is an accompanying relationship between the first target object and the second target object through the position trajectories between the first target object and the second target object. When it is determined that there is an accompanying relationship between the two, based on the target pair with the accompanying relationship, it is identified whether the first target object in the target video frame is tethered, with high accuracy and strong adaptability.

[0171] In some embodiments, after step 140, the method may further include: outputting an alarm message when it is determined that the first target object is not tethered.

[0172] In this embodiment, when it is detected through step 140 that the first target object is not tethered, an alarm message is output.

[0173] The alarm message is used to prompt that the first target object is not tethered.

[0174] It should be noted that in the actual execution process, based on step 140, image recognition can be performed on each frame of the target video to detect whether the first target object is tethered; or, image recognition can also be performed on specific video frames in the target video to detect whether the first target object is tethered.

[0175] In some embodiments, the method may further include: outputting an alarm message when no tethering is recognized in the target video frames of the target number of frames with the shortest target first pixel distance.

[0176] In this embodiment, the target number of frames can be user-defined. For example, the target number of frames can be set to 5 frames or 7 frames, etc., and the present invention does not make a limitation.

[0177] In the actual execution process, continue to refer to Figure 2, the first pixel distance corresponding to each target pair in each frame of the target video can be calculated respectively, and the frame of the target video with the shortest distance among the multiple first pixel distances can be selected, and based on the target pair corresponding to the frame of the target video, image extraction is performed to generate the target image block corresponding to the target pair for the frame number of the target.

[0178] Then, image recognition is performed on the target image blocks for the frame number of the target respectively. When it is recognized that none of the target image blocks for the frame number of the target include a tether, it is determined that there is no tether between the second target object and the first target object corresponding to the target pair.

[0179] When there is one second target object, when it is recognized that none of the target image blocks for the frame number of the target include a tether, an alarm message is output.

[0180] When there is one second target object, when a tether is recognized in any one or more frames of the target image blocks, it is determined that the first target object is tethered.

[0181] In this embodiment, by detecting each frame of the target video with the shortest first pixel distance, when no tether is recognized in each frame of the target video, it is determined that there is no tether between the first target object and the second target object, significantly improving the accuracy of the detection result.

[0182] In some other embodiments, when there are multiple second target objects and multiple target pairs, for example, when there are P second target objects and Q second target pairs, where P≥Q>1, taking the frame number of the target as 5 frames as an example, this embodiment is specifically described.

[0183] In this embodiment, the first pixel distance between each second target object in each frame of the target video and the first target object can be calculated respectively, and the first pixel distance can be determined as the first pixel distance corresponding to the second target object in the frame of the target video.

[0184] For each second target object, 5 target video frames corresponding to the 5 shortest first pixel distances are selected respectively from the multiple first pixel distances corresponding to the second target object, and 5 frames of the target video corresponding to the second target object are obtained.

[0185] Then, based on the target pair corresponding to the second target object, image extraction is performed on the 5 frames of the target video corresponding to the second target object respectively to generate 5 target image blocks corresponding to the second target object.

[0186] Based on the above steps, a total of 5Q target image blocks corresponding to Q second target objects can be obtained.

[0187] Perform image recognition on 5Q target image patches respectively. In the case where at least one target image patch recognizes a leash, determine that the first target object has a leash.

[0188] In the case where no leash is recognized in all 5Q target image patches, determine that there is no leash relationship between the Q second target objects and the first target object, and then output an alarm message.

[0189] According to the pet leash detection method provided by the embodiments of the present invention, on the basis of recognizing each video frame in the target video, the target image patches corresponding to each pair of target pairs in the target video frame are also recognized to determine whether there is a leash between each second target object and the first target object in each target video frame, further improving the accuracy of the recognition process; in the case where a leash is recognized for any pair of target pairs in any target video frame, determine that the first target object has a leash; in the case where no leash is recognized for all target pairs in all target video frames, output an alarm message, significantly reducing the error generated by video judgment and improving the accuracy and accuracy of the recognition result.

[0190] Next, a pet leash detection device provided by the present invention will be described. The pet leash detection device described below can be correspondingly referred to the pet leash detection method described above.

[0191] As Figure 3 shown, the pet leash detection device includes: a first recognition module 310, a first determination module 320, a second determination module 330, and a first detection module 340.

[0192] The first recognition module 310 is configured to recognize target objects from the target video frames of the target video, where the target objects include a first target object corresponding to a pet and a second target object corresponding to a person;

[0193] The first determination module 320 is configured to determine a first pixel distance between the first target object and the second target object in the target video frame based on the position information of the first target object and the second target object in the target video frame;

[0194] The second determination module 330 is configured to determine the first target object and the second target object as a target pair when the first pixel distances corresponding to consecutive multiple target video frames in the target video are all less than a first target threshold;

[0195] The first detection module 340 is configured to detect whether the first target object has a leash based on the target pair.

[0196] According to the detection device for pet leashes provided by the embodiments of the present invention, it is determined whether there is an accompanying relationship between the first target object and the second target object based on the position trajectories between the first target object and the second target object. In the case where it is determined that there is an accompanying relationship between the two, the first target object in the target video frame is further identified as to whether it is leashed, with high accuracy and strong adaptability.

[0197] In some embodiments, the first detection module 340 can also be used for:

[0198] Based on the target pair, perform image extraction on the target video frame to generate a target image block, where the target image block includes at least a pixel region between at least part of the first target object and the second target object;

[0199] Perform image recognition on the target image block to detect whether the first target object is leashed.

[0200] In some embodiments, the first detection module 340 can also be used for:

[0201] Based on the target pair, determine the target matte region corresponding to the target video frame;

[0202] Based on the target matte region, perform image extraction on the target video frame to generate a target image block.

[0203] In some embodiments, the target video includes consecutive first video frames and second video frames. The first recognition module 310 can also be used for:

[0204] Respectively perform feature extraction on the first video frame and the second video frame to generate the first candidate object feature corresponding to the first video frame and the second candidate object feature corresponding to the second video frame;

[0205] Based on the position information of the first candidate object feature in the first video frame and the position information of the second candidate object feature in the second video frame, generate a second pixel distance;

[0206] In the case where the similarity between the first candidate object feature and the second candidate object feature is greater than the second target threshold and the second pixel distance is less than the third target threshold, determine the first candidate object feature and the second candidate object feature as the same target object.

[0207] In some embodiments, the first recognition module 310 can also be used for: labeling the target object to generate labeling information, where the labeling information includes at least one of the identifier of the target object, the features of the target object in the target video frame, and the position information of the target object in the target video frame.

[0208] In some embodiments, the device may further include a first output module, configured to: after detecting whether a first target object is tethered based on a target pair, output an alarm message when it is determined that the first target object is not tethered.

[0209] Figure 4 The figure illustrates a schematic physical structure diagram of an electronic device, as Figure 4 shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the method for detecting pet tethering, and the method includes: identifying target objects from target video frames of a target video, where the target objects include a first target object corresponding to a pet and a second target object corresponding to a person; determining a first pixel distance between the first target object and the second target object in the target video frame based on the position information of the first target object and the second target object in the target video frame; determining the first target object and the second target object as a target pair when the first pixel distances corresponding to consecutive multiple target video frames in the target video are all less than a first target threshold; detecting whether the first target object is tethered based on the target pair.

[0210] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0211] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the pet leash detection method provided by each of the above methods. The method includes: identifying target objects from target video frames of a target video, where the target objects include a first target object corresponding to a pet and a second target object corresponding to a person; determining a first pixel distance between the first target object and the second target object in the target video frame based on the position information of the first target object and the second target object in the target video frame; when the first pixel distances corresponding to a continuous plurality of target video frames in the target video are all less than a first target threshold, determining the first target object and the second target object as a target pair; and detecting whether the first target object is leashed based on the target pair.

[0212] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the pet leash detection method provided by each of the above. The method includes: identifying target objects from target video frames of a target video, where the target objects include a first target object corresponding to a pet and a second target object corresponding to a person; determining a first pixel distance between the first target object and the second target object in the target video frame based on the position information of the first target object and the second target object in the target video frame; when the first pixel distances corresponding to a continuous plurality of target video frames in the target video are all less than a first target threshold, determining the first target object and the second target object as a target pair; and detecting whether the first target object is leashed based on the target pair.

[0213] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0214] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A detection method for a pet leash, characterized in that, including: identifying a target object from a target video frame of a target video, where the target object includes a first target object corresponding to a pet and a second target object corresponding to a person; determining a first pixel distance between the first target object and the second target object in the target video frame based on position information of the first target object and the second target object in the target video frame; when first pixel distances corresponding to a continuous plurality of the target video frames in the target video are all less than a first target threshold, determining the first target object and the second target object as a target pair, where the first target threshold is used to determine whether the first target object and the second target object are in an accompanying state, and the target pair is determined based on a trajectory feature formed by first pixel distances corresponding to a continuous plurality of the target video frames and is used to represent that there is an accompanying relationship between the first target object and the second target object; detecting whether the first target object is tethered based on the target pair; wherein, the target video includes continuous first video frames and second video frames, and the identifying a target object from a target video frame of the target video includes: performing feature extraction on the first video frame and the second video frame respectively to generate a first candidate object feature corresponding to the first video frame and a second candidate object feature corresponding to the second video frame; generating a second pixel distance based on position information of the first candidate object feature in the first video frame and position information of the second candidate object feature in the second video frame; when the similarity between the first candidate object feature and the second candidate object feature is greater than a second target threshold and the second pixel distance is less than a third target threshold, determining the first candidate object feature and the second candidate object feature as the same target object.

2. The detection method of the pet leash according to claim 1, wherein The detecting whether the first target object is tethered based on the target pair includes: performing image extraction on the target video frame based on the target pair to generate a target image block, where the target image block includes at least a pixel region between the first target object and the second target object; performing image recognition on the target image block to detect whether the first target object is tethered.

3. The detection method of the pet leash according to claim 2, wherein The performing image extraction on the target video frame based on the target pair to generate a target image block includes: determining a target matte region corresponding to the target video frame based on the target pair; performing image extraction on the target video frame based on the target matte region to generate the target image block.

4. The detection method of the pet leash according to any one of claims 1-3, characterized in that, The identifying a target object from a target video frame of the target video includes: labeling the target object to generate labeling information, where the labeling information includes at least one of an identifier of the target object, a feature of the target object in the target video frame, and position information of the target object in the target video frame.

5. The detection method of the pet leash according to any one of claims 1-3, characterized in that, After the detecting whether the first target object is tethered based on the target pair, the method further includes: outputting an alarm message when it is determined that the first target object is not tethered.

6. A detection device for a pet leash, characterized in that, including: A first recognition module, configured to recognize a target object from a target video frame of a target video, where the target object includes a first target object corresponding to a pet and a second target object corresponding to a person; A first determination module, configured to determine a first pixel distance between the first target object and the second target object in the target video frame based on position information of the first target object and the second target object in the target video frame; A second determination module, configured to determine the first target object and the second target object as a target pair when first pixel distances corresponding to a plurality of consecutive target video frames in the target video are all less than a first target threshold, where the first target threshold is used to determine whether the first target object and the second target object are in an accompanying state, the target pair is determined based on a trajectory feature formed by first pixel distances corresponding to a plurality of consecutive target video frames, and is used to represent that there is an accompanying relationship between the first target object and the second target object; A first detection module, configured to detect whether the first target object is leashed based on the target pair; Wherein, the target video includes consecutive first video frames and second video frames, and the first recognition module is specifically configured to: Extract features from the first video frame and the second video frame respectively, to generate a first candidate object feature corresponding to the first video frame and a second candidate object feature corresponding to the second video frame; Generate a second pixel distance based on position information of the first candidate object feature in the first video frame and position information of the second candidate object feature in the second video frame; Determine the first candidate object feature and the second candidate object feature as the same target object when the similarity between the first candidate object feature and the second candidate object feature is greater than a second target threshold and the second pixel distance is less than a third target threshold.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein When the processor executes the program, the steps of the method for detecting whether a pet is leashed according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method for detecting whether a pet is leashed according to any one of claims 1 to 5 are implemented.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method for detecting whether a pet is leashed according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Image recognition method, electronic equipment and computer readable storage medium

    CN110163029A

  • Illegal dog walking event detection method and device based on monitoring video

    CN112906678A