A pet traction detection method and device, an electronic device, and a storage medium

CN122090384BActive Publication Date: 2026-08-11ZHEJIANG UNIVIEW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]现有的宠物有效牵引的检测,通过检测目标宠物和牵引绳,计算牵引绳区域与目标宠物区域的相交比判断目标宠物是否被有效牵引,但是这种方式在牵引绳细小、透明或与背景相似时,检测准确率较低,并且对于目标宠物虽佩戴牵引绳但无人牵引等无效牵引状态无法识别

Benefits of technology

[0010]本发明实施例的技术方案,通过识别视频帧图像中的目标宠物、目标人物以及牵引绳,并在微观层面确定目标人物的手部握持状态、在中观层面物理量化牵引绳的绳索张力状态,以及在宏观层面耦合目标宠物、目标人物与牵引绳之间的运动耦合特征。通过三级联动判断目标宠物的宠物牵引状态,避免了单一特征失效导致的误判,不受单一视觉特征的限制,显著提高了复杂场景下的鲁棒性。同时,能够精确感知目标人物的握持姿态,量化牵引绳的绳索张力状态,分析目标人物操控与牵引绳、目标宠物动态响应之间的运动耦合关系,实现了对宠物牵引行为有效性的智能、精确、可靠判别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090384B_ABST
    Figure CN122090384B_ABST
Patent Text Reader

Abstract

This invention discloses a pet leash detection method, device, electronic device, and storage medium. The method includes: determining a target pet, a target person matching the target pet, and a leash matching the target pet and the target person based on video frame images; determining the target person's hand grip state, the leash tension state, and the motion coupling characteristics between the target pet, the target person, and the leash based on the video frame images; and determining the pet's leash state based on the hand grip state, the leash tension state, and the motion coupling characteristics. This invention enables intelligent, precise, and reliable determination of the effectiveness of pet leash behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, and in particular to a pet leash detection method, device, electronic device, and storage medium. Background Technology

[0002] With the acceleration of urbanization and the improvement of residents' living standards, the number of pets is increasing daily, and the safe management of pets' outdoor activities has become an important issue in urban governance and community services. Among these measures, leashing pets is crucial to preventing them from getting lost, injuring people, or causing traffic accidents. Therefore, utilizing computer vision and artificial intelligence technologies to automatically detect whether pets are on leashes has become an important research direction in the fields of smart security and property management.

[0003] Existing methods for detecting effective pet leash use the detection of the target pet and the leash, calculating the intersection ratio between their respective areas to determine if the pet is effectively leashed. However, this method has low accuracy when the leash is thin, transparent, or blends into the background, and it fails to detect invalid leash situations where the pet is on a leash but not being led. Another method relies on the consistency between the leash trajectory, the target pet's trajectory, and the target person's trajectory to determine effective leash traction. However, this method is prone to failure in scenarios with occlusion, rapid movement, or low frame rates, and it cannot effectively detect situations where the pet lunges forward and pulls on the target person, posing a safety risk.

[0004] Therefore, there is an urgent need for a testing method to accurately and reliably determine the effectiveness of pet leashing. Summary of the Invention

[0005] This invention provides a pet leash detection method, device, electronic device, and storage medium to achieve intelligent, precise, and reliable judgment of the effectiveness of pet leash behavior.

[0006] In a first aspect, embodiments of the present invention provide a pet leash detection method, the method comprising: Based on video frame images, identify the target pet, the target person matching the target pet, and the leash matching the target pet and the target person; Based on the video frame images, determine the hand gripping state of the target person, the rope tension state of the leash, and the motion coupling characteristics between the target pet, the target person, and the leash; The pet traction state is determined based on the hand grip state, the rope tension state, and the motion coupling characteristics.

[0007] Secondly, embodiments of the present invention also provide a pet leash detection device, the device comprising: The leash detection module is used to determine the target pet, the target person matching the target pet, and the leash matching the target pet and the target person based on video frame images. The multi-feature extraction module is used to determine the hand gripping state of the target person, the rope tension state of the leash, and the motion coupling features between the target pet, the target person, and the leash based on video frame images. The pet leash state determination module is used to determine the pet leash state based on the hand grip state, the rope tension state, and the motion coupling characteristics.

[0008] Thirdly, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the pet leash detection method as described in any of the embodiments of the present invention.

[0009] Fourthly, embodiments of the present invention also provide a storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform the pet leash detection method as described in any of the embodiments of the present invention.

[0010] The technical solution of this invention identifies the target pet, the target person, and the leash in a video frame image. It determines the target person's hand grip at a microscopic level, physically quantifies the leash tension at a mesoscopic level, and couples the motion coupling characteristics between the target pet, the target person, and the leash at a macroscopic level. This three-level linkage approach to determine the pet's leash status avoids misjudgments caused by the failure of a single feature, is not limited by a single visual feature, and significantly improves robustness in complex scenarios. Simultaneously, it can accurately perceive the target person's grip posture, quantify the leash tension, and analyze the motion coupling relationship between the target person's control and the leash and the target pet's dynamic response, achieving intelligent, accurate, and reliable judgment of the effectiveness of pet leash behavior.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a pet leash detection method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a pet leash detection method provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of key points of the hand provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the structure of a pet leash detection device provided in Embodiment 3 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. In the embodiments of this application, certain software, components, models, and other existing industry solutions may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0016] The acquisition, transmission, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0017] Example 1 Figure 1The flowchart of a pet leash detection method is provided in Embodiment 1 of the present invention. This embodiment can be applied to the situation of determining the validity of a pet's leash status. The method can be executed by a pet leash detection device, which can be implemented in hardware and / or software. The pet leash detection device can be configured in a server or electronic device and used in conjunction with a monitoring camera or other shooting device.

[0018] like Figure 1 As shown, the method includes: S110. Based on the video frame images, determine the target pet, the target person matching the target pet, and the leash matching the target pet and the target person.

[0019] In this context, "video frame image" refers to the video frame currently being processed. "Target pet" refers to the pet appearing in the video frame image; the target pet type includes cats, dogs, and other pets that require a leash. "Target person" refers to the person deemed to have a preliminary pairing relationship with the target pet and therefore bears responsibility for leashing it. It should be noted that this person can be a natural person, and with the development of artificial intelligence technology, it can also be extended to include AI robots, etc. The leash is the rope tool used by the target person to lead the target pet; typically, one end is tied to the target pet's chest or back, and the other end is held by the target person.

[0020] In this embodiment, a target object detection algorithm or model can be used to detect pets in continuous video frame images. When a target pet is detected, the detection of the target pet's leash status in this embodiment is triggered. Simultaneously, after acquiring the video frame images, to improve subsequent processing effects, image preprocessing such as noise reduction, illumination intensity normalization, and resolution adjustment can be performed before pet detection and pet leash status detection.

[0021] In this embodiment, determining the target person matching the target pet can be achieved by performing pet and person detection on consecutive video frames, assigning identity identifiers to the detected target pet and person, and generating the movement trajectories of the target pet and person. Based on spatial proximity and movement trajectory consistency, the movement trajectory of the person matching the target pet's movement trajectory is determined. The identity identifier corresponding to this movement trajectory is paired and stored with the target pet's identity identifier. When the target pet and person are detected in the current video frame, the target person can be directly determined based on the person identity identifier paired with the target pet's identity identifier.

[0022] In this embodiment, determining the leash that matches the target pet and the target person can be achieved based on edge detection algorithms, contour extraction algorithms, etc., or it can be achieved by semantic segmentation of the search area in the video frame image using a deep learning model. The search area must cover the area where the target pet and the target person are located.

[0023] This embodiment provides a method for accurately extracting a leash from video frame images. Furthermore, determining a leash that matches the target pet and target person based on the video frame images may include: S111. Determine the center point of the target person's hand grip based on the video frame image; S112. Determine the search area based on the center point of the hand grip and the center point of the target pet. S113. Determine a binary region based on the search region, and perform depth estimation and curve fitting on the binary region to obtain a three-dimensional leash curve that matches the target pet and the target person.

[0024] In this embodiment, the center point of the target person's hand grip can be determined by further detecting the hand region within the target person's area using a detection algorithm or model. The centroid of the hand region or the center point of the smallest bounding rectangle of the hand region can be used as the center point of the hand grip. Alternatively, key points of the hand can be extracted, and the average coordinates of the extracted key points can be used as the coordinates of the center point of the hand grip. Furthermore, after extracting the key points, the average coordinates of the index finger key point and the thumb key point can be determined as the coordinates of the center point of the hand grip. This embodiment does not limit the specific method for determining the center point of the hand grip.

[0025] In this embodiment, similarly, the centroid of the target pet region or the center point of the smallest bounding rectangle of the target pet region can be used as the target pet's center point. Alternatively, key points of the target pet can be extracted, and the average coordinates of the extracted key points or the average coordinates of the key points on the target pet's head can be used as the target pet's center point. Furthermore, after extracting the key points, the key points on the target pet's neck can be directly used as the target pet's center point. This embodiment does not limit the specific method for determining the target pet's center point.

[0026] Based on the center point of the hand grip and the center point of the target pet, the search area is determined. Specifically, a high-resolution image block is defined centered on the center point of the hand grip. Pixel-level semantic analysis is performed on this image block to distinguish three semantic components: the hand, the leash, and the background. Starting from the center point of the hand grip, the contour of the contact area between the leash area and the hand area is extracted. The search area is constructed along the leash contour until the center point of the target pet is the endpoint.

[0027] Furthermore, the width of the search area can be set to a fixed value, which can be determined based on experience or as a multiple of the historical average width of the traction rope. An adaptive width setting is also possible; for example, the search area width can be determined using the following formula: ,in, Indicates the width of the search area. , As a coefficient, it can be Set to 0.1, Set to 10 pixels. Indicates the center point of the hand grip. Indicates the center point of the target pet.

[0028] The binary region is determined based on the search region. Specifically, a directionally controllable Gabor filter bank (4 directions, 3 scales) is applied to the search region to enhance the linear structure response, resulting in a filtered response map. Adaptive thresholding is then performed on the filtered response map to obtain the binary region.

[0029] Depth estimation and curve fitting are performed on the binary region to obtain a 3D leash curve matching the target pet and the target person. Specifically, morphological refinement is performed on the binary region to obtain a single-pixel wide centerline C_2d(t). A depth estimation network is used to process the binary region to obtain a scene depth map D(t). Based on the scene depth map, the centerline is mapped to 3D space, resulting in P_3d(x,y,z)=backproject(C_2d(x,y), D(x,y)). B-spline curve fitting is then used to process the mapped P_3d(x,y,z) to obtain a smooth 3D leash curve L_3d(t): s→(x,y,z), where s is the arc length parameter.

[0030] In this embodiment, by using filtering extraction, depth estimation, and curve fitting, the detection problems of bent, jittery, thin, and weakly textured ropes can be solved, thus achieving accurate detection of traction ropes.

[0031] S120. Based on the video frame images, determine the hand gripping state of the target person, the rope tension state of the leash, and the motion coupling characteristics between the target pet, the target person, and the leash.

[0032] The target person's hand grip directly reflects their grip posture, indicating at a microscopic level whether they intend to actively hold the leash. Hand grip status can be represented by scores, levels, etc. For example, a score of 0-10 can be used, with higher scores indicating a stronger intention to actively hold the leash. Different hand grip status types can also be pre-set, represented by probabilities or confidence levels corresponding to each type. For example, hand grip status types can include a firm grip, a loose grip, and no grip.

[0033] In an optional embodiment, determining the hand-holding state of a target person based on video frame images can be achieved using a hand-holding state recognition model pre-trained from a deep learning model. For example, the hand-holding state of sample images can be labeled, and the sample images can be input into a pre-set deep learning model. The predicted hand-holding state values ​​of the deep learning model are compared with the labels, and the model is fine-tuned and iteratively trained until the model's recognition accuracy meets the training requirements, thus obtaining the hand-holding state recognition model. The search region is then input into the hand-holding state recognition model to obtain the hand-holding state of the target person.

[0034] In another optional embodiment, the hand gripping state of the target person can be determined based on the video frame image. This can be achieved by extracting key points of the target person's hand, performing feature analysis on each key point, such as feature analysis of finger curvature, hand contour, etc., and classifying and recognizing the hand gripping state based on the feature analysis results to obtain the hand gripping state of the target person.

[0035] The leash tension state is a quantitative representation of the physical tension of the leash, reflecting the traction control between the person and the pet at a meso-level. Similarly, leash tension can be represented by scores, levels, etc. For example, using a score of 0-10, a higher score indicates greater leash tension and a tighter leash, while a lower score indicates less leash tension and a looser or even tensionless leash. Different leash tension state types can also be pre-set, represented by probabilities or confidence levels corresponding to different types. For example, leash tension state types can include high tension, low tension, and no tension, or taut, loose, and dragging states, etc.

[0036] Similarly, in an optional embodiment, determining the rope tension state of the traction rope based on video frame images can be achieved using a rope tension state recognition model pre-trained from a deep learning model. For example, rope tension state labels can be applied to sample images of the traction rope. These images can then be input into a pre-set deep learning model. The model's predicted rope tension state values ​​are compared with the labels, and the model is fine-tuned and iteratively trained until its recognition accuracy meets the training requirements, thus obtaining the rope tension state recognition model. The three-dimensional traction rope curve is then input into the rope tension state recognition model to obtain the rope tension state of the traction rope.

[0037] In another optional embodiment, the rope tension state of the traction rope can be determined based on the video frame image. This can also be achieved by performing curve feature analysis on the three-dimensional traction rope curve, such as curvature features and shape features. Based on the curve feature recognition results, the rope tension state is classified and identified to obtain the rope tension state of the traction rope.

[0038] The kinematic coupling characteristics between the target pet, the target person, and the leash directly reflect the dynamic mechanical interaction between them, demonstrating the causal coupling between the person holding the leash and the pet's dynamic response at a macroscopic level. Similarly, kinematic coupling characteristics can be represented by scores or levels. For example, using a score of 0-10, a higher score indicates a more pronounced causal coupling between the person holding the leash and the pet's dynamic response, and a higher effectiveness of the pet's response to the person pulling it by the leash.

[0039] First, the motion characteristics of the target person, target pet, and leash can be determined, including speed and acceleration. Based on the speed of the target person and the leash, the transfer entropy from the target person to the leash can be calculated, measuring the amount of information transferred from the target person's motion to the leash's motion. Similarly, the transfer entropy from the target pet to the leash can be calculated, measuring the amount of information transferred from the target pet's motion to the leash's motion. Based on the two transfer entropies, the causal dominant features of the motion between the target pet, target person, and leash can be calculated. The causal dominant features, the correlation between the target pet's speed and the target person's speed, and the leash tension state characteristics can be determined. These features are then combined to form a holistic representation of the motion coupling features.

[0040] In this embodiment, the hand gripping state of the target person is determined at the micro level, the rope tension state is determined at the meso level, and the motion coupling characteristics between the person, rope, and pet are determined at the macro level. Combining these three different dimensions of features for comprehensive decision-making, and combining visual, mechanical, and kinematic laws, can avoid overall misjudgment caused by the failure of a single feature. For example, even if the rope area is not segmented due to shadows, the traction state can still be accurately judged by the hand gripping state and the overall motion coupling characteristics.

[0041] S130. Determine the pet traction state based on the hand grip state, the rope tension state, and the motion coupling characteristics.

[0042] The pet leash status indicates whether the target pet is effectively leashed by the target person. Effective leashing means that the target person can effectively control the target pet and play a leading role in the movement of both the target person and the target pet. Situations such as the target pet suddenly lunging and the target person passively following, the target person not effectively controlling the leash, or the target pet not wearing a leash are not considered effective leashing.

[0043] Furthermore, the pet's leash status can be represented by a score or level, with higher scores or levels indicating a higher degree of effectiveness in the pet's leash being controlled by the target person. The pet's leash status can also be represented by different leash status types and their corresponding confidence levels, with higher confidence levels indicating a higher probability that the pet's leash status belongs to that type. For example, pet leash status types can include effective leash, passive following, ineffective leash holding, and no leash. Passive following refers to the state where the target person is passively followed due to the target pet's sudden movement; ineffective leash holding refers to the state where the target person is not effectively controlling the leash; and no leash refers to the state where the target pet is not wearing a leash. The different types of pet leash status can be flexibly set according to the actual application scenario; this embodiment does not impose any limitations on this.

[0044] In an optional embodiment, determining the pet traction state based on the hand grip state, the rope tension state, and the motion coupling characteristics may include: quantifying the hand grip state, rope tension state, and motion coupling characteristics respectively; specifically, the hand grip state and rope tension state can be represented by scores or different types of corresponding confidence levels, and the motion coupling degree is determined based on the motion coupling characteristics. Similarly, the motion coupling degree is represented by scores or confidence levels; the higher the motion coupling degree, the greater the degree to which the target person's active movement drives the target pet's movement. The weighted sum of the hand grip state, rope tension state, and motion coupling degree yields the pet traction effectiveness; the higher the pet traction effectiveness, the higher the effectiveness of the pet traction state.

[0045] In another optional embodiment, determining the pet traction state based on the hand grip state, the rope tension state, and the motion coupling features may further include: inputting the hand grip state, the rope tension state, and the motion coupling features into a pre-trained temporal classification model to obtain at least two pet traction states output by the temporal classification model and a confidence level matching the pet traction state.

[0046] The temporal classification model can be trained using a bidirectional LSTM (Long Short-Term Memory) network. Specifically, sample images are pre-set, and the pet's leash state is labeled for each sample image. Using the same method as in this embodiment, the hand grip state, rope tension state, and motion coupling features corresponding to each sample image are determined. The hand grip state, rope tension state, and motion coupling features corresponding to each sample image are input into the bidirectional LSTM model. The model loss is calculated based on the pet's leash state predicted by the bidirectional LSTM model and the pet's leash state labeled in the sample images, and the model is fine-tuned. The above process is repeated until the prediction accuracy of the bidirectional LSTM model meets the pre-set accuracy requirements, thus obtaining the temporal classification model.

[0047] In this embodiment, the hand gripping state, rope tension state, and motion coupling features are concatenated to obtain a concatenated feature vector. The concatenated feature vector is then input into a time-series classification model to output the confidence level corresponding to each pet's traction state.

[0048] In this embodiment, the confidence scores of each pet leash state output by the time-series classification model are used to represent the probability that the pet leash state belongs to that type. Since there are multiple types of pet leash states, after obtaining the confidence scores of each type of pet leash state output by the time-series classification model, it is necessary to further determine the current pet leash state based on the confidence scores of each type of pet leash state.

[0049] Specifically, in this embodiment, the method further includes: determining the traction force based on the center point of the target person's hand grip and the center point of the target pet.

[0050] In this embodiment, the traction force can be estimated based on the center point of the target person's hand grip and the center point of the target pet. Specifically, it can be expressed by the following formula: , .in, Indicates traction force. The quality of the target pet can be estimated by performing visual volume mapping on the target pet in the video frame image. This represents the instantaneous acceleration of the target pet at time t. Indicates the center point of the target pet.

[0051] In this embodiment, the purpose of determining the traction force is to combine it with the confidence level of each pet's traction status as an auxiliary basis for judging the pet's traction status, and to jointly determine the current traction status of the pet.

[0052] After obtaining at least two pet leash states output by the time-series classification model and the confidence levels matching the pet leash states, the method further includes: if the confidence level for determining that the pet leash state is without a leash is greater than or equal to a preset confidence threshold, then the current pet leash state is determined to be without a leash, and a target pet leash state alarm is triggered; otherwise, the current pet leash state is determined to be the pet leash state with the highest confidence level; if the current pet leash state is determined to be with an invalid leash, or if the current pet leash state is determined to be passively following, and the traction force is greater than or equal to a preset tension threshold, then a target pet leash state alarm is triggered.

[0053] In this embodiment, if the confidence level of the pet's leash status being unleashed is greater than or equal to a preset confidence threshold, such as 0.8, then to ensure safety, regardless of the confidence level of other pet leash status types, the current pet leash status is directly determined to be unleashed.

[0054] When the confidence level for a pet's leash-free status is less than a confidence threshold, the leash status with the highest confidence level is taken as the current pet leash status. Furthermore, the current pet leash status can be smoothed over time based on historical pet leash statuses to avoid frequent jumps in leash status.

[0055] In this embodiment, when the current pet leash is ineffective, especially if the pet's historical leash status was effective, an alarm for the target pet's leash status is required. Furthermore, when the current pet is passively following, the leash force needs to be considered to further determine whether an alarm for the target pet's leash status is necessary. It is understandable that if the current pet is passively following and the leash force is high, it indicates that the target person is passively following due to the target pet's sudden acceleration, and the target person has weak control over the target pet, thus requiring an alarm for the target pet's leash status.

[0056] Furthermore, the alarm methods for the target pet's leash status when the traction force is greater than or equal to the tension threshold in the leashless state, the invalid leash holding state, and the passive following state can be the same or different. Preferably, the alarm method for the target pet's leash status when the traction force is greater than or equal to the tension threshold in the passive following state can be set to the lightest alarm method, such as an attention prompt; the alarm method for the invalid leash holding state can be set to a more severe alarm method, such as a safety risk warning; and the alarm method for the leashless state can be set to the most severe alarm method, such as an emergency alarm prompt. However, this embodiment does not limit the alarm method for the target pet's leash status corresponding to different states in a multi-level alarm scenario.

[0057] In this embodiment, by setting early warning mechanisms or even tiered early warning mechanisms for different pet leashing states, it is possible to make comprehensive decisions based on three levels of characteristics and provide different levels of alarm prompts, thereby achieving precise and differentiated safety management for pet leashing states with different risk levels.

[0058] Furthermore, after determining the pet's leash status and issuing an alarm, manual spot checks and verifications can be performed. If the target pet's leash status corresponding to the target video frame image is corrected, the target video frame image and the corrected pet leash status are added to the incremental learning queue. The temporal classification model is then fine-tuned based on the incremental learning queue, thereby continuously improving the classification accuracy of the temporal classification model and the precision of judging pet leash behavior.

[0059] The technical solution of this invention identifies the target pet, the target person, and the leash in a video frame image. It determines the target person's hand grip at a microscopic level, physically quantifies the leash tension at a mesoscopic level, and couples the motion coupling characteristics between the target pet, the target person, and the leash at a macroscopic level. This three-level linkage approach to determine the pet's leash status avoids misjudgments caused by the failure of a single feature, is not limited by a single visual feature, and significantly improves robustness in complex scenarios. Simultaneously, it can accurately perceive the target person's grip posture, quantify the leash tension, and analyze the motion coupling relationship between the target person's control and the leash and the target pet's dynamic response, achieving intelligent, accurate, and reliable judgment of the effectiveness of pet leash behavior.

[0060] Example 2 Figure 2 This is a flowchart of a pet leash detection method provided in Embodiment 2 of the present invention. Based on the above embodiments, the present invention further specifies the process of determining the hand gripping state of the target person, the process of determining the rope tension state of the leash, and the process of determining the motion coupling characteristics between the target pet, the target person, and the leash.

[0061] like Figure 2As shown, the method includes: S210. Based on the video frame images, determine the target pet and the target person matching the target pet.

[0062] The process of pet and human identification based on video frame images, and matching pets and humans, has been described in the above embodiments and will not be repeated here.

[0063] S220. Determine the center point of the target person's hand grip based on the video frame image.

[0064] Specifically, key points of the hand can be extracted from the target person area in the video frame image, and the midpoint of the line connecting the key points of the index finger and the key points of the thumb can be used as the center point of the hand grip.

[0065] For example, Figure 3 A schematic diagram of key hand points is provided, such as... Figure 3 As shown, there are 21 key points for the hand.

[0066] Furthermore, the contact area of ​​the target person's hands can be detected in advance, and key points can be extracted for hands with traction ropes in the contact area.

[0067] S230. Determine the search area based on the center point of the hand grip and the center point of the target pet.

[0068] S240. Determine a binary region based on the search region, and perform depth estimation and curve fitting on the binary region to obtain a three-dimensional leash curve that matches the target pet and the target person.

[0069] The process of determining the search area based on the center point of the target person's hand grip and the center point of the target pet, and then determining the three-dimensional leash curve based on the search area, has been described in the above embodiments and will not be repeated in this embodiment.

[0070] S250. Based on the video frame image, extract the key points of the target person's hand and determine the bounding box of the key points of the hand.

[0071] In this context, a bounding box refers to the smallest or simplest rectangle or cuboid that can enclose the target object. In this embodiment, the bounding box of the hand key points refers to the smallest cuboid that can enclose each hand key point.

[0072] In this embodiment, the bounding box of the key points of the hand can be determined by the AABB (Axis-Aligned Bounding Box) method or the OBB (Oriented Bounding Box) method. The specific process will not be described in detail in this embodiment.

[0073] S260. Based on the bounding boxes of the hand key points, determine the bounding box tightness feature, and based on the hand key points, determine the finger curvature feature, palm depth feature, and hand contour feature.

[0074] The bounding box tightness feature can be represented by the ratio of the volume to the surface area of ​​the bounding box of the hand keypoint. Specifically, the side lengths of the bounding box of the hand keypoint in the x, y, and z dimensions are calculated, and the volume and surface area of ​​the bounding box are calculated based on the side lengths. The ratio of volume to surface area is then used as the bounding box tightness feature.

[0075] It is understandable that when the hand is bent and clenched, the fingers are bent and the bounding box tightness feature is small. Therefore, the bounding box tightness feature can be used to reflect whether the hand is in a clenched state.

[0076] Finger curvature features are used to quantify the degree of finger curvature, and can be represented by the angle between the line connecting the distal and proximal phalanges of each finger and the back of the hand. Specifically, for each finger, the angle between the line connecting the distal and proximal phalanges and the back of the hand is calculated. Taking the index finger as an example... Figure 3 As shown, the feature points of the index finger include points 6, 7, and 8. The index finger vector formed by the line connecting the distal and proximal phalanges can be represented by P8-P6, and the back-of-hand vector can be represented by P0-P9. The cosine of the included angle can be expressed by the following formula: ,in, Indicates the amount pointed to by the index finger. Let represent the vector on the back of the hand. The angle between the line connecting the distal and proximal phalanges of the fingers and the back of the hand can be expressed by the following formula: .

[0077] Understandably, when the hand is bent and clenched, the angle between the line connecting the distal and proximal knuckles of the fingers and the back of the hand is smaller.

[0078] Palm depth features are used to quantify the depth of the palm region of the hand. Specifically, the average depth value of the palm region can be calculated as the palm depth feature based on key points in the palm region using the following formula: .

[0079] It is understandable that the average depth of the palm area will be greater when the hand is bent and clenched than when the hand is relaxed.

[0080] Hand contour features are used to quantify convexity defects in the hand contour. Convexity defects refer to the concave areas between the hand contour and its convex hull. The hand contour shape is not strictly a convex set; the portion of the hand contour that is concave inward and extends beyond the boundary of the convex hull is considered a convexity defect. Specifically, convexity defects can be calculated as hand contour features using the following formula: ,in, Describes the outline features of the hand. This indicates the number of significantly concave areas in the hand's outline relative to its convex hull. This represents the deepest point of the k-th convex defect. express to line segment The orthogonal projection points of , where, This represents the starting point of the convex enclosure of the k-th convex defect. This represents the endpoint of the convex enclosure of the k-th convex defect. This represents the ideal boundary when the hand's outline is convex.

[0081] Understandably, when the hand is bent and clenched, the number of convex defects is less, but the area of ​​the convex defects is larger.

[0082] S270. Determine the hand gripping state of the target person based on the tightness characteristics of the enclosing box, the curvature characteristics of the fingers, the depth characteristics of the palm, and the contour characteristics of the hand.

[0083] In an optional embodiment, the hand gripping state of the target person is determined based on bounding box tightness features, finger curvature features, palm depth features, and hand contour features. This can be achieved by quantifying the bounding box tightness features, finger curvature features, palm depth features, and hand contour features, and determining the positive or negative weights and magnitudes of these features based on their correlation with the hand gripping state (positive or negative correlation) and their contribution. Then, the bounding box tightness features, finger curvature features, palm depth features, and hand contour features are weighted and summed to obtain the quantified hand gripping state.

[0084] In another optional embodiment, the hand gripping state of the target person is determined based on the bounding box tightness features, finger curvature features, palm depth features, and hand contour features. Alternatively, the bounding box tightness features, finger curvature features, palm depth features, and hand contour features can be input into a pre-trained hand gripping state classification model to obtain the hand gripping state type and its corresponding confidence level output by the hand gripping state classification model.

[0085] Specifically, the training process of the hand grip state classification model may include: determining sample hand images pre-labeled with hand grip state types, and calculating the bounding box tightness features, finger curvature features, palm depth features, and hand contour features of each sample hand image using the same method described above; inputting the bounding box tightness features, finger curvature features, palm depth features, and hand contour features of each sample hand image into a pre-determined lightweight classification model; calculating the model loss and fine-tuning the model based on the hand grip state type predicted by the lightweight classification model and the hand grip state type labeled in the sample hand images; repeating the above process until the accuracy of the model prediction meets the pre-set accuracy requirements, etc., to obtain the hand grip state classification model.

[0086] The hand grip state classification model outputs hand grip state types, which can include tight grip, loose grip, and no grip, etc. The confidence score is used to represent the probability of belonging to that hand grip state type.

[0087] In this embodiment, by combining multiple dimensions such as bounding box tightness, finger curvature, palm depth, and hand contour, the hand gripping state of the target person is determined, which can improve the accuracy of hand gripping state detection and enhance the anti-interference ability of hand gripping state detection.

[0088] S280. Based on the three-dimensional traction rope curve, determine the curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics of the traction rope.

[0089] The curvature feature represents the degree of bending of the traction rope, and it can be calculated based on the average curvature of the three-dimensional traction rope curve. Specifically, the curvature at each point on the three-dimensional traction rope curve is calculated, and the average curvature at each point is calculated to obtain the average curvature. The curvature feature can be calculated using the following formula: ,in, Indicates curvature characteristics. The adjustment parameter representing the average curvature. It represents the mean curvature.

[0090] Understandably, the tighter the traction rope, the lower the degree of bending, the smaller the average curvature, and the closer the curvature characteristic is to 1.

[0091] The vibration spectrum characteristics represent the lateral vibration frequency of the traction rope, which can be calculated based on the dominant lateral vibration frequency of the three-dimensional traction rope curve. Specifically, equidistant sampling is performed on the three-dimensional traction rope curve to obtain a predetermined number of sampling points, such as three. The lateral displacement perpendicular to the main direction of the rope at each sampling point is determined, and a short-time Fourier transform is performed on the lateral displacement of each sampling point to extract the dominant frequency and frequency band energy distribution. The vibration spectrum characteristics can be calculated using the following formula: ,in, Indicates the characteristics of the vibration spectrum. , The adjustment parameter representing the transverse vibration frequency. Indicates the dominant frequency of lateral vibration. This indicates the reference frequency for slack ropes.

[0092] The curve morphology can be calculated based on the ratio of the total length of the three-dimensional traction rope curve to the straight-line distance between its endpoints. Specifically, first calculate the ratio of the total length of the three-dimensional traction rope curve to the straight-line distance between its endpoints, and then calculate the curve morphology using the following formula: ,in, Indicates the morphological characteristics of the curve. This represents the adjustment coefficient. It is the ratio of the total length of the three-dimensional traction rope curve to the straight-line distance between the endpoints.

[0093] Understandably, when the traction rope is fully taut, the ratio of the total length of the three-dimensional traction rope curve to the straight-line distance between its endpoints is 1. Close to 1.

[0094] S290. Determine the rope tension state of the traction rope based on its curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics.

[0095] Similarly, the rope tension state of the traction rope can be determined based on its curvature characteristics, vibration spectrum characteristics, and curve shape characteristics. This can be achieved by quantifying the curvature characteristics, vibration spectrum characteristics, and curve shape characteristics, and determining the positive or negative weights and magnitudes of these weights based on their correlation with the rope tension state (positive or negative correlation) and their contribution. Finally, the curvature characteristics, vibration spectrum characteristics, and curve shape characteristics are weighted and summed to obtain the quantified rope tension state.

[0096] In another optional embodiment, the rope tension state of the traction rope is determined based on the curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics of the traction rope. Alternatively, the curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics can be input into a pre-trained rope tension state classification model to obtain the rope tension state type and its corresponding confidence level output by the rope tension state classification model.

[0097] Specifically, the training process of the rope tension state classification model may include: determining sample traction rope images pre-labeled with rope tension state types, and calculating the curvature features, vibration spectrum features, and curve morphology features of each sample traction rope image using the same method described above; inputting the curvature features, vibration spectrum features, and curve morphology features of each sample traction rope image into a pre-determined lightweight classification model; calculating the model loss and fine-tuning the model based on the rope tension state type predicted by the lightweight classification model and the rope tension state type labeled in the sample traction rope images; repeating the above process until the accuracy of the model prediction meets the pre-set accuracy requirements, thus obtaining the rope tension state classification model.

[0098] The rope tension state type output by the rope tension state model can include taut, slack, and tension-free states, and the confidence level is used to represent the probability of belonging to that rope tension state type.

[0099] Furthermore, to improve the accuracy of rope tension state detection, the maximum curvature of the three-dimensional traction rope curve and the dominant lateral vibration frequency can be added to the input of the rope tension state model. Specifically, the tension state feature vector can be constructed using the following formula as the input of the rope tension state model: ,in, This represents the maximum curvature of the three-dimensional traction rope curve.

[0100] In this embodiment, mechanical quantitative indicators based on physical laws, such as curvature, vibration spectrum, and morphological energy, are introduced. These indicators are the essential reflection of the force state of the traction rope and are not affected by surface visual features such as color, texture, and brightness. They can significantly improve the accuracy of traction rope detection and state judgment, and improve the robustness in complex scenarios (such as night, rainy and foggy days, and grassy areas).

[0101] S2100. Based on the video frame image, determine the target pet speed at the target pet center point, the gripping point speed at the target person's hand gripping center point, and the lateral swing speed at the midpoint of the leash, and determine the swing angle of the leash relative to the target pet and the target person.

[0102] In this embodiment, in order to analyze the motion coupling characteristics between the target pet, the target person, and the leash, multi-source motion signal extraction is first performed.

[0103] Specifically, the target pet's speed at the center point of the target pet is determined using the following formula: , Indicates the target pet's speed. Indicates the center point of the target pet.

[0104] The gripping speed at the center point of the target person's hand can be determined using the following formula: ,in, G(t) represents the gripping point velocity, and G(t) represents the center point of the target person's hand gripping the target person.

[0105] The lateral swing velocity at the midpoint of the traction rope is determined by the following formula: ,in, This indicates the midpoint of the traction rope.

[0106] The swing angle of the leash relative to the target pet and the target person is determined using the following formula: .

[0107] S2110. Based on the target pet's speed, gripping point speed, and lateral swing speed, determine the causal dominant characteristics of the target person's movement on the target pet's movement.

[0108] Specifically, the transfer entropy from the grip point velocity to the lateral swing velocity can be calculated using the grip point velocity and the lateral swing velocity. This entropy is used to measure the amount of information transferred from the target person's movement to the traction rope's movement. It can be calculated using the following formula: .in, Indicates the delay time.

[0109] Similarly, by calculating the transfer entropy from the target pet's speed to its lateral swing speed using the target pet's speed and lateral swing speed, we can measure the amount of information transferred from the target pet's movement to the leash's movement. This can be calculated using the following formula: .

[0110] Based on the transfer entropy from the gripping point velocity to the lateral swing velocity, and the transfer entropy from the target pet's velocity to the lateral swing velocity, the causal dominant feature of the target character's movement on the target pet's movement is calculated. Specifically, this can be calculated using the following formula: ,in, This represents the causal correlation coefficient.

[0111] S2120. Determine the cross-correlation coefficient between the grip point speed and the lateral swing speed, and determine the peak time delay based on the cross-correlation coefficient.

[0112] In this embodiment, the cross-correlation coefficient between the grip point velocity and the lateral swing velocity can be expressed by the following formula: ,in, This represents the standard deviation of the grip point velocity sequence within the time window. It represents the standard deviation of the lateral oscillation velocity sequence within the time window.

[0113] In this embodiment, a time delay is used. As the independent variable, with Maximize as the objective, determine Maximize the corresponding peak time delay .

[0114] Understandably, when a pet is on a leash, the person's hand movements will slightly precede the leash's response, so the peak time delay should be a small positive value.

[0115] S2130. Determine the speed correlation based on the target pet's speed and the gripping point speed; determine the swing angle standard deviation based on the swing angle; determine the average tension index based on the rope tension state of the leash; and determine the rope motion entropy based on the lateral swing speed.

[0116] In this embodiment, the speed correlation between the target pet's speed and the gripping point speed can be calculated using the following formula: Where w represents the number of frames corresponding to the time window length. For example, if the time window is 1 second, w can take the value 30.

[0117] The standard deviation of the swing angle can be calculated using the following formula: .

[0118] The average tension index can be calculated using the following formula: .

[0119] The entropy of rope motion can be calculated using the following formula: .in, This represents the probability of the i-th speed interval occurring. The speed interval can be determined as follows: determine the unit direction vector of the rope, and based on this, determine the vertical unit vector; based on the vertical unit vector, determine the vertical velocity at the midpoint of the rope; divide the vertical velocity at the midpoint of the rope into equal-width bins to obtain a predetermined number of speed intervals, for example, 8; and calculate the probability of occurrence of each speed interval. The corresponding speed range The sum is 1.

[0120] S2140. Based on the causal dominant characteristics, time delay peak, velocity correlation, swing angle standard deviation, average tension index, and rope motion entropy, determine the motion coupling characteristics between the target pet, the target person, and the leash.

[0121] In this embodiment, a time-based motion coupling feature is constructed, and the feature is concatenated within a sliding window to obtain the motion coupling feature. w represents the length of the sliding window, which can be 1 second for example.

[0122] Specifically, the motion coupling characteristics can be represented by the following formula: .

[0123] S2150. Input the hand gripping state, the rope tension state, and the motion coupling features into a pre-trained temporal classification model to obtain at least two pet traction states output by the temporal classification model and the confidence level matching the pet traction state.

[0124] The three-level linkage pet leash status detection process based on different dimensions, and the specific process of subsequent risk warning based on pet leash status have been described in the above embodiments, and will not be repeated here.

[0125] It should be noted that in this embodiment, the determination of the hand grip state, the rope tension state, and the motion coupling characteristics are performed separately and simultaneously, and are not limited by the order of steps in this embodiment.

[0126] The technical solution of this embodiment achieves microscopic hand contact determination by classifying the target person's hand gripping state based on key hand points at the visual level; by classifying rope tension state based on curvature and vibration spectrum at the mechanical level, it achieves mesoscopic rope tension quantification; and by determining the coupled motion characteristics based on the cross-correlation of transfer entropy and time delay at the kinematic level, it achieves macroscopic motion coupling analysis of the person, rope, and pet. By combining vision, mechanics, and kinematics, a multi-dimensional feature fusion multi-layer comprehensive decision-making mechanism is achieved, avoiding overall misjudgment caused by the failure of a single feature, improving accuracy and robustness in complex scenarios, and enabling precise judgment of the pet's leash state.

[0127] Example 3 Figure 4 This is a schematic diagram of a pet leash detection device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: The leash detection module 310 is used to determine, based on video frame images, the target pet, the target person matching the target pet, and the leash matching the target pet and the target person. The multi-feature extraction module 320 is used to determine the hand gripping state of the target person, the rope tension state of the leash, and the motion coupling features between the target pet, the target person, and the leash based on the video frame images. The pet leash state determination module 330 is used to determine the pet leash state based on the hand grip state, the rope tension state, and the motion coupling characteristics.

[0128] Based on the above embodiments, optionally, the multi-feature extraction module 320 includes: The bounding box determination unit is used to extract the key points of the hand of the target person based on the video frame image, and determine the bounding box of the key points of the hand. The hand feature determination unit is used to determine the bounding box tightness feature based on the bounding box of the hand key points, and to determine the finger curvature feature, palm depth feature and hand contour feature based on the hand key points. The hand grip state determination unit is used to determine the hand grip state of the target person based on the tightness characteristics of the enclosure box, the curvature characteristics of the fingers, the depth characteristics of the palm, and the contour characteristics of the hand.

[0129] Based on the above embodiments, optionally, the traction rope detection module 310 includes: The hand grip center point determination unit is used to determine the hand grip center point of the target person based on video frame images; The search area determination unit is used to determine the search area based on the center point of the hand grip and the center point of the target pet. The three-dimensional leash curve determination unit is used to determine a binary region based on the search area, and to perform depth estimation and curve fitting on the binary region to obtain a three-dimensional leash curve that matches the target pet and the target person.

[0130] Based on the above embodiments, optionally, the multi-feature extraction module 320 includes: The traction rope feature determination unit is used to determine the curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics of the traction rope based on the three-dimensional traction rope curve. The rope tension state determination unit is used to determine the rope tension state of the traction rope based on the curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics of the traction rope.

[0131] Based on the above embodiments, optionally, the multi-feature extraction module 320 includes: The speed and swing angle determination unit is used to determine the target pet speed at the target pet center point, the gripping point speed at the target person's hand gripping center point, and the lateral swing speed at the midpoint of the leash based on the video frame image, and to determine the swing angle of the leash relative to the target pet and the target person. The causal dominant feature determination unit is used to determine the causal dominant features of the target person's movement on the target pet's movement based on the target pet's speed, gripping point speed, and lateral swing speed. The time delay peak determination unit is used to determine the cross-correlation coefficient between the grip point speed and the lateral swing speed, and to determine the time delay peak based on the cross-correlation coefficient; The motion characteristic determination unit is used to determine the speed correlation based on the target pet's speed and the gripping point speed, determine the swing angle standard deviation based on the swing angle, determine the average tension index based on the rope tension state of the leash, and determine the rope motion entropy based on the lateral swing speed. The motion coupling feature determination unit is used to determine the motion coupling features between the target pet, the target person, and the leash based on causal dominant features, peak time delay, velocity correlation, swing angle standard deviation, average tension index, and rope motion entropy.

[0132] Based on the above embodiments, optionally, the pet leash status determination module 330 includes: The pet leash state and confidence determination unit is used to input the hand grip state, the rope tension state and the motion coupling features into a pre-trained temporal classification model to obtain at least two pet leash states output by the temporal classification model and the confidence levels that match the pet leash states.

[0133] Optionally, based on the above embodiments, the apparatus further includes: The traction force determination unit is used to determine the traction force based on the center point of the target person's hand grip and the center point of the target pet. The no-leash status alarm unit is used to determine that the current pet is in a no-leash state if the confidence level of determining that the pet is in a no-leash state is greater than or equal to a preset confidence level threshold, and to trigger a target pet leash status alarm. The current pet leash status determination unit is used to determine the current pet leash status as the pet leash status with the highest confidence level; The target pet leash status alarm unit is used to trigger a target pet leash status alarm if it is determined that the current pet leash status is invalid, or if it is determined that the current pet leash status is passive following, and the traction force is greater than or equal to a preset tension threshold.

[0134] The pet leash detection device provided in this embodiment of the invention can execute the pet leash detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0135] Example 4 Figure 5A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0136] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as pet leash detection methods.

[0139] In some embodiments, the pet leash detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the pet leash detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the pet leash detection method by any other suitable means (e.g., by means of firmware).

[0140] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0141] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable pet leash detection device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0142] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0143] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0144] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0145] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0146] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0147] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for detecting pet leashes, characterized in that, include: Based on video frame images, identify the target pet, the target person matching the target pet, and the leash matching the target pet and the target person; Based on the video frame images, determine the hand gripping state of the target person, the rope tension state of the leash, and the motion coupling characteristics between the target pet, the target person, and the leash; Determine the motion coupling characteristics between the target pet, the target person, and the leash, including: Based on the video frame images, determine the target pet speed at the target pet center point, the gripping point speed at the target person's hand gripping center point, and the lateral swing speed at the midpoint of the leash, and determine the swing angle of the leash relative to the target pet and the target person. Based on the target pet's speed, gripping point speed, and lateral swing speed, determine the causal dominant characteristics of the target person's movement on the target pet's movement; Determine the cross-correlation coefficient between the grip point speed and the lateral swing speed, and determine the peak time delay based on the cross-correlation coefficient; Based on the target pet's speed and the gripping point speed, determine the speed correlation; based on the swing angle, determine the swing angle standard deviation; based on the rope tension state, determine the average tension index; and based on the lateral swing speed, determine the rope motion entropy. Based on the causal dominant characteristics, time delay peak, velocity correlation, swing angle standard deviation, average tension index, and rope motion entropy, the motion coupling characteristics between the target pet, the target person, and the leash are determined. The pet traction state is determined based on the hand grip state, the rope tension state, and the motion coupling characteristics.

2. The method according to claim 1, characterized in that, Determining the target person's hand gripping state based on video frame images includes: Based on the video frame images, extract the key points of the target person's hands and determine the bounding box of the key points of the hands; Based on the bounding boxes of key hand points, determine the bounding box tightness features, and based on the key hand points, determine the finger curvature features, palm depth features, and hand contour features. The hand gripping state of the target person is determined based on the tightness of the enclosing box, the curvature of the fingers, the depth of the palm, and the contour of the hand.

3. The method according to claim 1, characterized in that, Based on video frame images, determine the leash that matches the target pet and the target person, including: Based on the video frame images, determine the center point of the target person's hand grip; The search area is determined based on the center point of the hand grip and the center point of the target pet. A binary region is determined based on the search area, and depth estimation and curve fitting are performed on the binary region to obtain a three-dimensional leash curve that matches the target pet and the target person.

4. The method according to claim 3, characterized in that, Determining the rope tension state of the traction rope based on video frame images includes: Based on the three-dimensional traction rope curve, the curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics of the traction rope are determined. The rope tension state of the traction rope is determined based on its curvature characteristics, vibration spectrum characteristics, and curve morphology characteristics.

5. The method according to any one of claims 1-4, characterized in that, The pet traction state is determined based on the hand grip state, the rope tension state, and the motion coupling characteristics, including: The hand gripping state, the rope tension state, and the motion coupling features are input into a pre-trained temporal classification model to obtain at least two pet traction states output by the temporal classification model and the confidence level matching the pet traction state.

6. The method according to claim 5, characterized in that, The method further includes: The traction force is determined based on the center point of the target person's hand grip and the center point of the target pet. After obtaining at least two pet leash states output by the time-series classification model and the confidence scores matching the pet leash states, the method further includes: If the confidence level that the pet is in an unleashed state is greater than or equal to the preset confidence threshold, then the current pet is determined to be in an unleashed state, and an alarm for the target pet's leash status is triggered. Otherwise, determine the current pet leash status as the pet leash status with the highest confidence level; If it is determined that the current pet leash status is invalid, or if it is determined that the current pet leash status is passive following and the traction force is greater than or equal to the preset tension threshold, then the target pet leash status alarm will be triggered.

7. A pet leash detection device, characterized in that, include: The leash detection module is used to determine the target pet, the target person matching the target pet, and the leash matching the target pet and the target person based on video frame images. The multi-feature extraction module is used to determine the hand gripping state of the target person, the rope tension state of the leash, and the motion coupling features between the target pet, the target person, and the leash based on video frame images. The multi-feature extraction module includes: The speed and swing angle determination unit is used to determine the target pet speed at the target pet center point, the gripping point speed at the target person's hand gripping center point, and the lateral swing speed at the midpoint of the leash based on the video frame image, and to determine the swing angle of the leash relative to the target pet and the target person. The causal dominant feature determination unit is used to determine the causal dominant features of the target person's movement on the target pet's movement based on the target pet's speed, gripping point speed, and lateral swing speed. The time delay peak determination unit is used to determine the cross-correlation coefficient between the grip point speed and the lateral swing speed, and to determine the time delay peak based on the cross-correlation coefficient; The motion characteristic determination unit is used to determine the speed correlation based on the target pet's speed and the gripping point speed, determine the swing angle standard deviation based on the swing angle, determine the average tension index based on the rope tension state of the leash, and determine the rope motion entropy based on the lateral swing speed. The motion coupling feature determination unit is used to determine the motion coupling features between the target pet, the target person, and the leash based on the causal dominant features, time delay peak, velocity correlation, swing angle standard deviation, average tension index, and rope motion entropy. The pet leash state determination module is used to determine the pet leash state based on the hand grip state, the rope tension state, and the motion coupling characteristics.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the pet leash detection method as described in any one of claims 1-6.

9. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the pet leash detection method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Detection method and device of pet pulling rope, monitoring equipment and storage medium

    CN120510179A

  • Pet guy rope state detection method and system based on collaborative motion analysis

    CN121686322A