Method and device for identifying box lock catch inspection behavior, and electronic equipment

Automatically identify the box lock inspection behavior through computer vision technology, solving the problem of poor subjectivity and real-time human inspection in the existing technology, achieving efficient and accurate lock inspection, and improving the security of bank vault management.

CN120544261APending Publication Date: 2025-08-26HANVON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536482.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the monitoring method for checking box locks relies on manual observation, which has subjectivity, easy to miss inspection and misjudgment, and poor real-time performance, making it difficult to meet the needs of high-frequency and large-scale management.

Method used

Using a computer vision-based method, the lock inspection behavior in the video image is automatically identified through box key point detection, hand detection and feature analysis, and combined with the recognition of human posture key point recognition, to improve the accuracy and efficiency of inspection.

Benefits of technology

It realizes efficient, accurate and automatic identification of box lock inspection behavior, reduces negligence and subjective misjudgment of manual inspection, and improves real-time and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544261A_ABST
    Figure CN120544261A_ABST
Patent Text Reader

Abstract

The invention discloses a recognition method and device for a box lock catch inspection behavior and electronic equipment. The method comprises the following steps: respectively carrying out box body key point detection and hand detection on a video image of a box body lock catch inspection scene to obtain position information of a lock catch point and a hand detection frame; based on the position information, obtaining a matching result of the hand detection frame and the lock catch point; if the matching is successful, performing hand key point detection on a target image area determined according to the position of the successfully matched hand detection frame in the video image, and obtaining position information of hand key points; performing feature analysis on the hand key points based on the position information to obtain an analysis result indicating whether a check behavior aiming at the successfully matched lock catch exists in the video image, and if yes, performing human body posture key point identification on the video image to obtain human body posture key points; and based on the human body posture key points, the target locking point and the hand key points, obtaining an identification result of the box body locking inspection behavior in the video image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of behavior recognition technology, and in particular to a method, device, electronic device, and computer-readable storage medium for identifying box lock inspection behavior. Background Art

[0002] With the rapid development of artificial intelligence (AI), computer vision-based behavior recognition technology has been widely applied in various fields, such as security monitoring, behavioral analysis, and workflow management. For example, in the management of cash drawers in bank vaults, ensuring the integrity of the locks is crucial for regulating work practices and preventing financial risks. Existing techniques for checking locks still rely on manual observation and judgment. The method involves: personnel manually inspecting the locks before delivery to determine if they are damaged or abnormal. Furthermore, management personnel use monitoring systems to observe whether personnel have inspected the locks in accordance with regulations, thereby monitoring the lock inspection process.

[0003] As can be seen, existing methods for monitoring cash box lock inspections have at least the following drawbacks: manual inspection and monitoring are highly subjective, prone to omissions or misjudgments due to fatigue, negligence, and other factors; they are labor-intensive and lack real-time performance, making them difficult to meet the needs of high-frequency and large-scale management. Existing methods for identifying cash box lock inspections still need improvement. Summary of the Invention

[0004] The embodiments of the present application provide a method and device for identifying box lock buckle inspection behavior, which helps to improve the accuracy and efficiency of box lock buckle inspection behavior monitoring, thereby effectively improving the real-time and safety of the box lock buckle inspection link.

[0005] In a first aspect, an embodiment of the present application provides a method for identifying a box lock check behavior, comprising:

[0006] Performing box key point detection on a video image of a box lock inspection scene to obtain position information of the box key points, and performing hand detection on the video image to obtain position information of a hand detection frame, wherein the box key points include: lock points;

[0007] Based on the position information of the locking point and the position information of the hand detection frame, obtaining a matching result between the hand detection frame and the locking point;

[0008] In response to the matching result indicating the successful matching of the hand detection frame and the lock point, performing hand key point detection on a target image area of ​​the video image to obtain position information of the hand key points, wherein the target image area is determined according to the position of the successfully matched hand detection frame;

[0009] Based on the position information of the hand key points, feature analysis is performed on the hand key points to obtain an analysis result indicating whether there is an inspection behavior for a target lock buckle in the video image, the target lock buckle being: the lock buckle corresponding to the lock buckle point that has been successfully matched;

[0010] In response to the analysis result indicating that there is an inspection behavior for the target lock in the video image, performing human posture key point recognition on the video image to obtain the human posture key points;

[0011] Based on the human body posture key points, the target locking points and the hand key points, a recognition result of the box locking inspection behavior in the video image is obtained.

[0012] In a second aspect, an embodiment of the present application provides a device for identifying a box lock check behavior, comprising:

[0013] The first detection module is configured to perform box key point detection on a video image of a box lock inspection scene to obtain position information of the box key points, and to perform hand detection on the video image to obtain position information of a hand detection frame, wherein the box key points include lock points;

[0014] a matching module, configured to obtain a matching result between the hand detection frame and the locking point based on the position information of the locking point and the position information of the hand detection frame;

[0015] a second detection module, configured to, in response to the matching result indicating the successful matching of the hand detection frame and the lock point, perform hand key point detection on a target image area of ​​the video image to obtain position information of the hand key points, wherein the target image area is determined based on the position of the successfully matched hand detection frame;

[0016] a first behavior recognition module, configured to perform feature analysis on the hand key points based on the position information of the hand key points, and obtain an analysis result indicating whether there is an inspection behavior for a target lock buckle in the video image, wherein the target lock buckle is the lock buckle corresponding to the lock buckle point that has been successfully matched;

[0017] a human posture key point acquisition module, configured to, in response to the analysis result indicating that there is an inspection behavior for the target lock in the video image, perform human posture key point recognition on the video image to acquire human posture key points;

[0018] The second behavior recognition module is used to obtain the recognition result of the box lock inspection behavior in the video image based on the human body posture key points, the target lock points and the hand key points.

[0019] On the third aspect, an embodiment of the present application also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the computer program, it implements the method for identifying the box lock inspection behavior described in the embodiment of the present application.

[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method for identifying the box lock inspection behavior disclosed in the embodiment of the present application are implemented.

[0021] The embodiment of the present application discloses a method for identifying a box lock inspection behavior, which performs box key point detection on a video image of a box lock inspection scene to obtain position information of the box key points, and performs hand detection on the video image to obtain position information of a hand detection frame, wherein the box key points include: lock points; based on the position information of the lock points and the position information of the hand detection frame, obtain a matching result of the hand detection frame and the lock points; in response to the matching result indicating a successful match between the hand detection frame and the lock point, perform hand key point detection on a target image area of ​​the video image to obtain position information of the hand key points, wherein, The target image area is determined based on the position of the successfully matched hand detection frame; based on the position information of the hand key points, the hand key points are subjected to feature analysis to obtain an analysis result indicating whether an inspection action for a target lock buckle is present in the video image, where the target lock buckle is the lock buckle corresponding to the successfully matched lock buckle point; in response to the analysis result indicating the presence of an inspection action for the target lock buckle in the video image, human posture key points are identified on the video image to obtain human posture key points; and based on the human posture key points, the target lock buckle points, and the hand key points, an identification result of the case lock buckle inspection action in the video image is obtained. This method combines image processing and target detection technologies to automatically identify whether a case lock buckle inspection action is present in a video image, thereby improving the efficiency of the case lock buckle inspection and enhancing inspection accuracy and real-time performance. This method effectively avoids the problem of manual inspection methods that are prone to missed inspections or misjudgments due to operator negligence or subjective factors, which affects the accuracy of inspection results and the security of vault management.

[0022] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] Figure 1 This is one of the step flow charts of the method for identifying the box lock inspection behavior disclosed in the embodiment of the present application;

[0025] Figure 2 is a schematic diagram of a video image of a box lock inspection behavior disclosed in an embodiment of the present application;

[0026] Figure 3 The identification method of the box lock inspection behavior disclosed in the embodiment of this application is Figure 2 Schematic diagram of the result of box key point detection on the video image shown;

[0027] Figure 4 The identification method of the box lock inspection behavior disclosed in the embodiment of this application is Figure 2 Schematic diagram of the result of hand detection on the video image shown;

[0028] Figure 5 The identification method of the box lock inspection behavior disclosed in the embodiment of this application is based on Figure 3 and Figure 4 Schematic diagram of distance matching effect of detection results;

[0029] Figure 6 This is a schematic diagram of the comprehensive recognition principle of the identification method for the box lock inspection behavior disclosed in the embodiment of the present application;

[0030] Figure 7 This is a schematic diagram of the structure of the device for identifying the box lock inspection behavior disclosed in the embodiment of the present application;

[0031] Figure 8 A block diagram schematically shows an electronic device for executing the method according to the present application; and

[0032] Figure 9 The figure schematically shows a storage unit for storing or carrying a program code for implementing the method according to the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] The following combination Figure 1 The flowchart shown illustrates the specific implementation of the present application in detail.

[0035] Reference Figure 1 , a method for identifying box lock inspection behavior disclosed in an embodiment of the present application includes: steps 102 to 112.

[0036] Step 102: perform box key point detection on the video image of the box lock inspection scene to obtain position information of the box key points, and perform hand detection on the video image to obtain position information of the hand detection frame, wherein the box key points include: lock points.

[0037] The box can be a cash box or a box of other standard shapes. The video image can be an image frame in a surveillance video of a box lock inspection scene. In the embodiment of the present application, there is no limitation on the specific implementation method of obtaining a video image of a box lock inspection scene. The video image is an image of the front of the box (i.e., the side where the lock is located). The schematic diagram of the video image is as follows Figure 2 shown.

[0038] Optionally, the performing of box key point detection on the video image of the box lock inspection scene to obtain the position information of the box key points includes: performing box detection on the video image of the box lock inspection scene to obtain the position information of the box detection frame; based on the position information of the box detection frame, enlarging the box detection frame to obtain an enlarged image area; performing box key point detection on the enlarged image area in the video image to obtain the position information of the box key points.

[0039] In some optional embodiments, a pre-trained box object detection model can be used to perform box detection on the video image, and a box detection result output by the box object detection model can be obtained. The box detection result includes: the location information and confidence level of the box detection frame. Furthermore, the box detection frame with the highest confidence level that meets a preset confidence threshold is obtained as the final box detection frame detected from the video image, and the location information of the box detection frame is recorded. The box object detection model can be trained based on YOLOv8 (You Only Look Once version 8).

[0040] Then, based on the position information of the box detection frame, the box detection frame is enlarged to obtain an enlarged image area, which covers more image content to improve the accuracy of box key point detection. For example, the width and height of the image area of ​​the box detection frame in the video image are enlarged by 2 / 3 of the original size to obtain the enlarged image area.

[0041] In some optional embodiments, taking the position information of the original box detection frame as [x1, y1, x2, y2] as an example, where x1, y1 represent the coordinates of the upper left corner of the box detection frame, x2, y2 represent the coordinates of the lower right corner of the box detection frame, and the width and height of the video image are width and height, the position information of the enlarged image area can be calculated using the following formula:

[0042]

[0043] in, Indicates the coordinates of the upper left corner of the enlarged image area, Indicates the coordinates of the lower right corner of the enlarged image area.

[0044] Afterwards, a pre-trained box key point detection model is used to perform box key point detection on the image content in the enlarged image area to obtain the box key point detection result. The box key point detection model can be trained based on the YOLOv8-pose (an improved model based on YOLOv8). The box key point detection result includes: the position information and confidence of the preset box key points. The box key points include: the four corner points and locking points of the box, wherein the four corner points are: the first corner point, the second corner point, the third corner point, and the fourth corner point; taking the box as an example, there are two locking points, namely: the first locking point and the second locking point. Figure 3 Taking the box key point detection results shown as an example, the first corner point 312 is the upper left corner point, the second corner point 314 is the upper right corner point, the third corner point 316 is the lower left corner point, the fourth corner point 318 is the lower right corner point, the first locking point 322 is the left locking point, and the second locking point 324 is the right locking point.

[0045] In order to reduce the adverse effects of errors in key point prediction on subsequent behavior recognition based on box key points, in some embodiments of the present application, geometric constraints are introduced to correct the predicted box key points to limit the positions of the key points. The geometric constraints include: constraint 1, the lower left corner point, the lower right corner point, the left lock point and the right lock point are on the same straight line; constraint 2, the distance from the lower left corner point to the left lock point is equal to the distance from the lower right corner point to the right lock point. For constraint 1, the weighted least squares fitting method can be used to first fit the four points to a straight line. For constraint 2, the left lock point and the right lock point of the lock can be translated using only the distance relationship.

[0046] Optionally, after performing box key point detection on the enlarged image area in the video image and obtaining the position information of the box key points, it also includes: using a weighted least squares fitting method to fit the third corner point, the fourth corner point, the first locking point and the second locking point to a straight line; adjusting the position information of the first locking point and the second locking point with the distance between the third corner point to the first locking point being equal to the distance between the fourth corner point to the second locking point, and the locking position adjustment amount being minimized as a constraint.

[0047] As mentioned above, when the box key point detection is performed, the detection result obtained also includes: the confidence level of the box key point. In the embodiment of the present application, the fitting error is based on the confidence level of the box key point.

[0048] Optionally, for constraint 1, you can use a weighted least squares fit to fit the four points to a straight line. For constraint 2, you can translate the left and right lock points using only the distance relationship.

[0049] Below, the coordinates of the lower left corner, left lock point, lower right corner, and right lock point are (x1, y1, p1), (x2, y2, p2), (x3, y3, p3), and (x4, y4, p4), respectively. i ,y i ) is the coordinate, p i is the confidence level predicted by the box key point detection model. Taking i=1, 2, 3, and 4 as examples, the correction method for the box key points is described.

[0050] When the lower left corner point, the left locking point, the lower right corner point, and the right locking point are fitted to the same straight line, the equation of the straight line can be expressed as: y=kx+b.

[0051] In the prediction results of the box key point detection model, points with higher confidence are usually predicted more accurately. Based on this, in the embodiment of the present application, the weighted least squares method is introduced to fit the box key points. The error function E(k, b) is defined to represent the sum of the squares of the vertical distances from the point to the line. The error function is minimized to fit the parameters of the line equation. The error function E(k, b) is calculated by the following formula:

[0052]

[0053] Among them, i and j represent the numbers of the key points of the box, p i and p j Indicates the confidence of the key points of the corresponding box. The optimal parameter k is obtained by numerical calculation method. * and b * .

[0054] For constraint 2, define the distance from the lower left corner to the left locking point as d1, and the distance from the lower right corner to the right locking point as d2, where:

[0055]

[0056] In order to satisfy the constraint condition d1=d2, the coordinates of the left and right locking points are adjusted by translation. Let the adjustment amount of the left locking point be (Δx2, Δy2), and the adjustment amount of the right locking point be (Δx3, Δy3). The distance relationship after adjustment is:

[0057]

[0058] By introducing the constraint min((Δx2) 2 +(Δy2) 2 +(Δx3) 2 +(Δy3) 2 ), so that the distance relationship holds true while keeping the translation as small as possible.

[0059] The final correction result can be obtained through geometric optimization so that the key points of the box meet the constraints of constraints 1 and 2.

[0060] By correcting the key points of the box based on the least squares method, the positioning accuracy of the key points and locking points of the box can be improved.

[0061] On the other hand, a pre-trained hand detection model can be used to detect hands on the video image to obtain the position information of the hand detection frame. The hand detection model can be trained based on YOLOv8. Figure 2 As an example, the video image shown in Figure 2 The hand detection of the video image shown in , can be obtained as follows Figure 4The test results shown.

[0062] The specific training methods of the box object detection model, the box key point detection model and the hand detection model can be found in the prior art and will not be repeated in the embodiments of this application.

[0063] In other optional embodiments, other target detection methods may be used to detect the video image to obtain box detection results, box key point detection results, and header detection results.

[0064] It should be noted that when box object detection and box key point detection are required for different boxes, different box images can be used as training samples to train the corresponding models.

[0065] This step performs box detection, box key point detection, and hand detection on the video image. This yields the positional information for the box detection frames, the positional information and confidence scores for the box key points within each box detection frame, and the positional information for the hand detection frames. Next, based on this information, the hand detection frames are matched with the lock points.

[0066] Step 104: Based on the position information of the locking point and the position information of the hand detection frame, a matching result between the hand detection frame and the locking point is obtained.

[0067] In some optional embodiments, the obtaining of the matching result of the hand detection frame and the locking point based on the position information of the locking point and the position information of the hand detection frame includes: calculating the relative distances from each locking point to each hand detection frame based on the position information of the locking point and the position information of the hand detection frame; when the minimum value of the relative distances meets a preset distance condition, determining the hand detection frame and the locking point corresponding to the minimum value of the relative distances as the successfully matched hand detection frame and the locking point, and taking the successfully matched hand detection frame and the locking point as the matching result; when the minimum relative distance does not meet the preset distance condition, obtaining a matching result indicating that the matching of the hand detection frame and the locking point fails.

[0068] Optionally, the relative distance from the locking point to the hand detection frame is: the ratio of the distance between the locking point and the center point of the hand detection frame to half the diagonal distance of the hand detection frame.

[0069] For example, if n hand detection frames are detected in a video image, the distance between the center point of each hand detection frame and each locking point of the box is calculated. A greedy algorithm is then used to associate and match the hand detection frames with the locking points. Based on the matching results, the distance between the center point of each hand detection frame and the corresponding locking point is calculated. If the distance is greater than or equal to the set threshold, the locking check is not performed. If the distance is less than the threshold, a check may have occurred, and further analysis of hand features is required.

[0070] The n hand detection frames detected in the video image are represented as ...as an example, the center point of the hand detection box is calculated using the following formula:

[0071]

[0072] Among them, xc i and yc i Represents the coordinates of the center point of the i-th hand detection box.

[0073] Since the distance of the video image shooting angle will cause the size of the target object in the video image to vary, in the embodiment of the present application, the relative distance is used to match the hand detection frame and the lock point. Optionally, the relative distance valid_d1 from the lock point to the hand detection frame can be calculated by the following formula:

[0074] Among them, x j and y j is the position coordinate of the locking point j.

[0075] Optionally, the preset distance condition may be: a relative distance less than 1. That is, when the relative distance valid_d1 between the hand detection frame and a certain lock point is less than 1, the hand detection frame is considered to be a valid hand detection frame; otherwise, the hand detection frame is considered to be an invalid hand detection frame.

[0076] In the actual box lock inspection scenario, the left and right hands of multiple people may approach the box at the same time. Therefore, in the embodiment of the present application, a greedy algorithm is adopted to simplify the allocation, and only the hand corresponding to the valid hand detection frame with the minimum relative distance valid_d1 is taken as the hand for checking the lock. That is, the hand detection frame and the lock point corresponding to the minimum value of the relative distance valid_d1 are taken as the successfully matched hand detection frame and the lock point, and the successfully matched hand detection frame and the lock point are taken as the matching result. The other hand detection frames are confirmed as having no matching lock points. The matching effect of the hand detection frame and the lock point is as follows. Figure 5 shown. Figure 5The middle hand detection frame 512 matches the left lock point 522 , the hand detection frame 514 matches the left lock point 524 , and the hand detection frame 516 does not match any lock point.

[0077] If the result of the distance-based matching judgment in the current video image is that there is a successfully matched hand detection frame and locking point, step 106 is executed to further determine whether the hand action is an action to check the locking point based on the image features; if the result of the distance-based matching judgment in the current video image is that there is no successfully matched hand detection frame and locking point, the recognition process executed on the current video image is terminated.

[0078] By matching the lock points with the hand detection frame, the amount of data processed in subsequent steps can be reduced, while providing high-precision data support for subsequent behavior judgment.

[0079] In some optional embodiments, in response to the matching result indicating that the hand detection frame and the lock point fail to match, the next frame of video image is acquired, and step 102 and subsequent steps are repeated.

[0080] Step 106: In response to the matching result indicating the successful matching of the hand detection frame and the locking point, hand key point detection is performed on the target image area of ​​the video image to obtain the position information of the hand key points, wherein the target image area is determined according to the position of the successfully matched hand detection frame.

[0081] Optionally, the target image area is the hand detection frame that is matched successfully (such as Figure 5 The image area is obtained by enlarging the hand detection frames 512 and 514 in the image by a preset ratio. The preset ratio may be 2 / 3.

[0082] In some optional embodiments, a pre-trained hand key point detection model may be used to perform hand key point detection on the image content of the target image area in the video image to obtain position information of the hand key points of the hand corresponding to the hand detection frame. The hand key point detection model may be fine-tuned based on YOLOv8-pose, and the detected hand key points may include 21 hand key points, including fingertips, knuckles, and wrists.

[0083] In other optional embodiments, other methods may be used to detect hand key points in the target image area of ​​the video image to obtain the position information of the hand key points. The embodiments of the present application do not limit the specific implementation method of performing hand key point detection.

[0084] Step 108: Based on the position information of the hand key points, feature analysis is performed on the hand key points to obtain an analysis result indicating whether there is an inspection behavior for a target lock buckle in the video image, and the target lock buckle is: the lock buckle corresponding to the lock buckle point that has been successfully matched.

[0085] Next, the position information of the hand key points detected from the video image is encoded to obtain the hand key point features, and based on the hand key point features, it is identified whether the inspection behavior is the target lock.

[0086] In some optional embodiments, the feature analysis of the hand key points is performed based on the position information of the hand key points to obtain an analysis result indicating whether there is an inspection behavior for the target lock in the video image, including: obtaining hand key point features based on the position information of the hand key points; matching the hand key point features with the pre-acquired hand motion feature templates of the standard box lock inspection behavior to obtain feature matching results; in response to the feature matching result indicating a successful match, obtaining an analysis result indicating that there is an inspection behavior for the target lock in the video image, wherein the analysis result includes: the hand detection box and the target lock corresponding to the hand key point; in response to the feature matching result indicating a failed match, obtaining an analysis result indicating that there is no lock inspection behavior in the video image.

[0087] The target lock is: the lock corresponding to the lock point that matches the hand detection frame corresponding to the target image area.

[0088] Optionally, the position information includes: coordinates, and obtaining the hand key point features based on the position information of the hand key points includes: normalizing the coordinates of the position information of the hand key points respectively to obtain the relative position coordinates of the hand key points; performing preset feature encoding processing on the relative position coordinates to obtain the hand key point features corresponding to each of the hand detection frames.

[0089] For example, a pre-trained key point feature encoding network KPNet can be used to perform feature encoding, and the relative position coordinates of the hand key points detected in the above steps can be encoded into a feature vector of a preset dimension as the hand key point features.

[0090] The relative position coordinates of the hand key points obtained after normalization are used for feature encoding processing, which can effectively characterize the hand behavior characteristics.

[0091] During the specific implementation, it is necessary to first collect hand images of the standardized box lock inspection behavior in advance, then detect the hand key points in the hand image, and encode the position information of the hand key points. After that, the obtained hand key point features are stored in the template library as hand action feature templates for subsequent feature matching.

[0092] Optionally, the hand key point features are matched with pre-acquired hand motion feature templates of standard box lock inspection behaviors to obtain feature matching results, including: calculating the cosine similarity between the hand key point features and the pre-acquired hand motion feature templates of standard box lock inspection behaviors; in response to the cosine similarity being greater than or equal to a preset distance threshold, obtaining a feature matching result indicating a successful match; in response to the cosine similarity being less than a preset distance threshold, obtaining a feature matching result indicating a failed match. The value of the preset distance threshold is set based on experimental results. For example, the value of the preset distance threshold can be set to 0.6.

[0093] Optionally, in response to the cosine similarity being greater than or equal to a preset distance threshold, it is determined that the hand key points corresponding to the hand key point features match the hand movements of the standard box lock inspection behavior, and a feature matching result indicating a successful match is obtained; in response to the cosine similarity being less than the preset distance threshold, it is determined that the hand key points corresponding to the hand key point features do not match the hand movements of the standard box lock inspection behavior, and a feature matching result indicating a failed match is obtained.

[0094] When comparing the hand key point features extracted from the current video image with the hand motion feature templates in the template library, respectively calculate the cosine similarity between the hand key point features and the hand motion feature templates in the template library, and then take the maximum value of the cosine similarity for comparison with the preset distance threshold. If the maximum value of the cosine similarity is greater than or equal to the preset distance threshold, it is considered that the gesture in the current video image matches the hand motion of the standard box lock inspection behavior, and thus it is determined that there is a behavior of checking the lock in the current video image. If the maximum value of the cosine similarity is less than the preset distance threshold, it is considered that the gesture in the current video image does not match the hand motion of the standard box lock inspection behavior, and thus it is determined that although the hand is detected in the current video image, there is no behavior of checking the lock. For example, after feature comparison, it can be concluded that Figure 5 The hand key points in the middle hand detection frame 512 match the hand movements of the standard box lock inspection behavior, and the hand key points in the outer hand detection frame 514 do not match the hand movements of the standard box lock inspection behavior.

[0095] In some optional embodiments, in response to the analysis result indicating that there is no inspection behavior of the target lock in the video image, the next frame of the video image is acquired, and step 102 and subsequent steps are repeated.

[0096] Step 110 : In response to the analysis result indicating that there is an inspection behavior for the target lock in the video image, human posture key points are recognized on the video image to obtain human posture key points.

[0097] In some embodiments of the present application, after determining that a hand in a video image is performing an action of checking a lock, it is also necessary to identify the identity of the person, so as to use comprehensive identification to determine the compliance of the lock checking action.

[0098] Among them, the specific implementation method of performing human posture key point recognition on the video image and obtaining human posture key points can be referred to the prior art and will not be repeated in the embodiments of this application. For example, a pre-trained YOLOv8-pose model can be used to detect human posture key points to obtain 17 human posture key points, wherein the obtained human posture key points include but are not limited to: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle.

[0099] Step 112: Based on the human body posture key points, the target locking points, and the hand key points, obtain a recognition result of the box locking inspection behavior in the video image.

[0100] In some optional embodiments, based on the human body posture key points, the target locking points and the hand key points, the recognition results of the box lock inspection behavior in the video image are obtained, including: obtaining the identity recognition result of the video image based on the human body posture key points; obtaining the left and right hand recognition results based on the distance between the wrist key point in the human body posture key points and the wrist key point in the hand key points; obtaining the recognition result of the box lock inspection behavior in the video image based on the identity recognition result, the left and right hand recognition results and the target locking points.

[0101] During implementation, the key point feature templates of the human posture of the occupants are pre-entered and stored in a personnel identity information database. During the identification phase, the key points of the human posture detected from the current video image are feature-encoded and matched with the key point feature templates to obtain the identity information of the person corresponding to the successfully matched key point feature templates. The specific implementation method for obtaining the identity recognition result of the video image based on the key points of human posture is described in the prior art and will not be further described in the present embodiments.

[0102] The obtained identity recognition results include two situations: a match failure, or a match success. In the case of a match success, the number or other identity information of the person who has successfully matched can be further obtained based on the person identity information database.

[0103] Because the human wrist key points and the hand wrist key points are in the same position, the wrist key points (such as the left wrist key point or the right wrist key point) obtained by performing human posture key point detection and the wrist key points obtained by performing hand key point detection can be determined to belong to the same person based on the distance between the wrist key points, thereby determining the human body to which the hand belongs, and whether the hand is the left hand or the right hand, that is, determining the person identity information and left and right hand information matched by the hand detection frame.

[0104] Optionally, obtaining a recognition result of the box lock check behavior in the video image based on the identity recognition result, the left and right hand recognition results, and the target lock point includes: in response to the identity recognition result indicating successful identity recognition, obtaining identity information obtained by the identity recognition; and generating a recognition result of the box lock check behavior in the video image based on the identity information, the left and right hand recognition results, and the lock corresponding to the target lock point. The recognition result indicates which person and which hand checked which lock.

[0105] by Figure 6 Taking the video image shown as an example, if the hand key point detection and feature analysis results indicate that the hand in the hand detection frame performed a lock check operation on the right lock, the identity recognition result is that the person in the video image is: person s, and the distance between the wrist key point obtained by the hand key point detection and the right wrist key point among the human posture key points of person s is the smallest, then the following recognition result of the box lock check behavior can be obtained: person s checked the left lock of the box with his left hand.

[0106] Optionally, obtaining the recognition result of the box lock inspection behavior in the video image based on the identity recognition result, the left and right hand recognition results and the target locking point includes: in response to the identity recognition result indicating that identity recognition has failed, obtaining the recognition result of the presence of abnormal lock inspection behavior in the video image.

[0107] At this point, the recognition results of the box lock inspection behavior in a single-frame video image are completed.

[0108] The comprehensive analysis method of multi-target matching can ensure the accuracy and completeness of the lock inspection behavior judgment.

[0109] In practice, the aforementioned identification process for checking the locks of the box can be continuously executed for each frame of surveillance video within a specified time period. If any violation is detected during the identification process, the monitoring system will issue a timely warning.

[0110] In some optional embodiments, after obtaining the recognition result of the box lock buckle inspection behavior in the video image based on the human body posture key points, the target lock points and the hand key points, it also includes one or more of the following early warning detection methods: in response to the recognition result of the box lock buckle inspection behavior in the video image within a preset time period meeting the preset violation condition, alarm processing is performed based on the recognition result, wherein the preset violation condition includes: the presence of an recognition result of abnormal lock buckle inspection behavior in the video image; in response to obtaining the box lock buckle inspection behavior of the box lock buckle inspection scene meeting the preset compliance condition, obtaining the recognition result of the box lock buckle inspection behavior of the box lock buckle inspection scene being normal, wherein the preset compliance condition includes: there is an inspection behavior of the target person inspecting two locks of the same box, and the number of the target persons is greater than or equal to the preset number threshold.

[0111] For example, if the recognition result of the box lock check behavior in the video image acquired within the preset time period includes the inspection result that the box lock check behavior is performed by a non-personnel, an alarm process can be performed.

[0112] For another example, when N persons are required to check the lock of a certain box, it is necessary to identify in the video images collected within a specified time period that all N persons have performed the operation of checking the two locks of the specified box.

[0113] This method continuously tracks lock inspections by detecting and analyzing the behavior of lock inspections within a specified time period. Non-compliant behaviors based on the detection results can be detected in real time, providing efficient and secure monitoring capabilities for the application system.

[0114] In summary, the method for identifying the box lock inspection behavior disclosed in the embodiment of the present application performs box key point detection on the video image of the box lock inspection scene to obtain the position information of the box key point, and performs hand detection on the video image to obtain the position information of the hand detection frame, wherein the box key point includes: a lock point; based on the position information of the lock point and the position information of the hand detection frame, a matching result of the hand detection frame and the lock point is obtained; in response to the matching result indicating that the hand detection frame and the lock point are successfully matched, hand key point detection is performed on the target image area of ​​the video image to obtain the position information of the hand key point, wherein In this method, the target image area is determined based on the position of the successfully matched hand detection frame; based on the position information of the hand key points, the hand key points are feature analyzed to obtain an analysis result indicating whether an inspection behavior for a target lock buckle exists in the video image, where the target lock buckle is the lock buckle corresponding to the successfully matched lock buckle point; in response to the analysis result indicating the presence of an inspection behavior for the target lock buckle in the video image, human posture key points are identified on the video image to obtain human posture key points; and based on the human posture key points, the target lock buckle points, and the hand key points, an identification result of the box lock buckle inspection behavior in the video image is obtained. This method combines image processing and target detection technologies to automatically identify whether a video image contains a box lock buckle inspection behavior, thereby improving the inspection efficiency of the box lock buckle inspection behavior and enhancing the accuracy of the inspection, with higher real-time performance. This method effectively avoids the problem that manual inspection methods are prone to missed inspections or misjudgments due to operator negligence or subjective factors, which affects the accuracy of inspection results and the security of vault management.

[0115] Taking the example of cash box lock inspection using this method for the first time in the field of bank vault management, by introducing target detection and key point detection technology based on deep learning, the shortcomings of insufficient intelligence of traditional solutions can be compensated, and high-precision and high-real-time behavior recognition can be achieved, further improving the security of bank vault management.

[0116] This method integrates advanced target detection and key point detection technologies to accurately detect the hand position and lock area of ​​the person, which can significantly reduce the false detection and missed detection rates and improve the robustness and practicality of the lock inspection behavior recognition method.

[0117] Reference Figure 7 , the embodiment of the present application further discloses a device for identifying box lock inspection behavior, the device comprising:

[0118] The first detection module 702 is configured to perform box key point detection on a video image of a box lock inspection scene to obtain position information of the box key points, and to perform hand detection on the video image to obtain position information of a hand detection frame, wherein the box key points include lock points;

[0119] A matching module 704 is configured to obtain a matching result between the hand detection frame and the locking point based on the position information of the locking point and the position information of the hand detection frame;

[0120] a second detection module 706 configured to, in response to the matching result indicating the successful matching of the hand detection frame and the lock point, perform hand key point detection on a target image area of ​​the video image to obtain position information of the hand key points, wherein the target image area is determined based on the position of the successfully matched hand detection frame;

[0121] A first behavior recognition module 708 is configured to perform feature analysis on the hand key points based on the position information of the hand key points to obtain an analysis result indicating whether there is an inspection behavior for a target lock buckle in the video image, where the target lock buckle is the lock buckle corresponding to the lock buckle point that has been successfully matched;

[0122] A human posture key point acquisition module 710 is configured to, in response to the analysis result indicating that there is an inspection behavior for a target lock in the video image, perform human posture key point recognition on the video image to acquire human posture key points;

[0123] The second behavior recognition module 712 is used to obtain the recognition result of the box lock inspection behavior in the video image based on the human body posture key points, the target lock points and the hand key points.

[0124] Optionally, obtaining a recognition result of the box lock inspection behavior in the video image based on the human body posture key points, the target lock points, and the hand key points includes:

[0125] Acquire an identity recognition result of the video image based on the human body posture key points;

[0126] Obtaining left and right hand recognition results based on the distance between the wrist key point in the human body posture key points and the wrist key point in the hand key points;

[0127] Based on the identity recognition result, the left and right hand recognition results and the target locking point, the recognition result of the box lock inspection behavior in the video image is obtained.

[0128] Optionally, the device further includes:

[0129] an early warning detection module (not shown in the figure), configured to, in response to a recognition result of a box lock check behavior in the video image within a preset time period meeting a preset violation condition, perform an alarm process based on the recognition result, wherein the preset violation condition includes: a recognition result of an abnormal lock check behavior in the video image;

[0130] The early warning detection module is used to obtain an identification result that the box lock buckle inspection behavior of the box lock buckle inspection scenario is normal in response to obtaining that the box lock buckle inspection behavior of the box lock buckle inspection scenario meets the preset compliance conditions, wherein the preset compliance conditions include: there is an inspection behavior of the target person inspecting two locks of the same box, and the number of the target persons is greater than or equal to a preset number threshold.

[0131] Optionally, the performing feature analysis on the hand key points based on the position information of the hand key points to obtain an analysis result indicating whether there is an inspection behavior for the target lock in the video image includes:

[0132] Acquire hand key point features based on the position information of the hand key points;

[0133] Matching the hand key point features with pre-acquired standard hand motion feature templates of box lock inspection behavior to obtain feature matching results;

[0134] In response to the feature matching result indicating a successful match, obtaining an analysis result indicating that an inspection behavior for a target buckle exists in the video image, wherein the analysis result includes: the hand detection frame corresponding to the hand key point and the target buckle;

[0135] In response to the feature matching result indicating a matching failure, an analysis result is obtained indicating that no lock check behavior exists in the video image.

[0136] Optionally, performing box key point detection on the video image of the box lock inspection scene to obtain position information of the box key points includes:

[0137] Perform box detection on the video image of the box lock inspection scene to obtain the position information of the box detection frame;

[0138] Based on the position information of the box detection frame, the box detection frame is enlarged to obtain an enlarged image area;

[0139] Perform box key point detection on the enlarged image area in the video image to obtain position information of the box key points.

[0140] Optionally, the box key points include: a first corner point, a second corner point, a third corner point, a fourth corner point, a first locking point, and a second locking point. After performing box key point detection on the enlarged image area in the video image and obtaining position information of the box key points, the method further includes:

[0141] Fitting the third corner point, the fourth corner point, the first locking point, and the second locking point to a straight line using a weighted least squares fitting method;

[0142] The position information of the first locking point and the second locking point is adjusted with the constraints that the distance between the third corner point and the first locking point is equal to the distance between the fourth corner point and the second locking point, and the locking position adjustment amount is minimized.

[0143] Optionally, obtaining a matching result between the hand detection frame and the locking point based on the position information of the locking point and the position information of the hand detection frame includes:

[0144] Calculating the relative distances from each of the lock points to each of the hand detection frames based on the position information of the lock points and the position information of the hand detection frames;

[0145] When the minimum value of the relative distance meets a preset distance condition, determining the hand detection frame and the lock point corresponding to the minimum value of the relative distance as the successfully matched hand detection frame and the lock point, and taking the successfully matched hand detection frame and the lock point as a matching result;

[0146] When the minimum relative distance does not meet the preset distance condition, a matching result indicating that the hand detection frame and the lock point fail to match is obtained.

[0147] The device for identifying the box lock inspection behavior disclosed in the embodiment of the present application is used to implement the method for identifying the box lock inspection behavior described in the embodiment of the present application. The specific implementation methods of each module of the device will not be repeated here, and reference can be made to the specific implementation methods of the corresponding steps in the method embodiment.

[0148] The embodiment of the present application discloses a device for identifying the lock inspection behavior of a box body. The device performs box key point detection on a video image of a box lock inspection scene to obtain the position information of the box key points, and performs hand detection on the video image to obtain the position information of the hand detection frame, wherein the box key points include: lock points; based on the position information of the lock points and the position information of the hand detection frame, a matching result of the hand detection frame and the lock points is obtained; in response to the matching result indicating that the hand detection frame and the lock point are successfully matched, hand key point detection is performed on the target image area of ​​the video image to obtain the position information of the hand key points, wherein The target image area is determined based on the position of the successfully matched hand detection frame; based on the position information of the hand key points, the hand key points are feature analyzed to obtain an analysis result indicating whether an inspection behavior for a target lock buckle exists in the video image, where the target lock buckle is the lock buckle corresponding to the successfully matched lock buckle point; in response to the analysis result indicating the presence of an inspection behavior for the target lock buckle in the video image, human posture key points are identified on the video image to obtain human posture key points; and based on the human posture key points, the target lock buckle points, and the hand key points, an identification result of the box lock buckle inspection behavior in the video image is obtained. This device combines image processing and target detection technologies to automatically identify whether a video image contains a box lock buckle inspection behavior, thereby improving the inspection efficiency of the box lock buckle inspection behavior and enhancing the accuracy of the inspection, with higher real-time performance. This effectively avoids the problem that manual inspection methods are prone to missed inspections or misjudgments due to operator negligence or subjective factors, which affects the accuracy of inspection results and the security of vault management.

[0149] This device integrates advanced target detection and key point detection technologies to accurately detect the position of a person's hand and the lock area, which can significantly reduce the false detection and missed detection rates and improve the robustness and practicality of the lock inspection behavior recognition method.

[0150] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between the various embodiments can be referred to in conjunction with each other. For the device embodiments, since they are generally similar to the method embodiments, their description is relatively simple, and for relevant parts, reference can be made to the description of the method embodiments.

[0151] The above is a detailed introduction to the method and device for identifying the box lock inspection behavior provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0153] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It will be appreciated by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the electronic device according to the embodiment of the present application. The application can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for performing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0154] For example, Figure 8An electronic device that can implement the method according to the present application is shown. The electronic device can be a PC, a mobile terminal, a personal digital assistant, a tablet computer, etc. The electronic device conventionally includes a processor 810 and a memory 820, and a program code 830 stored on the memory 820 and executable on the processor 810. When the processor 810 executes the program code 830, the method described in the above embodiments is implemented. The memory 820 can be a computer program product or a computer-readable medium. The memory 820 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. The memory 820 has a storage space 8201 for program code 830 of a computer program for executing any of the method steps described above. For example, the storage space 8201 for program code 830 can include individual computer programs for implementing various steps in the above method. The program code 830 is computer-readable code. These computer programs can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The computer program includes a computer-readable code, and when the computer-readable code is run on an electronic device, the electronic device is caused to execute the method according to the above embodiment.

[0155] The embodiment of the present application further discloses a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for identifying the box lock inspection behavior as described in the embodiment of the present application are implemented.

[0156] Such a computer program product may be a computer-readable storage medium having a computer program product. Figure 8 The memory 820 in the electronic device shown is similarly arranged as a storage segment, storage space, etc. The program code can be compressed and stored in the computer readable storage medium in an appropriate form. The computer readable storage medium is generally as shown in FIG. Figure 9 The portable or fixed storage unit generally includes computer-readable code 830', which is a code read by a processor and implements the steps of the above-described method when executed by the processor.

[0157] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0158] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0159] In the claims, any reference signs placed between brackets shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for identifying box lock inspection behavior, characterized in that: The method comprises: Performing box key point detection on a video image of a box lock inspection scene to obtain position information of the box key points, and performing hand detection on the video image to obtain position information of a hand detection frame, wherein the box key points include: lock points; Based on the position information of the locking point and the position information of the hand detection frame, obtaining a matching result between the hand detection frame and the locking point; In response to the matching result indicating the successful matching of the hand detection frame and the lock point, performing hand key point detection on a target image area of ​​the video image to obtain position information of the hand key points, wherein the target image area is determined according to the position of the successfully matched hand detection frame; Based on the position information of the hand key points, feature analysis is performed on the hand key points to obtain an analysis result indicating whether there is an inspection behavior for a target lock buckle in the video image, the target lock buckle being: the lock buckle corresponding to the lock buckle point that has been successfully matched; In response to the analysis result indicating that there is an inspection behavior for the target lock in the video image, performing human posture key point recognition on the video image to obtain the human posture key points; Based on the human body posture key points, the target locking points and the hand key points, a recognition result of the box locking inspection behavior in the video image is obtained.

2. The method according to claim 1, characterized in that The obtaining of a recognition result of the box lock inspection behavior in the video image based on the human body posture key points, the target lock points, and the hand key points includes: Acquire an identity recognition result of the video image based on the human body posture key points; Obtaining left and right hand recognition results based on the distance between the wrist key point in the human body posture key points and the wrist key point in the hand key points; Based on the identity recognition result, the left and right hand recognition results and the target locking point, the recognition result of the box lock inspection behavior in the video image is obtained.

3. The method according to claim 1, characterized in that After obtaining the recognition result of the box lock inspection behavior in the video image based on the human body posture key points, the target lock point and the hand key points, the system further includes one or more of the following early warning detection methods: In response to a recognition result of a box lock check behavior in the video image within a preset time period meeting a preset violation condition, an alarm process is performed based on the recognition result, wherein the preset violation condition includes: a recognition result of an abnormal lock check behavior in the video image; In response to obtaining that the box lock inspection behavior of the box lock inspection scenario meets the preset compliance conditions, an identification result that the box lock inspection behavior of the box lock inspection scenario is normal is obtained, wherein the preset compliance conditions include: there is an inspection behavior of the target person inspecting two locks of the same box, and the number of the target persons is greater than or equal to the preset number threshold.

4. The method according to claim 1, wherein The performing feature analysis on the hand key points based on the position information of the hand key points to obtain an analysis result indicating whether there is an inspection behavior for the target lock in the video image, including: Acquire hand key point features based on the position information of the hand key points; Matching the hand key point features with pre-acquired standard hand motion feature templates of box lock inspection behavior to obtain feature matching results; In response to the feature matching result indicating a successful match, obtaining an analysis result indicating that an inspection behavior for a target buckle exists in the video image, wherein the analysis result includes: the hand detection frame corresponding to the hand key point and the target buckle; In response to the feature matching result indicating a matching failure, an analysis result is obtained indicating that no lock check behavior exists in the video image.

5. The method according to claim 1, wherein The method of detecting key points of the box body on the video image of the box body lock inspection scene to obtain position information of the key points of the box body includes: Perform box detection on the video image of the box lock inspection scene to obtain the position information of the box detection frame; Based on the position information of the box detection frame, the box detection frame is enlarged to obtain an enlarged image area; Perform box key point detection on the enlarged image area in the video image to obtain position information of the box key points.

6. The method according to claim 5, characterized in that The box key points include: a first corner point, a second corner point, a third corner point, a fourth corner point, a first locking point, and a second locking point. After performing box key point detection on the enlarged image area in the video image and obtaining position information of the box key points, the method further includes: Fitting the third corner point, the fourth corner point, the first locking point, and the second locking point to a straight line using a weighted least squares fitting method; The position information of the first locking point and the second locking point is adjusted with the constraints that the distance between the third corner point and the first locking point is equal to the distance between the fourth corner point and the second locking point, and the locking position adjustment amount is minimized.

7. The method according to claim 1, characterized in that The obtaining, based on the position information of the locking point and the position information of the hand detection frame, a matching result between the hand detection frame and the locking point includes: Calculating the relative distances from each of the lock points to each of the hand detection frames based on the position information of the lock points and the position information of the hand detection frames; When the minimum value of the relative distance meets a preset distance condition, determining the hand detection frame and the lock point corresponding to the minimum value of the relative distance as the successfully matched hand detection frame and the lock point, and taking the successfully matched hand detection frame and the lock point as a matching result; When the minimum relative distance does not meet the preset distance condition, a matching result indicating that the hand detection frame and the lock point fail to match is obtained.

8. A device for identifying box lock inspection behavior, characterized in that: The device comprises: The first detection module is configured to perform box key point detection on a video image of a box lock inspection scene to obtain position information of the box key points, and to perform hand detection on the video image to obtain position information of a hand detection frame, wherein the box key points include lock points; a matching module, configured to obtain a matching result between the hand detection frame and the locking point based on the position information of the locking point and the position information of the hand detection frame; a second detection module, configured to, in response to the matching result indicating the successful matching of the hand detection frame and the lock point, perform hand key point detection on a target image area of ​​the video image to obtain position information of the hand key points, wherein the target image area is determined based on the position of the successfully matched hand detection frame; a first behavior recognition module, configured to perform feature analysis on the hand key points based on the position information of the hand key points, and obtain an analysis result indicating whether there is an inspection behavior for a target lock buckle in the video image, wherein the target lock buckle is the lock buckle corresponding to the lock buckle point that has been successfully matched; a human posture key point acquisition module, configured to, in response to the analysis result indicating that there is an inspection behavior for the target lock in the video image, perform human posture key point recognition on the video image to acquire human posture key points; The second behavior recognition module is used to obtain the recognition result of the box lock inspection behavior in the video image based on the human body posture key points, the target lock points and the hand key points.

9. An electronic device comprising a memory, a processor, and a program code stored in the memory and executable on the processor, wherein: When the processor executes the program code, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having program code stored thereon, characterized in that: When the program code is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.