Climbing behavior identification method, electronic equipment and system

By analyzing the key points of the human body in the video image and judging climbing behavior, the problem of difficulty in identifying special climbing postures in the existing technology is solved, and early recognition and blocking of climbing behavior is achieved.

CN119992650APending Publication Date: 2025-05-13HANGZHOU HIKVISION SYST TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064674.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to detect and identify special climbing postures, and managers may not have time to stop them when people climb the fence, affecting subway operation.

Method used

Determine whether to perform climbing behavior by determining the key points of the first human body in a frame of the target image of the video file, such as the elbow, shoulder and hip angle, shoulder, hip and knee angle and knee height difference.

Benefits of technology

Accurately identify various climbing actions before personnel cross the fence, promptly remind managers to prevent climbing behavior and avoid personnel entering the management area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992650A_ABST
    Figure CN119992650A_ABST
Patent Text Reader

Abstract

The invention discloses a climbing behavior identification method, electronic equipment and a system, relates to the technical field of image processing, and is used for accurately detecting various fence climbing behaviors before a human body crosses a fence. The method comprises the steps that based on human body key points of a first human body in a frame of target image of a video file, action features of the first human body are determined, and the action features comprise one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle and the knee height difference; and identifying whether the first human body in the target image performs a climbing behavior or not based on the action features of the first human body.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202411216177.5, and the application date of the original application is August 30, 2024. The entire content of the original application can be incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of image processing technology, and in particular to a climbing behavior recognition method, electronic equipment and system. Background Art

[0003] For some management areas, managers do not want people to enter the management area illegally. For example, the track area of ​​the subway is a management area. Passengers climbing the subway screen door to enter the track area is a violation of the rules, which will cause great inconvenience to the work of subway staff and affect the operation of the subway. Therefore, it is necessary to detect the fences in the management area to identify whether there are people climbing the fences.

[0004] Currently, whether a person climbs over a fence to enter a management area is often determined by judging whether the line connecting the person's two feet intersects with the upper edge of the fence. However, when the person's posture for climbing over the fence is special and there is no state where the line connecting the person's two feet intersects with the upper edge of the fence, this method cannot detect the person's special climbing posture. Moreover, when a person climbing over a fence is detected by this method, the line connecting the person's two feet intersects with the upper edge of the fence, and part of the person's body has entered the management area. At this time, the detected climbing behavior is reported to the management personnel, and the management personnel may not have time to stop the person climbing the fence. Summary of the invention

[0005] The present application provides a climbing behavior recognition method, electronic device and system for accurately detecting various fence climbing behaviors before a human body crosses the fence.

[0006] In order to achieve the above technical objectives, this application adopts the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a climbing behavior recognition method, the method comprising:

[0008] Based on the human body key points of the first human body in a frame of target image of the video file, the motion features of the first human body are determined, and the motion features include one or more of elbow-shoulder-hip angle, shoulder-hip-knee angle, and knee height difference; the elbow-shoulder-hip angle is an angle formed by the shoulder key point on the same side of the first human body as a vertex, and the line segment formed by the shoulder key point and the elbow key point, and the line segment formed by the shoulder key point and the hip key point as edges; the shoulder-hip-knee angle is an angle formed by the hip key point on the same side of the first human body as a vertex, and the line segment formed by the hip key point and the shoulder key point, and the line segment formed by the hip key point and the knee key point as edges; the knee height difference is the pixel height difference between the two knee key points of the first human body;

[0009] It is identified based on the action feature of the first human body whether the first human body in the target image performs a climbing behavior.

[0010] The technical solution provided by the present application brings at least the following beneficial effects: firstly, the key points of the first human body in the target image are determined, and then the action characteristics of the human body are determined based on the key points of the human body, and one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference in the action characteristics of the first human body are used to determine whether the first human body performs climbing behavior, and it is not necessary to determine whether the human body crosses the fence to perform climbing behavior by the intersection of the two feet of the human body with the fence. In other words, when the related art determines that the human body performs climbing behavior by the intersection of the two feet of the human body with the fence, the human body has already crossed the fence, and the management personnel may not have time to stop it, and cannot identify the special climbing posture, while the solution of the present application does not need to determine whether the human body crosses the fence to perform climbing behavior by the intersection of the two feet of the human body with the fence, and various climbing actions of the first human body can be identified before the first human body crosses the fence, so that the management personnel can be notified in time, which buys more time for the management personnel to prevent the first human body from crossing the fence, and the first human body can be prevented from crossing the fence in time to prevent the first human body from entering the management area.

[0011] In a possible implementation, identifying whether a first person in a target image performs a climbing behavior based on the motion characteristics of the first person includes: when the motion characteristics include elbow, shoulder and hip angles, and at least one elbow, shoulder and hip angle of the first person is greater than a first angle threshold, determining that the first person performs a climbing behavior; when the motion characteristics include shoulder, hip and knee angles, and at least one shoulder, hip and knee angle of the first person is less than a second angle threshold, determining that the first person performs a climbing behavior; when the motion characteristics include knee height difference, and the knee height difference of the first person is greater than a target pixel height, determining that the first person performs a climbing behavior; wherein the target pixel height is the product of the pixel distance between the shoulder key point and the hip key point among the body key points of the first person and a first preset coefficient.

[0012] In a possible implementation, the human body in the target image that meets the screening conditions is the first human body; wherein the screening conditions include at least one of the following: the pixel distance between at least one hand key point among the human body key points and the upper edge of the fence is less than the first pixel distance; the first pixel distance is used to represent the pixel distance threshold where the human body hand is connected to the upper edge of the fence; the pixel distance between at least one foot key point among the human body key points and the lower edge of the fence is less than the second pixel distance; the second pixel distance is used to represent the pixel distance threshold where the human body foot is connected to the lower edge of the fence.

[0013] In one possible implementation, the confidence of at least one hand key point is greater than a first preset confidence; the confidence of at least one foot key point is greater than a second preset confidence; the first pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points and the second preset coefficient; the second pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points and the third preset coefficient.

[0014] In a possible implementation, before determining the motion characteristics of the first human body based on the human body key points of the first human body, the method also includes: obtaining a climbing behavior score of the first human body based on the human body key points of the first human body, the climbing behavior score being used to characterize the possibility of the first human body performing a climbing behavior; determining the motion characteristics of the first human body based on the human body key points of the first human body, including: determining the motion characteristics of the first human body based on the human body key points of the first human body when the climbing behavior score of the first human body is within a preset range.

[0015] In a possible implementation, the method further includes: in a congestion detection mode, if there is any first human body performing climbing behavior in the target image, recording the target image as a first state; in a non-congestion detection mode, if the first human body in the target image performs climbing behavior, recording the first human body as a second state; in a congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, outputting a first alarm message, the first alarm message being used to prompt that there is a human body performing climbing behavior; or, in a congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, and the first alarm message is not transmitted for the last time In the non-congestion detection mode, if the time interval between the second alarm information and the target image is greater than the first time interval, the first alarm information is output; in the non-congestion detection mode, if the same first human body exists in the target image and a second preset number of frame images that are continuous with the target image before the target image, and the first human body is in the second state, the second alarm information is output, and the second alarm information is used to prompt the first human body to perform a climbing behavior; or, in the non-congestion detection mode, if the same first human body exists in the target image and a second preset number of frame images that are continuous with the target image before the target image, and the first human body is in the second state, and the time interval between the second alarm information indicating the first human body issued last time is greater than the second time interval, the second alarm information is output.

[0016] In a possible implementation, the method further includes: when the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a third preset number of frame images that are continuous with the target image before the target image is less than or equal to the first value, and there is no human body performing climbing behavior in the target image, switching the congestion detection mode to the non-congestion detection mode; or, when the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a fourth preset number of frame images that are continuous with the target image before the target image is less than or equal to the second value, and there is no human body performing climbing behavior in the recognition area, switching the congestion detection mode to the non-congestion detection mode. The control unit switches the detection mode from the non-congestion detection mode to the non-congestion detection mode; in the case where the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the fifth preset number of frame images that are continuous with the target image before the target image is greater than the first value, and there is no human body in a climbing state in the target image, the non-congestion detection mode is switched to the congestion detection mode; or, in the case where the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the sixth preset number of frame images that are continuous with the target image before the target image is greater than the second value, and there is no human body in a climbing state in the target image, the non-congestion detection mode is switched to the congestion detection mode.

[0017] In a possible implementation, the method further includes: in a congestion detection mode, if a target key point among the human key points of any human body in the target image falls into the management area, the target image is recorded as the third state; in a non-congestion detection mode, if a target key point among the human key points of a second human body in the target image falls into the management area, the second human body is recorded as the fourth state, and the second human body is any human body in the target image; in a congestion detection mode, if the target image and the seventh preset number of frame images continuous with the target image before the target image are all in the third state, and the first alarm information has been output within the first preset time length before the target image, the third alarm information is output, and the third alarm information is used to prompt that there is a human body in the target image that illegally enters the management area; or, in a congestion detection mode, if the target image and the seventh preset number of frame images continuous with the target image before the target image are all in the third state, and the first alarm information has been output within the first preset time length before the target image If a first alarm message has been issued and the time interval with the last issuance of the third alarm message is greater than the third time interval, a third alarm message is output; in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image and the second human body is in the fourth state, and the second alarm message indicating the second human body has been output within the second preset time length before the target image, a fourth alarm message is output, and the fourth alarm message is used to prompt that the second human body has illegally entered the management area; or, in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image and the second human body is in the fourth state, the second alarm message indicating the second human body has been output within the second preset time length before the target image, and the time interval with the last issuance of the fourth alarm message indicating the second human body is greater than the fourth time interval, a fourth alarm message is output.

[0018] In a second aspect, the present application provides a climbing behavior recognition device, comprising:

[0019] The processing module is used to determine the motion features of the first human body based on the human body key points of the first human body in a frame of target image of the video file, and the motion features include one or more of elbow-shoulder-hip angle, shoulder-hip-knee angle, and knee height difference; the elbow-shoulder-hip angle is an angle formed by the shoulder key point on the same side of the first human body as the vertex, and the line segment formed by the shoulder key point and the elbow key point, and the line segment formed by the shoulder key point and the hip key point as the sides; the shoulder-hip-knee angle is an angle formed by the hip key point on the same side of the first human body as the vertex, and the line segment formed by the hip key point and the shoulder key point, and the line segment formed by the hip key point and the knee key point as the sides; the knee height difference is the pixel height difference between the two knee key points of the first human body;

[0020] The processing module is further used to identify whether the first person in the target image performs a climbing behavior based on the action characteristics of the first person.

[0021] In one possible implementation, the processing module is specifically used to: determine that the first human body has performed a climbing behavior when the action characteristics include elbow, shoulder and hip angles, and at least one elbow, shoulder and hip angle of the first human body is greater than a first angle threshold; determine that the first human body has performed a climbing behavior when the action characteristics include shoulder, hip and knee angles, and at least one shoulder, hip and knee angle of the first human body is less than a second angle threshold; determine that the first human body has performed a climbing behavior when the action characteristics include knee height difference, and the knee height difference of the first human body is greater than a target pixel height; wherein the target pixel height is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points of the first human body and a first preset coefficient.

[0022] In a possible implementation, the human body in the target image that meets the screening conditions is the first human body; wherein the screening conditions include at least one of the following: the pixel distance between at least one hand key point among the human body key points and the upper edge of the fence is less than the first pixel distance; the first pixel distance is used to represent the pixel distance threshold where the human body hand is connected to the upper edge of the fence; the pixel distance between at least one foot key point among the human body key points and the lower edge of the fence is less than the second pixel distance; the second pixel distance is used to represent the pixel distance threshold where the human body foot is connected to the lower edge of the fence.

[0023] In one possible implementation, the confidence of at least one hand key point is greater than a first preset confidence; the confidence of at least one foot key point is greater than a second preset confidence; the first pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points and the second preset coefficient; the second pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points and the third preset coefficient.

[0024] In a possible implementation, the processing module is further used to obtain a climbing behavior score of the first human body based on the human body key points of the first human body, and the climbing behavior score is used to characterize the possibility of the first human body performing a climbing behavior; the processing module is specifically used to: when the climbing behavior score of the first human body is within a preset range, determine the action characteristics of the first human body based on the human body key points of the first human body.

[0025] In a possible implementation, the processing module is further used for: in a congestion detection mode, if there is any first human body in the target image performing a climbing behavior, recording the target image as a first state; in a non-congestion detection mode, if the first human body in the target image performs a climbing behavior, recording the first human body as a second state; the climbing behavior recognition device also includes a transmission module, the transmission module is used for: in a congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, outputting a first alarm message, the first alarm message is used to prompt that there is a human body performing a climbing behavior; or, in a congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, and are consistent with the previous state. When the time interval between the first alarm information being sent out for the first time is greater than the first time interval, the first alarm information is output; the transmission module is also used for: in a non-congestion detection mode, if the same first human body exists in the target image and a second preset number of frame images that are continuous with the target image before the target image, and the first human body is in the second state, the second alarm information is output, and the second alarm information is used to prompt the first human body to perform a climbing behavior; or, in a non-congestion detection mode, if the same first human body exists in the target image and a second preset number of frame images that are continuous with the target image before the target image, and the first human body is in the second state, and the time interval between the last second alarm information indicating the first human body being sent out is greater than the second time interval, the second alarm information is output.

[0026] In one possible implementation, the processing module is further used for: when the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a third preset number of frame images that are continuous with the target image before the target image is less than or equal to the first value, and there is no human body performing climbing behavior in the target image, switching the congestion detection mode to the non-congestion detection mode; or, when the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a fourth preset number of frame images that are continuous with the target image before the target image is less than or equal to the second value, and there is no human body performing climbing behavior in the identification area, switching the congestion detection mode. to the non-congestion detection mode; the processing module is also used for: when the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the fifth preset number of frame images that are continuous with the target image before the target image is greater than the first value, and there is no human body in a climbing state in the target image, switching the non-congestion detection mode to the congestion detection mode; or, when the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the sixth preset number of frame images that are continuous with the target image before the target image is greater than the second value, and there is no human body in a climbing state in the target image, switching the non-congestion detection mode to the congestion detection mode.

[0027] In a possible implementation, the processing module is also used for: in a congestion detection mode, if a target key point among the human key points of any human body in the target image falls into the management area, recording the target image as the third state; in a non-congestion detection mode, if a target key point among the human key points of a second human body in the target image falls into the management area, recording the second human body as the fourth state, and the second human body is any human body in the target image; the transmission module is also used for: in a congestion detection mode, if the target image and the seventh preset number of frame images continuous with the target image before the target image are all in the third state, and the first alarm information has been output within the first preset time length before the target image, outputting a third alarm information, and the third alarm information is used to prompt that there is a human body in the target image that illegally enters the management area; or, in a congestion detection mode, if the target image and the seventh preset number of frame images continuous with the target image before the target image are all in the third state, and the first alarm information has been output within the first preset time length before the target image The transmission module is also used for: in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image, and the second human body is in the fourth state, and the second alarm information indicating the second human body has been output within the second preset time length before the target image, output the fourth alarm information, and the fourth alarm information is used to prompt that the second human body has illegally entered the management area; or, in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image, and the second human body is in the fourth state, the second alarm information indicating the second human body has been output within the second preset time length before the target image, and the time interval from the last issuance of the fourth alarm information indicating the second human body is greater than the fourth time interval, output the fourth alarm information.

[0028] In a third aspect, the present application provides a climbing behavior recognition system, comprising:

[0029] An image acquisition device, used for acquiring video files;

[0030] An electronic device, used to execute any one of the climbing behavior identification methods provided in the first aspect above;

[0031] A prompting device is used to receive and display a first alarm message, a second alarm message, a third alarm message and a fourth alarm message sent by an electronic device, the first alarm message is used to indicate that a human body performs climbing behavior in a frame of a target image in a video file, the second alarm message is used to indicate that the first human body performs climbing behavior, the third alarm message is used to indicate that a human body illegally enters a management area in a frame of a target image in a video file, and the fourth alarm message is used to indicate that a second human body illegally enters a management area.

[0032] In a fourth aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any one of the climbing behavior identification methods provided in the first aspect.

[0033] In a fifth aspect, the present application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement any one of the climbing behavior identification methods provided in the first aspect above.

[0034] In a sixth aspect, the present application provides a computer program product, comprising computer instructions, which, when executed by a processor, implement any one of the climbing behavior identification methods provided in the first aspect above.

[0035] For the specific description of the second to sixth aspects and their various implementations in the present application, reference may be made to the detailed description in the first aspect and its various implementations; and for the beneficial effects of the second to sixth aspects and their various implementations, reference may be made to the beneficial effects analysis in the first aspect and its various implementations, which will not be repeated here.

[0036] These and other aspects of the present application will become more apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of a climbing scene provided in an embodiment of the present application;

[0038] Figure 2 A schematic diagram of the structure of a climbing behavior recognition system provided in an embodiment of the present application;

[0039] Figure 3 A schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application;

[0040] Figure 4 A flowchart of a climbing behavior identification method provided in an embodiment of the present application;

[0041] Figure 5 A schematic diagram of human body key point recognition provided in an embodiment of the present application;

[0042] Figure 6 A schematic diagram of a network structure of a human key point detection model provided in an embodiment of the present application;

[0043] Figure 7 A schematic diagram of the detection results of a human key point detection model provided in an embodiment of the present application;

[0044] Figure 8Schematic diagram of an application scenario of a climbing behavior recognition method provided in an embodiment of the present application Figure 1 ;

[0045] Fig. 9 Schematic diagram of an application scenario of a climbing behavior recognition method provided in an embodiment of the present application Figure 2 ;

[0046] Fig.10 Schematic diagram of an application scenario of a climbing behavior recognition method provided in an embodiment of the present application Figure 3 ;

[0047] Fig.11 Schematic diagram of an application scenario of a climbing behavior recognition method provided in an embodiment of the present application Figure 4 ;

[0048] Fig.12 Schematic diagram of an application scenario of a climbing behavior recognition method provided in an embodiment of the present application Figure 5 ;

[0049] Fig.13 A logical diagram of a climbing behavior recognition process and a detection process of a person entering a management area provided in an embodiment of the present application;

[0050] Fig.14 A logical schematic diagram of a climbing behavior identification event reporting process and a detection event reporting process of a person entering a management area provided in an embodiment of the present application;

[0051] Fig.15 A schematic diagram of the structure of a climbing behavior recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0053] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way. The terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, "multiple" means two or more.

[0054] For some management areas, managers do not want people to enter the management area illegally. For example, the track area of ​​the subway is a management area. Passengers climbing the subway screen door to enter the track area is a violation of the rules, which will cause great inconvenience to the work of subway staff and affect the operation of the subway. Therefore, it is necessary to detect the fences in the management area to identify whether there are people climbing the fences.

[0055] Currently, the method of determining whether a person has climbed a fence is to determine whether the line connecting the person's two feet intersects with the upper edge of the fence. However, this method cannot detect special climbing behaviors, such as jumping over the fence, climbing over the fence, or back-jumping over the fence. Figure 1 In the climbing behavior shown, the line connecting the two feet of person 1 does not intersect with the upper edge of the fence, so person 1 will not be judged as climbing based on this method. Moreover, when person 1 is detected climbing the fence in this way, the line connecting the two feet of person 1 intersects with the upper edge of the fence, and part of the body of person 1 has entered the management area. At this time, the detected climbing behavior is reported to the management personnel, and the management personnel may not have time to stop the person climbing the fence.

[0056] In this regard, an embodiment of the present application provides a climbing behavior recognition method, which first determines the climbing behavior score of the first human body in the target image. For the first human body whose climbing behavior score is within a preset interval, one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference of the first human body are used to determine whether the first human body has performed a climbing behavior. It is not necessary to determine whether the human body has crossed the fence to perform a climbing behavior by the intersection of the two feet of the human body with the fence. In other words, when the related art determines whether the human body has performed a climbing behavior by the intersection of the two feet of the human body with the fence, the human body has already crossed the fence, and the management personnel may not be able to stop it in time, and the special climbing posture cannot be identified. However, the solution of the present application does not need to determine whether the human body has crossed the fence to perform a climbing behavior by the intersection of the two feet of the human body with the fence. The various climbing actions of the first human body can be identified before the first human body crosses the fence, so the management personnel can be notified in time, which buys more time for the management personnel to stop the first human body from crossing the fence, and the first human body can be stopped from crossing the fence in time to prevent the first human body from entering the management area.

[0057] Please refer to Figure 2 , which shows a schematic diagram of the structure of a climbing behavior identification system to which the climbing behavior identification method provided in this application is applicable. Figure 2 As shown, the climbing behavior recognition system 1 may include: an image acquisition device 10 , an electronic device 20 and a prompt device 30 .

[0058] The electronic device 20 establishes communication connections with the image acquisition device 10 and the prompt device 30 respectively. It should be understood that the connection method can be a wireless connection, such as a Bluetooth connection, a wireless fidelity (Wi-Fi) connection, etc.; or, the connection method can also be a wired connection, such as an optical fiber connection, etc., which is not limited. Exemplarily, the image acquisition device 10, the electronic device 20 or the prompt device 30 can be connected to the Internet through a router, thereby realizing the communication connection between the electronic device 20 and the image acquisition device 10 and the prompt device 30.

[0059] In some embodiments, the image acquisition device 10 is used to output a video file to the electronic device 20. For example, image acquisition devices are set up in the management area, and these image acquisition devices can capture the fence of the management area. When it is necessary to detect whether there is climbing behavior at the fence of the management area, the image acquisition device 10 can send the captured video file to the electronic device 20, and then the electronic device 20 detects the image in the video file to determine whether there is climbing behavior.

[0060] In some embodiments, the electronic device 20 is used to recognize images in a video file, determine a climbing behavior score of a first person in the image, and for the first person whose climbing behavior score is within a preset range, determine whether the first person has performed a climbing behavior through one or more of the first person's elbow-shoulder-hip angle, shoulder-hip-knee angle, and knee height difference.

[0061] In some embodiments, the electronic device 20 may include a processor. The processor is used to identify the climbing behavior score of the first human body in the image of the video file, and for the first human body whose climbing behavior score is within a preset interval, one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference of the first human body are used to determine whether the first human body performs climbing behavior. Among them, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose processor network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor can also be other devices with processing functions, such as circuits, devices, or software modules, and the embodiments of the present application do not impose any restrictions on this.

[0062] Optionally, the electronic device may further include a memory for storing images in a video file from the image acquisition device 10, and then the processor may search the memory for images in the video file to identify climbing behavior based on the images in the video file.

[0063] In some embodiments, the prompt device 30 is used to receive and display the first alarm information, the second alarm information, the third alarm information and the fourth alarm information sent by the electronic device 20, the first alarm information is used to indicate that there is a human body performing climbing behavior in a frame of the target image of the video file, the second alarm information is used to indicate that the first human body performs climbing behavior, the third alarm information is used to indicate that there is a human body illegally entering the management area in a frame of the target image of the video file, and the fourth alarm information is used to indicate that the second human body illegally enters the management area. Exemplarily, the prompt device 30 can be a voice prompt device used by the inspection personnel in the management area, and the prompt device 30 presents the alarm information to the inspection personnel by playing the audio indicating the alarm information. Alternatively, the prompt device 30 can also be a display device, and the prompt device 30 presents the alarm information to the user by displaying text or images indicating the alarm information on the display screen.

[0064] In some embodiments, the image acquisition device 10 and the electronic device 20 may be as follows: Figure 2 As shown, they are two independent devices, or the image acquisition device 10 and the electronic device 20 can be integrated into the same device.

[0065] In some embodiments, the electronic device 20 and the prompting device 30 may be as follows: Figure 2 As shown, they are two independent devices, or the electronic device 20 and the prompting device 30 can be integrated into the same device.

[0066] In some embodiments, the climbing behavior recognition system 1 may include one or more image acquisition devices 10 .

[0067] In some embodiments, the image acquisition device 10 can be any device that can transmit video files to the electronic device 20, such as a video camera, a still camera, a dome camera, etc. The embodiment of the present application does not limit the specific form of the image acquisition device 10.

[0068] In some embodiments, the electronic device 20 may be a single server or a server cluster, or the electronic device 20 may be a terminal device, such as a personal computer (PC), a notebook computer, a mobile device, a tablet computer, a laptop computer, etc. The specific form of the electronic device 20 is not limited in the embodiments of the present application.

[0069] The hardware structure of the electronic device 20 includes: Figure 3 The computing device shown in FIG. Figure 3 Taking the computing device shown as an example, the hardware structure of the electronic device 20 is introduced.

[0070] like Figure 3 As shown, the computing device may include a processor 201, a memory 202, a communication interface 203, and a bus 204. The processor 201, the memory 202, and the communication interface 203 may be connected via a bus 204.

[0071] The processor 201 is the control center of the computing device, and can be a processor or a general term for multiple processing elements. For example, the processor 201 can be a general-purpose central processing unit (CPU) or other general-purpose processors. Among them, the general-purpose processor can be a microprocessor or any conventional processor.

[0072] As an embodiment, the processor 201 may include one or more CPUs, such as Figure 3 CPU 0 and CPU 1 are shown in .

[0073] The memory 202 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0074] In a possible implementation, the memory 202 may exist independently of the processor 201, and the memory 202 may be connected to the processor 201 via a bus 204 to store instructions or program codes. When the processor 201 calls and executes the instructions or program codes stored in the memory 202, the model deployment method provided in the embodiment of the present application can be implemented.

[0075] In another possible implementation, the memory 202 may also be integrated with the processor 201 .

[0076] The communication interface 203 is used for connecting the computing device to other devices through a communication network, which may be Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 203 may include a receiving unit for receiving data and a sending unit for sending data.

[0077] The bus 204 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0078] It should be pointed out that Figure 3 The structure shown in the figure does not constitute a limitation on the computing device, except Figure 3In addition to the components shown, the computing device may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0079] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0080] The climbing behavior recognition method provided in the embodiment of the present application can be executed by the electronic device 20 in the climbing behavior recognition system 1.

[0081] like Figure 4 As shown, the embodiment of the present application provides a climbing behavior recognition method, which includes the following steps:

[0082] S101, identifying a frame of a target image in a video file to obtain key points of a first human body in the target image.

[0083] The target image is a frame of image that the electronic device is currently performing climbing behavior recognition on.

[0084] In some embodiments, the video file is a video stream captured in real time by an image acquisition device and transmitted to an electronic device. In this way, the electronic device can identify climbing behavior in real time, and after reporting the real-time climbing behavior, the detection personnel can promptly stop the first person from crossing the fence and prevent the first person from entering the management area.

[0085] Optionally, the video file may also be a video stream of a previous time period collected by an image acquisition device. In this way, the electronic device identifies the climbing behavior, and the management personnel may also trace the climbing event.

[0086] In some embodiments, the electronic device inputs the target image into the target model to obtain the recognition result of the key points of the human body in the target image, and the target model is used to recognize the key points of the human body.

[0087] Preferably, the target model may be a YOLO-Pose human key point recognition model.

[0088] The following is an example of the training process of the YOLO-Pose human key point recognition model:

[0089] Before detecting human key points, you can use the public COCO person keypoints dataset to train the YOLO-Pose human key point recognition model. The COCO person keypoints dataset includes a large number of annotated sample images, such as Figure 5As shown in Table 1, 17 key points are annotated on the human body in each sample image. For example, the left shoulder of the human body corresponds to key point 5, the left elbow of the human body corresponds to key point 7, the left hand of the human body corresponds to key point 9, the right shoulder of the human body corresponds to key point 6, the right elbow of the human body corresponds to key point 8, and the right hand of the human body corresponds to key point 10.

[0090] Table 1 Human body key point annotation table

[0091] Human body key point number Names of key points of the human body 0 nose 1 Left Eye 2 Right eye 3 left ear 4 Right ear 5 Left shoulder 6 Right shoulder 7 left elbow 8 right elbow 9 Left Hand 10 Right hand 11 Left hip 12 Right hip 13 Left knee 14 Right knee 15 Left foot 16 Right foot

[0092] Among them, for each human body, 57 elements are predicted, including: the position x, y and confidence conf of each of the 17 key points of the human body, a total of 51 elements, and the center point coordinate C of the prediction box (the smallest rectangle covering the human body) where the human body is located x and C y , width W, height H, detection box confidence box conf And the confidence class of the detection point category in the detection box conf There are 6 elements in total, so all the elements that need to be predicted can be expressed as a prediction vector, that is:

[0093]

[0094] in, In training, it represents the horizontal coordinate of the nth key point. In training, it represents the ordinate of the nth key point. Indicates the visible flag of the nth keypoint during training. If Indicates that this key point is not marked. In this case , which means that the key point is marked but not visible (occluded). , which means that this key point is marked and visible at the same time. In the reasoning stage, the confidence is retained Key points greater than 0, that is, the marked key points are retained, and the unmarked key points are discarded. Since 17 key points are output for each human key point detection during the model training process, it is necessary to filter out the key points outside the field of view, otherwise there will be suspended key points, causing the human skeleton to deform. The COCO person keypoints human key point dataset is divided into a training dataset and a test dataset in a ratio of 9:1. Finally, before model training, the COCO person keypoints human key point dataset is enhanced according to resolution and light factors, and the image contrast is enhanced using histogram equalization. The image resolution is appropriately adjusted according to the resolution requirements during actual detection.

[0095] In practical applications, you can use Figure 6 The network structure of the YOLO-Pose human key point detection model shown in the figure uses the CSP-darknet53 backbone feature extraction network as the backbone. The CSP-darknet53 backbone feature extraction network can extract rich information features from the input image. The CSP-darknet53 backbone feature extraction network solves the gradient information duplication problem of network optimization in other large convolutional neural network frameworks, and integrates the gradient changes from beginning to end into the feature map, thereby reducing the number of parameters and floating-point operations per second (FLOPS) of the YOLO-Pose human key point detection model, which not only ensures the inference speed and accuracy, but also reduces the model size. The Neck network part of the YOLO-Pose human key point detection model adopts the PANet structure. The Neck network is mainly used to generate feature pyramids. The feature pyramid will enhance the model's detection of objects of different scales, so that the model can recognize the same object of different sizes and scales. The PANet structure adds a bottom-up enhancement based on the FPN algorithm, so that the top feature map can also enjoy the rich location information brought by the bottom layer, thereby improving the detection effect of large objects. The PANet structure is used to fuse features of various scales from the trunk, followed by the Head structure for final detection, outputting four detection heads of different scales, as well as the predicted boxes and key points in each detection head.

[0096] The YOLO-Pose human key point detection model uses complete-intersection over union (CIOU) as the regression positioning loss function of the prediction box (bounding box, bbox), where the intersection over union (IOU) means that for any prediction box A and B, the intersection and union are calculated respectively, and the ratio of the two is calculated. The expression of IOU is: The regression positioning loss of the YOLO-Pose human key point detection model considers three geometric parameters: overlapping area, center point distance, and aspect ratio. CIOU adds a penalty term (loss) for the scale of the predicted box on the basis of distance-intersection over union (DIOU), so that the predicted box will be more consistent with the real box. The expression of CIOU is as follows:

[0097]

[0098] Where α is the weight function, And υ is used to measure the similarity of aspect ratio, defined as

[0099] Then the complete prediction box regression positioning loss function is defined as:

[0100] Target keypoint similarity (OKS) is a commonly used metric for evaluating keypoints. By defining the regressed keypoints directly as fixed centers, the concept of IOU loss can be extended from bbox to keypoints. In the presence of keypoints, OKS is treated as IOU, so the OKS loss is essentially scale-invariant and tilted for specific keypoints. In other words, keypoints on a person's head, such as ears, nose, and eyes, will receive more error penalties at the pixel level than keypoints on the body, such as shoulders, knees, and hips. Unlike the standard IOU loss, the gradient of IOU loss disappears when there is no overlap, while the OKS loss does not have a vanishing gradient. Therefore, the OKS loss is more similar to the DIOU loss.

[0101] Corresponding to each bbox, the entire posture information is stored. Therefore, OKS is calculated for each key point separately, and then added together to obtain the final OKS loss or key point IOU loss:

[0102]

[0103] Among them, d n is the Euclidean distance between the predicted nth key point and its true value, k n is the penalty weight factor of the key point itself, s is the scale of the target, δ(v n ) is the visible mark of each key point, which is also mentioned in the above content

[0104] Corresponding to each key point, a confidence parameter is learned, which shows whether there is a key point in the human body. The loss function is:

[0105]

[0106] in, is the confidence of the predicted nth key point.

[0107] Finally, the overall loss of the human key point prediction model is expressed as:

[0108]

[0109] Among them, λ cls =0.5,λ box=0.05,λ kpts =0.1,λ kpts_conf = 0.5 is a hyperparameter in the model, which is used to balance the losses between different scales.

[0110] In summary, the training process of the YOLO-Pose human key point detection model can be summarized as follows: Initialize the weights and parameters of the improved YOLO-Pose human key point detection model, including the convolutional layer parameter value, learning rate, number of iterations (epoch), and the number of data samples captured in one training (batch_size). Divide the COCO personkeypoints human key point data set into a training set and a test set and place them in an agreed directory, and run the program for training. Before training, the program will select 1 / 10 of the data from the training set as the validation set. Each iteration will be verified by the validation set, and difficult samples with poor performance will be recorded. After reaching the number of iterations, output the various indicators of the YOLO-Pose human key point detection model, including mean average precision (mAP), single category average precision (AP), precision and recall. When the number of iterations is reached, the training ends, and the parameters and weights of the YOLO-Pose human key point detection model are saved.

[0111] The YOLO-Pose human key point detection model is a universal human key point model trained end-to-end. It can simultaneously output the coordinates of the human detection frame, the confidence of the human detection frame, the coordinates of the human key points, and the confidence of the human key points. The model uses a variety of existing open source data sets for training. It has strong generalization capabilities for different scenarios, excellent performance, and low implementation costs. It does not require a large amount of simulation or sample data to support training. It detects human key points faster and has better performance. Using the YOLO-Pose human key point detection model to identify human key points in a frame of an image can be as follows: Figure 7 As shown in the figure, a frame of image is input into the YOLO-Pose human key point detection model, and the identified human key points will be marked in the output image.

[0112] In addition to the YOLO-Pose human key point detection model, the present application may also use other human key point detection models for human key point recognition, and use other human detection frame models for human detection frame recognition, and the present application does not impose any specific restrictions on this.

[0113] The YOLO-Pose human key point detection model is the preferred human key point detection model for this application. Other methods for detecting human key points include the Top-Down method, the Bottom-up method, etc. YOLO-Pose uses regression logic, so it can directly construct a loss function based on the key point "distance" and the IOU of the human frame to achieve end-to-end training. This is one of the reasons why YOLO-Pose has a better effect, and neither the Top-Down method nor the Bottom-up method can achieve end-to-end training of the model. YOLO-Pose can complete the reasoning of the human frame and key points under the complexity of target detection; generally, the Top-Down method detects the human body first and then cuts out the key points. When there are many targets, the reasoning speed is extremely slow and this method relies on an accurate human body detection model. When the human body detection is incomplete, the joints outside the human body cannot be detected; the Bottom-up method has a faster reasoning speed, but the key point clustering takes more time and there are situations where clustering cannot be performed or clustering errors occur (more likely to occur when there are many human bodies). Therefore, YOLO-Pose has the best human key point detection effect, and YOLO-Pose is the preferred human key point detection model for this application.

[0114] In some embodiments, the electronic device recognizes a frame of target image in a video file to obtain the human body key points of the first human body in the target image. This can be specifically implemented as follows: the electronic device first obtains the human body key points of all human bodies in the target image, and then filters and retains the human body key points of the first human body from all the human body key points.

[0115] As a possible implementation method, the electronic device recognizes the target image, obtains the human body key points of the human body in the target image, and the confidence of the human body key points of the human body; the electronic device determines the human body that meets the screening conditions as the first human body based on the confidence of the human body key points of the human body; the electronic device deletes the human body key points of the human body other than the first human body in the target image, and retains the human body key points of the first human body in the target image; wherein the screening conditions include at least one of the following: the confidence of at least one hand key point among the human body key points is greater than a first preset confidence, and the pixel distance between the hand key point greater than the first preset confidence and the upper edge of the fence is less than a first pixel distance; the first pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points of the human body and a second preset coefficient; the confidence of at least one foot key point among the human body key points is greater than the second preset confidence, and the pixel distance between the foot key point greater than the second preset confidence and the lower edge of the fence is less than the second pixel distance; the second pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points of the human body and a third preset coefficient.

[0116] In practical applications, the first preset reliability, the second preset reliability, the second preset coefficient and the third preset coefficient can be adaptively set by the user according to the specific application scenario, and the embodiment of the present application does not specifically limit it. For example, the typical value of the first preset reliability and the second preset reliability can be 0.5, the typical value of the second preset coefficient can be 0.3, and the typical value of the third preset coefficient can be 1.5.

[0117] For example, suppose the detection personnel set the first preset confidence to 0.5, the second preset confidence to 0.5, and there are human bodies 1, 2, and 3 in the target image, and the screening condition is set to satisfy one of the above two conditions. Among them, the confidence of the hand key point a1 of human body 1 is 0.3, the confidence of the hand key point a2 is 0.2, the confidence of the foot key point b1 of human body 1 is 0.8, the confidence of the foot key point b2 is 0.3, and the distance between the foot key point b1 and the lower edge of the fence is greater than the second pixel distance; the confidence of the hand key point c1 of human body 2 is 0.6, the confidence of the hand key point c2 is 0.4, and the pixel distance between the hand key point c1 and the upper edge of the fence is less than the first pixel distance. The confidence of the foot key point d1 of human 2 is 0.2, the confidence of the foot key point d2 is 0.7, and the distance between the foot key point d2 and the lower edge of the fence is greater than the second pixel distance; the confidence of the hand key point e1 of human 3 is 0.3, the confidence of the hand key point e2 is 0.4, the confidence of the foot key point f1 of human 3 is 0.5, the confidence of the foot key point f2 is 0.9, and the distance between the foot key point f2 and the lower edge of the fence is less than the second pixel distance. It can be seen that human 1 does not meet the screening conditions, human 2 and human 3 meet the screening conditions, human 2 and human 3 are the first human body, and the electronic device actually marks the human key points of human 2 and human 3 in the target image, and does not mark the human key points of human 1.

[0118] Another example, such as Figure 8 As shown, Figure 8 The pixel distance between the two hand key points of human body 1 and the upper edge of the fence is less than the first pixel distance, the confidence of the hand key points is also greater than the first preset confidence, the confidence of the left foot key point is greater than the second preset confidence, and the pixel distance between the left foot key point and the lower edge of the fence is less than the second pixel distance, then human body 1 meets the screening conditions and is determined to be the first human body.

[0119] In this way, by setting screening conditions, human key points with low credibility can be eliminated, that is, human key points with unreliable identification results are not considered in the first human body determination process. This can not only reduce the amount of calculation in the subsequent detection of whether a person has climbing behavior, save the detection time of climbing behavior, but also improve the recognition accuracy of climbing behavior.

[0120] It should be understood that when calculating the pixel distance between the shoulder key point and the hip key point, what is calculated is the pixel distance between the shoulder key point and the hip key point on the same side of the same human body.

[0121] In one example, the pixel distance between the shoulder key point and the hip key point on one side of the human body is selected as the pixel distance between the shoulder key point and the hip key point of the human body. For example, the user pre-sets the calculation of the pixel distance between the left shoulder key point and the left hip key point of the human body, and the electronic device determines that the pixel distance between the left shoulder key point and the left hip key point of the first human body is x1. Assuming that the second preset coefficient is 0.5, the first pixel distance is For another example, the user pre-sets the calculation of the pixel distance between the right shoulder key point and the right hip key point of the human body, and the electronic device determines that the pixel distance between the right shoulder key point and the right hip key point of the first human body is x2. Assuming that the third preset coefficient is 0.6, the second pixel distance is

[0122] As another example, the pixel distance between the shoulder key point and the hip key point of a certain human body can be the average of the pixel distances between the shoulder key points and the hip key points on both sides of the human body. For example, the pixel distance between the left shoulder key point and the left hip key point of the first human body is d1, and the pixel distance between the right shoulder key point and the right hip key point of the first human body is d2. Assuming that the second preset coefficient is 0.2, the first pixel distance is For another example, the pixel distance between the left shoulder key point and the left hip key point of the first human body is d1, and the pixel distance between the right shoulder key point and the right hip key point of the first human body is d2. Assuming that the third preset coefficient is 0.5, the second pixel distance is

[0123] Since the shoulders and hips of the human body are not prone to bending, the pixel distance between the shoulder key point and the hip key point is relatively easy to calculate for everyone, and the calculation result is relatively accurate. Therefore, the first pixel distance and the second pixel distance are determined by multiplying the pixel distance between the key point of the person and the key point of the hip by a preset coefficient. When the pixel distance between the key point of the hand of the human body that is greater than the first preset confidence level and the upper edge of the fence is less than the first pixel distance, it indicates that the person's hand may be resting on the upper edge of the fence and there is suspicion of climbing the fence. These people can be screened out, and in the subsequent process, the fence climbing behavior is further judged based on the climbing behavior score of the human body, and the final climbing behavior detection result is more accurate.

[0124] It should be understood that the hand key points are key points that characterize the position of the person's hand, and the hand key points can be a collection of multiple key points, for example, the hand key points include wrist key points, palm key points, finger key points, etc. The foot key points are key points that characterize the position of the person's foot, and the foot key points can be a collection of multiple key points, for example, the foot key points include ankle key points, sole key points, toe key points, etc.

[0125] As another possible implementation method, the electronic device receives a user's annotation operation on a target image, determines the area annotated by the user as a recognition area, and includes a fence in the recognition area, thereby determining a human body whose key points overlap with the recognition area as a first human body. Finally, the electronic device deletes the key points of human bodies other than the first human body in the target image, and retains the key points of the first human body in the target image.

[0126] For example, Fig. 9 As shown in the figure, the area enclosed by the solid circle is the recognition area marked by the user. Fig. 9 The image includes human body 1, human body 2 and human body 3. Human body 1 and human body 2 overlap with the recognition area, and human body 3 does not overlap with the recognition area. Then human body 1 and human body 2 are the first human bodies. The electronic device actually marks the human key points of human body 1 and human body 2 in the target image, but does not mark the human key points of human body 3.

[0127] The people in the identification area are close to the fence and are likely to climb the fence. Therefore, by marking the identification area, the key points of some people's bodies can be eliminated. This can not only reduce the amount of calculation in the subsequent detection of whether the person has climbing behavior, save the detection time of climbing behavior, but also improve the recognition accuracy of climbing behavior.

[0128] In some embodiments, the electronic device may perform human key point recognition on each frame of the video file, or may perform human key point recognition once every preset number of frames. For example, assuming that the preset number of frames is 4, the electronic device performs human key point recognition on the first received frame, does not perform human key point recognition on the second, third, fourth, and fifth received frames, and performs human key point recognition on the sixth received frame.

[0129] In this way, the electronic device recognizes the key points of the human body for each frame of the video file in real time, and the recognition result obtained is more accurate, reducing the omission of subsequent climbing behavior recognition. The electronic device recognizes the key points of the human body once every preset number of frames. For example, the general video stream is 25 frames / s. If 25 frames / s are detected, one computing card of the electronic device may only be able to detect 2 video streams at the same time; if the frame interval is 4, then when 5 frames / s are detected, 10 video streams can be detected; in the case of sufficient computing resources, there is no obvious difference in the alarm speed and detection time of 2 and 10 channels, mainly because the recognition effect is not significantly improved under the waste of resources, that is, there is no obvious difference in the detection effect of 25 frames / s and 5 frames / s. Therefore, the electronic device recognizes the key points of the human body once every preset number of frames, which can effectively reduce the number of image frames for the electronic device to recognize the key points of the human body, avoid the waste of computing resources of the electronic device, do not need to use more computing resources, reduce the implementation cost of the solution, speed up the recognition of climbing behavior, and save the detection time of climbing behavior.

[0130] S102: Obtain a climbing behavior score of the first person based on the key points of the first person.

[0131] The climbing behavior score is used to characterize the possibility of the first person performing the climbing behavior. In particular, the climbing behavior score has a value range between 0 and 1. Optionally, the climbing behavior score has a value range of any numerical range, and the upper limit and lower limit of the preset interval should be adaptively adjusted.

[0132] In some embodiments, step S102 can be specifically implemented as follows: the electronic device inputs the key points of the first human body into a climbing behavior scoring model constructed based on a machine learning binary classification algorithm, and obtains the climbing behavior score of the first human body output by the climbing behavior scoring model. The climbing behavior scoring model is obtained by training a preset training set, and the preset training set includes key points of human bodies that perform different climbing behaviors and key points of human bodies that do not perform climbing behaviors.

[0133] It should be understood that the number of key points of the human body that perform different climbing behaviors and the number of key points of the human body that do not perform climbing behaviors in the preset training set should be sufficient and balanced, so that the climbing behavior scoring model can achieve better training effects.

[0134] During the specific training process, the climbing behavior scoring model constructs multiple features, and based on the results of climbing behavior or non-climbing behavior corresponding to each sample in the preset training set, determines the degree of influence of different features on whether the human body in the image performs climbing behavior, and then determines whether the human body in the image performs climbing behavior based on the degree of influence of different features.

[0135] By pre-training the climbing behavior scoring model through the preset training set acquired in advance, when climbing behavior recognition is needed, the electronic device can quickly obtain the climbing behavior score of each first human body in the target image based on the pre-trained climbing behavior scoring model, and then the electronic device can quickly and accurately perform the subsequent climbing behavior determination process. In this way, there is no need to judge whether the human body has crossed the fence by the intersection of the human body's two feet with the fence, and more types of climbing behaviors can be identified.

[0136] Climbing behavior can be completed by a single person, and single-person posture classification is easy to implement. The climbing behavior scoring model can also be autonomously learned and continuously iterated based on the reported alarm information.

[0137] When the climbing behavior score of the first person is within the preset range, the electronic device sequentially executes step S103 and step S104. The upper limit and lower limit of the preset range can be set by the user based on actual application conditions. For example, a typical value of the upper limit of the preset range can be 0.9, and a typical value of the lower limit of the preset range can be 0.3.

[0138] S103: Determine action features of the first human body based on the human body key points of the first human body.

[0139] Among them, the action features include one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference; the elbow-shoulder-hip angle is an angle formed by the shoulder key point on the same side of the first human body as the vertex, and the line segment formed by the shoulder key point and the elbow key point, and the line segment formed by the shoulder key point and the hip key point as the sides; the shoulder-hip-knee angle is an angle formed by the hip key point on the same side of the first human body as the vertex, and the line segment formed by the hip key point and the shoulder key point, and the line segment formed by the hip key point and the knee key point as the sides; the knee height difference is the pixel height difference between the two knee key points of the first human body.

[0140] For example, in Figure 5In the schematic diagram of human key points shown, the elbow, shoulder and hip angle of the human body can be the angle formed by the key point represented by sequence number 5 as the vertex, the line segment formed by the key points represented by sequence numbers 5 and 7 as the edge, and the line segment formed by the key points represented by sequence numbers 5 and 11 as the edge; or, the elbow, shoulder and hip angle of the human body can be the angle formed by the key point represented by sequence number 6 as the vertex, the line segment formed by the key points represented by sequence numbers 6 and 8 as the edge, and the line segment formed by the key points represented by sequence numbers 6 and 12 as the edge. The shoulder, hip and knee angles of a human body can be the angle formed by taking sequence number 11 as the vertex, the line segment formed by the key points represented by sequence numbers 11 and 5 as the edge, and the line segment formed by the key points represented by sequence numbers 11 and 13 as the edge; or, the shoulder, hip and knee angles of a human body can be the angle formed by taking sequence number 12 as the vertex, the line segment formed by the key points represented by sequence numbers 12 and 6 as the edge, and the line segment formed by the key points represented by sequence numbers 12 and 14 as the edge. The height difference of the human body's knee is the pixel height difference between the pixel points of the key points represented by sequence numbers 13 and 14, that is, the difference in the pixel coordinates of the pixel points of the key points represented by sequence numbers 13 and 14 in the pixel height direction. For example, the pixel coordinates of the pixel points of the key points represented by sequence numbers 13 and 14 are (128, 365) and (150, 432) respectively, then the height difference of the human body's knee is 432-365=67.

[0141] S104: Identify whether the first person in the target image performs a climbing behavior based on the action feature of the first person.

[0142] In some embodiments, step S104 can be specifically implemented as follows: when the action characteristics include elbow, shoulder and hip angles, and at least one elbow, shoulder and hip angle of the first human body is greater than a first angle threshold, the electronic device determines that the first human body performs a climbing behavior; when the action characteristics include shoulder, hip and knee angles, and at least one shoulder, hip and knee angle of the first human body is less than a second angle threshold, the electronic device determines that the first human body performs a climbing behavior; when the action characteristics include knee height difference, and the knee height difference of the first human body is greater than a target pixel height, the electronic device determines that the first human body performs a climbing behavior; wherein the target pixel height is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points of the first human body and a first preset coefficient.

[0143] Whether a person is climbing can be judged by the elbow-shoulder-hip angle, the shoulder-hip-knee angle or the knee height difference. There are various ways to judge the climbing behavior of the human body, which can adapt to the judgment of the climbing behavior of the human body in various scenarios.

[0144] In actual applications, the first angle threshold, the second angle threshold, and the first preset coefficient can be adaptively set by the user based on the actual application scenario, and the embodiment of the present application does not specifically limit the values ​​of the first angle threshold, the second angle threshold, and the first preset coefficient. For example, the typical value of the first angle threshold is 75°, the typical value of the second angle threshold is 110°, and the typical value of the first preset coefficient is 0.3.

[0145] Optionally, the user may set the judgment condition that the first person performs a climbing behavior to include a plurality of sub-judgment conditions corresponding to each action feature, and the electronic device determines that the first person satisfies the judgment condition for performing a climbing behavior only when the sub-judgment conditions corresponding to each action feature are satisfied. For example, the user sets the judgment condition that the first person performs a climbing behavior to include the sub-judgment conditions corresponding to the elbow-shoulder-hip angle and the shoulder-hip-knee angle. If at least one elbow-shoulder-hip angle of the first person is greater than a first angle threshold, and at least one shoulder-hip-knee angle of the first person is less than a second angle threshold, the electronic device determines that the first person performs a climbing behavior. For another example, the user sets the judgment condition that the first person performs a climbing behavior to include the sub-judgment conditions corresponding to the elbow-shoulder-hip angle and the knee height difference. If at least one elbow-shoulder-hip angle of the first person is greater than the first angle threshold, and the knee height difference of the first person is greater than the target pixel height, the electronic device determines that the first person performs a climbing behavior. For another example, the user sets the judgment condition that the first person performs a climbing behavior to include the sub-judgment conditions corresponding to the shoulder-hip-knee angle and the knee height difference, then when at least one shoulder-hip-knee angle of the first person is less than the second angle threshold, and the knee height difference of the first person is greater than the target pixel height, the electronic device determines that the first person performs a climbing behavior. For another example, the user sets the judgment condition that the first person performs a climbing behavior to include the sub-judgment conditions corresponding to the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference, then when at least one elbow-shoulder-hip angle of the first person is greater than the first angle threshold, at least one shoulder-hip-knee angle of the first person is less than the second angle threshold, and the knee height difference of the first person is greater than the target pixel height, the electronic device determines that the first person performs a climbing behavior.

[0146] Optionally, the user sets the judgment condition that the first person performs a climbing behavior to include a plurality of sub-judgment conditions corresponding to each action feature, and when one of the sub-judgment conditions corresponding to each action feature is satisfied, the electronic device determines that the first person satisfies the judgment condition for performing a climbing behavior. For example, the user sets the judgment condition that the first person performs a climbing behavior to include the sub-judgment conditions corresponding to the elbow-shoulder-hip angle and the shoulder-hip-knee angle, then when at least one elbow-shoulder-hip angle of the first person is greater than a first angle threshold, or when at least one shoulder-hip-knee angle of the first person is less than a second angle threshold, the electronic device determines that the first person performs a climbing behavior. For another example, the user sets the judgment condition that the first person performs a climbing behavior to include the sub-judgment conditions corresponding to the elbow-shoulder-hip angle and the knee height difference, then when at least one elbow-shoulder-hip angle of the first person is greater than the first angle threshold, or when the knee height difference of the first person is greater than the target pixel height, the electronic device determines that the first person performs a climbing behavior. For another example, the user sets the judgment condition that the first person performs climbing behavior to include the sub-judgment conditions corresponding to the shoulder-hip-knee angle and the knee height difference, then when at least one shoulder-hip-knee angle of the first person is less than the second angle threshold, or the knee height difference of the first person is greater than the target pixel height, the electronic device determines that the first person performs climbing behavior. For another example, the user sets the judgment condition that the first person performs climbing behavior to include the sub-judgment conditions corresponding to the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference, then when at least one elbow-shoulder-hip angle of the first person is greater than the first angle threshold, or at least one shoulder-hip-knee angle of the first person is less than the second angle threshold, or the knee height difference of the first person is greater than the target pixel height, the electronic device determines that the first person performs climbing behavior.

[0147] Figure 4 The technical solution shown brings at least the following beneficial effects: firstly, the climbing behavior score of the first person in the target image is determined, and for the first person whose climbing behavior score is within the preset interval, one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference of the first person are used to determine whether the first person has performed a climbing behavior, and there is no need to determine whether the person has crossed the fence to perform a climbing behavior by the intersection of the two feet of the person and the fence. In other words, when the related art determines whether the person has performed a climbing behavior by the intersection of the two feet of the person and the fence, the person has already crossed the fence, and the management personnel may not be able to stop it in time, and the special climbing posture cannot be identified. However, the solution of the present application does not need to determine whether the person has crossed the fence to perform a climbing behavior by the intersection of the two feet of the person and the fence. The various climbing actions of the first person can be identified before the first person crosses the fence, so the management personnel can be notified in time, which buys more time for the management personnel to stop the first person from crossing the fence, and the first person can be stopped from crossing the fence in time to prevent the first person from entering the management area.

[0148] In some embodiments, when the climbing behavior score of the first person is greater than the upper limit of the preset interval, the electronic device determines that the first person has performed a climbing behavior; when the climbing behavior score of the first person is less than the lower limit of the preset interval, the electronic device determines that the first person has not performed a climbing behavior. For example, assuming that the preset interval is [0.3, 0.9], when the climbing behavior score of the first person is less than 0.3, the electronic device determines that the first person has not performed a climbing behavior, and when the climbing behavior score of the first person is greater than 0.9, the electronic device directly determines that the first person has performed a climbing behavior.

[0149] By setting the preset interval, the first human body whose climbing behavior score is less than the lower limit value of the preset interval, that is, the first human body with a lower possibility of performing climbing behavior, can be directly screened out, and the first human body whose climbing behavior score is greater than the upper limit value of the preset interval, that is, the first human body with a higher possibility of performing climbing behavior, can also be directly determined as the human body performing climbing behavior. Therefore, in the process of the electronic device determining the climbing behavior of the human body in the target image, the human body whose climbing behavior score is less than the lower limit value of the preset interval and the human body whose climbing behavior score is greater than the upper limit value of the preset interval can be directly screened out without determining the action characteristics, thereby reducing the workload that the electronic device needs to process and improving the detection speed of the electronic device.

[0150] Furthermore, the flow of people in the scene is constantly changing, and a large number of people or a small number of people has different effects on the detection effect. Therefore, the congestion detection mode or the non-congestion detection mode can be switched according to the number of people in the target image.

[0151] In some embodiments, in congestion detection mode, if there is any first human body performing climbing behavior in the target image, the electronic device records the target image as a first state; in non-congestion detection mode, if the first human body in the target image performs climbing behavior, the electronic device records the first human body as a second state.

[0152] In the congestion detection mode, if the target image and the first preset number of frame images that are continuous with the target image before the target image are all in the first state, the electronic device outputs a first alarm message, and the first alarm message is used to prompt that there is a human body performing climbing behavior; or, in the congestion detection mode, if the target image and the first preset number of frame images that are continuous with the target image before the target image are all in the first state, and the time interval with the last issuance of the first alarm message is greater than the first time interval, the electronic device outputs the first alarm message. Among them, the first preset number and the first time interval can be adaptively set by the user according to the specific application scenario, and the embodiment of the present application does not limit the specific values ​​of the first preset number and the first time interval. For example, the typical value of the first preset number can be 24, and the typical value of the first time interval can be 10 seconds.

[0153] It should be understood that the electronic device detects whether there is climbing behavior in the target image in the video file by performing a detection every target number of frames, that is, there are actually other images between two adjacent target frames, or by performing a detection on each frame in the video file, that is, each frame is the target image, then the first preset number of frames continuous with the target image can be the first preset number of frames continuous with the target image for detecting whether there is climbing behavior, or can be the first preset number of frames actually continuous with the target image in the video file, which can be adaptively set by the user. For example, the electronic device detects once every 4 frames, the first preset number is 4, and there are actually "image 1, image 2, image 3, ..., image 25, image 26, image 27" in the video file, and the target image currently detected is image 27, then the first preset number of frames continuous with the target image can be "image 7, image 12, image 17, image 22", or the first preset number of frames continuous with the target image can be "image 23, image 24, image 25, image 26".

[0154] Exemplarily, assuming that the first preset number is 24 and the first time interval is 10 seconds, if at 02:25:32, the electronic device detects that the target image and the 24 consecutive frames of target images before the target image are all in the first state, and the last time the first alarm information was output was 02:25:18, and the time interval is greater than the first time interval, the electronic device outputs the first alarm information at this time.

[0155] Optionally, the electronic device extracts one frame of image from the video file as the target image every target number, and when the continuous target images within the target duration set by the user are all in the first state, the electronic device outputs the first alarm information, and the first preset number can be determined by the target number within the target duration set by the user. For example, if the target number is 4, the electronic device extracts one frame of image from the video file as the target image every 4 frames of image, and within 1 second, the electronic device will determine 5 frames of target images. Assuming that the target duration is 5 seconds, the first preset number is 5×5-1, that is, when 25 consecutive frames of target images are all in the first state, the electronic device outputs the first alarm information.

[0156] In the non-congestion detection mode, if the target image and the second preset number of frame images that are continuous with the target image before the target image all have the same first human body and the first human body is in the second state, the electronic device outputs the second alarm information, and the second alarm information is used to prompt the first human body to perform climbing behavior; or, in the non-congestion detection mode, if the target image and the second preset number of frame images that are continuous with the target image before the target image all have the same first human body, the first human body is in the second state, and the time interval between the second alarm information indicating the first human body issued last time is greater than the second time interval, the electronic device outputs the second alarm information. Among them, the second preset number and the second time interval can be adaptively set by the user according to the specific application scenario, and the embodiment of the present application does not limit the specific values ​​of the second preset number and the second time interval. For example, the typical value of the second preset number can be 24, and the typical value of the second time interval can be 10 seconds. The specific implementation of the second preset number of frame images that are continuous with the target image can refer to the first preset number of frame images that are continuous with the target image, and will not be repeated here.

[0157] Exemplarily, assuming that the second preset number is 24 and the second time interval is 10 seconds, if at 02:32:52, the electronic device detects that the target image and the 24 consecutive frames of target images before the target image are all in the second state, and the last time the second alarm information was output was 02:31:36, and the time interval is greater than the second time interval, the electronic device outputs the second alarm information at this time.

[0158] Optionally, the electronic device extracts one frame of image from the video file as the target image every target number, and when the continuous target images within the target duration set by the user are all in the second state, the electronic device outputs the second alarm information, and the second preset number can be determined by the target number within the target duration set by the user. For example, if the target number is 4, the electronic device extracts one frame of image from the video file every 4 frames of image as the target image, and within 1 second, the electronic device will determine 5 frames of target images. Assuming that the target duration is 5 seconds, the second preset number is 5×5-1, that is, when 25 consecutive frames of target images are all in the second state, the electronic device outputs the second alarm information.

[0159] In this way, by dividing the detection mode into congestion detection mode and non-congestion detection mode, it is possible to adapt to changes in the number of human bodies in the scene, execute different detection modes for different numbers of human bodies, and the detection effect is more accurate.

[0160] In some embodiments, when climbing behavior detection is performed for the first time, climbing behavior detection is performed in a congestion detection mode to avoid missed detection caused by executing a non-congestion detection mode in a congested state. During the detection process, the electronic device switches between the congestion detection mode and the non-congestion detection mode according to the flow of people in the target image. The specific switching logic is as follows:

[0161] In the case where the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and the third preset number of frame images that are continuous with the target image before the target image is less than or equal to the first value, and there is no human body that performs climbing behavior in the target image, the electronic device switches the congestion detection mode to the non-congestion detection mode; or, in the case where the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and the fourth preset number of frame images that are continuous with the target image before the target image is less than or equal to the second value, and there is no human body that performs climbing behavior in the recognition area, the electronic device switches the congestion detection mode to the non-congestion detection mode. Among them, the third preset number, the first value, the fourth preset number and the second value can be adaptively set by the user according to the actual application scenario, and the embodiment of the present application does not specifically limit it. For example, the typical value of the third preset number can be 24, the typical value of the first value can be 15, the typical value of the fourth preset number can be 24, and the typical value of the second value can be 10. The specific implementation of the third preset number of frame images that are continuous with the target image can refer to the first preset number of frame images that are continuous with the target image, which will not be repeated here. The process of determining the third preset number may refer to the specific process of determining the first preset number or the second preset number, which will not be repeated here.

[0162] As an example, if the congestion detection mode is currently being executed, the number of human detection frames in the currently detected target image is 10, and the number of human detection frames in the 24 consecutive frames of target images of the electronic device before the target image are all less than 15, and no human body performing climbing behavior is detected in the target image, the electronic device switches the congestion detection mode to the non-congestion detection mode, and the target image in the next frame of the current target image corresponds to the non-congestion detection mode.

[0163] As another example, if the congestion detection mode is currently being executed, the number of human detection frames in the recognition area of ​​the currently detected target image is 8, and the number of human detection frames in the recognition area of ​​the electronic device in 24 consecutive frames of target images before the target image is less than 10, and no human body performing climbing behavior is detected in the target image, the electronic device switches the congestion detection mode to the non-congestion detection mode, and the target image in the next frame of the current target image corresponds to the non-congestion detection mode.

[0164] In the case where the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the fifth preset number of frame images that are continuous with the target image before the target image is greater than the first value, and there is no human body in a climbing state in the target image, the electronic device switches the non-congestion detection mode to the congestion detection mode; or, in the case where the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the sixth preset number of frame images that are continuous with the target image before the target image is greater than the second value, and there is no human body in a climbing state in the target image, the electronic device switches the non-congestion detection mode to the congestion detection mode. Among them, the fifth preset number and the sixth preset number can be adaptively set by the user according to the actual application scenario, and the embodiment of the present application does not specifically limit it. For example, the typical value of the fifth preset number can be 24, and the typical value of the sixth preset number can be 24. The specific implementation of the fifth preset number of frame images that are continuous with the target image can refer to the first preset number of frame images that are continuous with the target image, which will not be repeated here. The determination process of the fifth preset number can refer to the specific determination process of the first preset number or the second preset number, which will not be repeated here.

[0165] As an example, if the non-congestion detection mode is currently executed, the number of human detection frames in the currently detected target image is 16, and the number of human detection frames in the 24 consecutive frames of images of the electronic device before the target image is greater than 15, and no human body performing climbing behavior is detected in the target image, the electronic device switches the non-congestion detection mode to the congestion detection mode, and the target image in the next frame of the current target image corresponds to the congestion detection mode.

[0166] As another example, if the non-congestion detection mode is currently being executed, the number of human detection frames in the recognition area of ​​the currently detected target image is 16, and the number of human detection frames in the recognition area of ​​the electronic device in 24 consecutive frames of images before the target image is greater than 15, and no human body performing climbing behavior is detected in the target image, the electronic device switches the non-congestion detection mode to the congestion detection mode, and the target image in the next frame of the current target image corresponds to the congestion detection mode.

[0167] In this way, by switching between the congestion detection mode and the non-congestion detection mode, it can better adapt to the scenarios with more or fewer people in the management area, and the detection mode will only be switched when no human body performing climbing behavior is detected in the target image, which can avoid the situation where a human body performing climbing behavior is missed during the switching process.

[0168] When the detection mode is switched according to the number of human bodies in the recognition area in the target image, the human body has the possibility of climbing the fence only if there is an overlap between the human body and the recognition area. By setting the recognition area and the second value, the switching between the congestion detection mode and the non-congestion detection mode can be more accurately controlled to adapt to the scene with a large or small number of people in the recognition area.

[0169] In some embodiments, in the congestion detection mode, if the target key point of the human body key points of any human body in the target image falls into the management area, the electronic device records the target image as the third state; in the non-congestion detection mode, if the target key point of the human body key points of the second human body in the target image falls into the management area, the electronic device records the second human body as the fourth state, and the second human body is any human body in the target image. For example, the target key point can be a key point of the head of a person, a key point above the shoulders of a person, a key point of a hand of a person, a key point of a foot of a person, etc.

[0170] For example, a specific application scenario is the detection scenario of people climbing the platform screen door in a subway station, such as Fig.10 As shown, Fig.10 The area indicated by the solid line is the track area of ​​the subway, that is, the management area. The key points of the head of person 1 overlap with the track area, that is, the key points of the head of person 1 fall into the management area. If the current target image corresponds to the congestion detection mode, the electronic device records the target image as the third state. If the current target image corresponds to the non-congestion detection mode, the electronic device records person 1 as the fourth state.

[0171] In the congestion detection mode, if the target image and the seventh preset number of frame images that are continuous with the target image before the target image are both in the third state, and the first alarm information has been output within the first preset time length before the target image, the electronic device outputs the third alarm information, and the third alarm information is used to prompt that there is a human body in the target image that illegally enters the management area; or, in the congestion detection mode, if the target image and the seventh preset number of frame images that are continuous with the target image before the target image are both in the third state, the first preset time length before the target image has been output with the first alarm information, and the time interval with the last issuance of the third alarm information is greater than the third time interval, the electronic device outputs the third alarm information. Among them, the seventh preset number, the first preset time length and the third time interval can be adaptively set by the user according to the actual application scenario, and the embodiment of the present application does not specifically limit it. For example, the typical value of the seventh preset number can be 24, the typical value of the first preset time length can be 10 seconds, and the typical value of the third time interval can be 10 seconds. The specific implementation of the seventh preset number of frame images that are continuous with the target image can refer to the first preset number of frame images that are continuous with the target image, and will not be repeated here. The process of determining the seventh preset number may refer to the specific process of determining the first preset number or the second preset number, which will not be repeated here.

[0172] As an example, assuming that the seventh preset number is 24 and the first preset time is 10 seconds, in the congestion detection mode, in the target image corresponding to the time 02:32:21, the electronic device detects that a target key point of a human body falls into the management area, and in the 24 consecutive frames of images before the target image, there are target key points of the human body falling into the management area, and the electronic device outputs the first alarm information between 02:32:11-02:32:21, then the electronic device reports the third alarm information to prompt the user that there is a human body illegally entering the management area in the target image corresponding to the time 02:32:21.

[0173] As another example, assuming that the seventh preset number is 24 and the first preset duration is 10 seconds, in the congestion detection mode, in the target image corresponding to the time point 02:32:21, the electronic device detects that a target key point of a human body falls into the management area, and in the 24 consecutive frames of images before the target image, there are target key points of a human body falling into the management area, the electronic device outputs the first alarm information between 02:32:11-02:32:21, and the time when the electronic device last output the third alarm information is 02:31:56, and the time interval is greater than the third time interval, then the electronic device reports the third alarm information to prompt the user that there is a human body illegally entering the management area in the target image corresponding to the time point 02:32:21.

[0174] In the non-congestion detection mode, if the target image and the eighth preset number of frame images that are continuous with the target image before the target image all have the same second human body and the second human body is in the fourth state, and the second alarm information indicating the second human body has been output within the second preset time length before the target image, the electronic device outputs the fourth alarm information, and the fourth alarm information is used to prompt the second human body to enter the management area illegally; or, in the non-congestion detection mode, if the target image and the eighth preset number of frame images that are continuous with the target image before the target image all have the same second human body and the second human body is in the fourth state, the second alarm information indicating the second human body has been output within the second preset time length before the target image, and the time interval with the last issuance of the fourth alarm information indicating the second human body is greater than the fourth time interval, the electronic device outputs the fourth alarm information. Among them, the eighth preset number, the second preset time length and the fourth time interval can be adaptively set by the user according to the actual application scenario, and the embodiment of the present application does not specifically limit it. For example, the typical value of the eighth preset number can be 24, the typical value of the second preset time length can be 10 seconds, and the typical value of the fourth time interval can be 10 seconds. The specific implementation of the eighth preset number of frame images continuous with the target image can refer to the first preset number of frame images continuous with the target image, which will not be repeated here. The determination process of the eighth preset number can refer to the specific determination process of the first preset number or the second preset number, which will not be repeated here.

[0175] As an example, assuming that the eighth preset number is 24 and the second preset time is 10 seconds, in the non-congestion detection mode, in the target image corresponding to the time 02:32:21, the electronic device detects that the target key points of the human body 1 fall into the management area, and the target key points of the human body 1 in the 24 consecutive frames of images before the target image all fall into the management area, and between 02:32:11-02:32:21, the electronic device outputs the second alarm information indicating that the human body 1 has performed a climbing behavior, then the electronic device reports the fourth alarm information to prompt the user that the human body 1 in the target image corresponding to the time 02:32:21 has illegally entered the management area.

[0176] As another example, assuming that the eighth preset number is 24 and the second preset time is 10 seconds, in the non-congestion detection mode, in the target image corresponding to the time 02:53:11, the electronic device detects that the target key points of the human body 2 fall into the management area, and the target key points of the human body 2 in the 24 consecutive frames of images before the target image all fall into the management area, and the electronic device outputs the second alarm information indicating that the human body 2 has performed a climbing behavior between 02:53:01-02:53:11, and the time when the second alarm information indicating that the human body 2 has performed a climbing behavior was last issued is 02:52:35, and the time interval is greater than the fourth time interval, then the electronic device reports the fourth alarm information to prompt the user that the human body 2 in the target image corresponding to the time 02:53:11 has illegally entered the management area.

[0177] In this way, the fourth alarm information will be output only after the second alarm information indicating the second human body is output, taking the timing relationship between climbing the fence and entering the management area into consideration. Since people who enter the management area illegally will first climb the fence and then enter the management area, this alarm logic can effectively avoid false alarms caused by people who normally enter the management area (for example, entering the management area from the door on the fence).

[0178] The settings of the first time interval, the second time interval, the third time interval and the fourth time interval can prevent the electronic device from continuously outputting alarm information.

[0179] In some embodiments, in a non-congestion detection mode, the electronic device will continuously detect each human body in the target image. Specifically, the continuous detection of the human body can be achieved through the DeepSORT algorithm, and then the electronic device can determine the status of different human bodies in different images.

[0180] In some embodiments, the electronic device updates different states by recording an array.

[0181] In the congestion detection mode, the state recorded in the array is the state of each frame. In the non-congestion detection mode, each first person who performs climbing behavior corresponds to an array, and the state recorded in the array is the state of the first person who performs climbing behavior in each frame of the image corresponding to this array. The array records a set number of states. When the array has recorded a set number of states and needs to record the state of a new image, the earliest state in the array is deleted and the latest state is added.

[0182] During the climbing behavior detection process in the congestion detection mode, the status record update in the array can be specifically implemented as follows:

[0183] When a first human body performing climbing behavior is detected in the image, the state of the frame image is recorded as Y, and when the first human body performing climbing behavior does not exist in the image, the state of the frame image is recorded as N.

[0184] For example, when there is a first person performing climbing behavior in the target image, the updated array is A2=A1[1:]+[Y], where A2 is the updated array and A1 is the array before the update. "A2=A1[1:]+[Y]" means that when the number of states recorded in the array has reached the set number, the earliest state in A1 is deleted and the latest Y is added (at this time, Y represents that the image is in the first state) to obtain A2; if the first person performing climbing behavior is not detected in the image, the image state is updated to A2=A1[1:]+[N], that is, when the number of states recorded in the array has reached the set number, the earliest state in the array before the update is deleted and the latest N is added (at this time, N represents that there is no first person performing climbing behavior in the image) to obtain the updated array. In this way, by checking the array, it can be known whether there is a first person performing climbing behavior in the most recent frames of images, and then it is determined whether to output the first alarm information.

[0185] During the climbing behavior detection process in the non-congestion detection mode, the status record update in the array can be specifically implemented as follows:

[0186] If a first person performing climbing behavior is detected from the image, the state of the first person in the frame image is recorded as Y. If the first person does not perform climbing behavior in the image or the first person is not detected in the image, the state of the first person in the frame image is recorded as N. If the first person performing climbing behavior is detected for the first time in the non-congestion detection mode, a corresponding array is constructed for the first person, and the first person is continuously detected.

[0187] For example, if it is detected that person x in the target image is climbing, the array corresponding to person x is updated to T2 x =T1 x [1:]+[Y],T2 x is the array corresponding to person x after update, T1 x is the array corresponding to person x before the update, "T2 x =T1 x [1:]+[Y]" means to convert T1 x The earliest state in the image is deleted, and the latest state Y is added (at this time, Y indicates that person x in the image is in the second state), and T2 is obtained. x ; If person x in the image does not climb, update the array corresponding to person x to T2 x =T1 x[1:]+[N], that is, deleting the earliest state in the array corresponding to person x and adding the latest N (at this time N indicates that person x in the image has not performed any climbing behavior), to obtain the updated array corresponding to person x. In this way, by checking the array corresponding to person x, the state of person x in the most recent frames of images can be known, and then it can be determined whether it is necessary to output the second alarm information indicating that person x has performed a climbing behavior.

[0188] In this way, when there are too many human bodies in the scene, if the first human body is continuously detected, it may cause the first human body to be misidentified, the first human body to be missed, etc., and the detection success rate of the climbing behavior is reduced. By dividing the congestion detection mode and the non-congestion detection mode, in the congestion detection mode, only attention is paid to whether there is a first human body that performs climbing behavior in the image, and in the non-congestion detection mode, attention is paid to the state of each first human body that performs climbing behavior in each frame image, which can effectively improve the recall rate of detecting climbing behavior in a scene with a large number of human bodies and congestion, avoid the failure of continuous detection of the first human body that performs climbing behavior during congestion, and improve the detection success rate of climbing behavior. In addition, since the presence of the first human body that performs climbing behavior in a single frame image may be a coincidence, and the present application limits the array length, and then the electronic device can determine whether it is necessary to issue an alarm message based on the state represented by the array, and the detection result is more accurate.

[0189] It should be understood that the number of states recorded in the array can be adaptively set by the user according to the actual application scenario, and the length of the array can be different in different detection modes. For example, the user can set the array length to 5 in the congestion detection mode, that is, the array records the latest 5 states of the image, and the array length to 6 in the non-congestion detection mode, that is, the array records the latest 6 states of the first human body.

[0190] In some embodiments, there are four combinations of detection modes and detection contents: ① the detection process of climbing behavior in the congestion detection mode, ② the detection process of people entering the management area in the congestion detection mode, ③ the detection process of climbing behavior in the non-congestion detection mode, and ④ the detection process of people entering the management area in the non-congestion detection mode. Among them, ① and ② can be performed at the same time, and ③ and ④ can be performed at the same time.

[0191] Exemplarily, in the process of detecting climbing behavior in the congestion detection mode, the array length is set to 5. If the array record is {N; N; Y; Y; N}, it can be known that the first person who performs climbing behavior does not exist in the latest detected frame image; in the process of detecting climbing behavior in the non-congestion detection mode, the array length is set to 5. If the array record corresponding to person x is {Y; Y; Y; N; N}, it can be known that person x does not exist in the two latest detected frames or person x does not perform climbing behavior. In the first three frames, person x exists and person x performs climbing behavior. In the process of detecting people entering the management area in the congestion detection mode, the array length is set to 5. If the array record is {N; Y; Y; Y; Y}, it can be known that there are people entering the management area in the latest 4 frames. In the process of detecting people entering the management area in the non-congestion detection mode, the array length is set to 5. If the array record is {Y; N; Y; Y; N}, it can be known that there is no person entering the management area in the latest detected 1 frame image.

[0192] Optionally, the identifiers of different states recorded in the array in the above four cases may be different. For example, in the above ①, Y is used to indicate that the image is in the first state, and N is used to indicate that the image is not in the first state. In the above ②, T may be used to indicate that the image is in the third state, and F may be used to indicate that the image is not in the third state.

[0193] In some embodiments, each execution of the congestion detection mode corresponds to a new array. For example, the electronic device executes ① in period a, executes ③ in period b, and executes ① in period c. The array corresponding to ① in period a is different from the array corresponding to ① in period c. For another example, the electronic device executes ② in period a, executes ④ in period b, and executes ② in period c. The array corresponding to ② in period a is different from the array corresponding to ② in period c.

[0194] In the process of detecting people entering the management area in the congestion detection mode, the status record update in the array can be specifically implemented as follows:

[0195] When a human body entering the management area is detected in the image, the state of the frame image is recorded as Y. When no human body entering the management area is detected in the image, the state of the frame image is recorded as N.

[0196] For example, when there is a human body entering the management area in the target image, the updated array is B2=B1[1:]+[Y], where B2 is the updated array and B1 is the array before the update. "B2=B1[1:]+[Y]" means that when the number of states recorded in the array has reached the set number, the earliest state in B1 is deleted and the latest Y is added (at this time, Y represents the image as the third state) to obtain B2; if no human body entering the management area is detected in the image, the image state is updated to B2=B1[1:]+[N], that is, when the number of states recorded in the array has reached the set number, the earliest state in the array before the update is deleted and the latest N is added (at this time, N represents that there is no human body entering the management area in the image) to obtain the updated array. In this way, by checking the array, it is possible to know whether there is a human body entering the management area in the most recent frames of images, and then determine whether to output the third alarm information.

[0197] In the process of detecting people entering the management area in the non-congestion detection mode, the status record update in the array can be specifically implemented as follows:

[0198] If a second person entering the management area is detected from the image, the state of the second person in the frame image is recorded as Y. If the second person does not enter the management area in the image or the second person is not detected in the image, the state of the second person in the frame image is recorded as N. If the second person entering the management area is detected for the first time in the non-congestion detection mode, a corresponding array is constructed for the second person, and the second person is continuously detected.

[0199] For example, if person z in the target image is detected to enter the management area, the array corresponding to person z is updated to S2 z =S1 z [1:]+[Y],S2 z is the array corresponding to person z after update, S1 z is the array corresponding to person z before the update, "S2 z =S1 z [1:]+[Y]" means to convert S1 z The earliest state in the image is deleted, and the latest state Y is added (at this time, Y indicates that person z in the image is in the fourth state), and S2 is obtained. z ; If person z in the image does not enter the management area, update the array corresponding to person z to S2 z =S1 z[1:]+[N], that is, deleting the earliest state in the array corresponding to person z and adding the latest N (at this time N indicates that person z in the image has not entered the management area or person z has not been detected in the image), to obtain the updated array corresponding to person z. In this way, by checking the array corresponding to person z, the state of person z in the most recent frames of images can be known, and then it can be determined whether it is necessary to output the fourth alarm information indicating that person z has entered the management area.

[0200] In this way, when there are too many human bodies in the scene, if the second human body is continuously detected, it may cause the second human body to be misidentified, the second human body to be missed, etc., and the detection success rate of people entering the management area is reduced. By dividing the congestion detection mode and the non-congestion detection mode, in the congestion detection mode, only attention is paid to whether there are people entering the management area in the image, and in the non-congestion detection mode, attention is paid to the status of each person entering the management area in each frame of the image, which can effectively improve the recall rate of detecting people entering the management area in a large number of human bodies and congested scenes, avoid the failure of continuous detection of people entering the management area during congestion, and improve the detection success rate of people entering the management area. In addition, since the presence of people entering the management area in a single frame image may be a coincidence, and the present application limits the length of the array, and then the electronic device can determine whether it is necessary to issue an alarm message based on the state represented by the array, and the detection result is more accurate.

[0201] In some embodiments, under the above ③, if the first person in the continuous E frame images is not in the second state (the first person does not exist in the image or the first person exists but the first person does not perform climbing behavior), then the array corresponding to the first person is stopped from being updated, and E is a positive integer. For example, E is 5, and in the 5 frames of images after the detection of person a performing climbing behavior in the target image, person a does not exist, and the array corresponding to person a is updated from {Y; N; N; N; N} to {N; N; N; N; N}, then in the subsequent detection process, the array corresponding to person a is no longer updated.

[0202] In this way, by setting the array corresponding to the first human body not to update the array corresponding to the first human body when the F states most recently recorded in the array corresponding to the first human body are not the second state, it is possible to avoid endlessly updating the state of the first human body and record the effective state of the first human body's climbing behavior.

[0203] In some embodiments, under the above ④, if the second person in the consecutive F frame images is not in the fourth state (the second person does not exist in the image or the second person exists but has not entered the management area), then the array corresponding to the second person is stopped from being updated, and F is a positive integer. For example, F is 5, and in the 5 frames of images after the person a is detected to enter the management area in the target image, the person a does not exist, and the array corresponding to the person a is updated from {Y; N; N; N; N} to {N; N; N; N; N}, then in the subsequent detection process, the array corresponding to the person a is no longer updated.

[0204] In this way, by setting the array corresponding to the second human body to no longer update when the F states of the latest record in the array corresponding to the second human body are not the fourth state, it is possible to avoid endlessly updating the state of the second human body and record the effective state of the second human body entering the management area.

[0205] In some embodiments, the first warning information and the second warning information may include a target image of the first person who performs climbing behavior without being marked, or a target image of the first person who performs climbing behavior with being marked. The third warning information and the fourth warning information may include a target image of the second person who enters the management area without being marked, or a target image of the second person who enters the management area with being marked. For example, Fig.11 The display interface of the first warning information is shown, showing a target image of the first human body performing climbing behavior and a text prompt of "climbing behavior exists!".

[0206] In this way, after viewing the alarm information, the user can quickly respond to the climbing behavior or the behavior of entering the management area according to the image included in the alarm information, so as to maintain the safety of the management area.

[0207] In some embodiments, the relevant parameters of the above-mentioned climbing behavior identification process and the detection process of people entering the management area (for example, a preset interval, a first angle threshold, a second angle threshold, a first preset coefficient, a first preset reliability, a second preset coefficient, a second preset reliability, a third preset coefficient, a first preset number, a first time interval, a second preset number, a second time interval, a third preset number, a first numerical value, a fourth preset number, a second numerical value, a fifth preset number, a sixth preset number, a seventh preset number, a first preset duration, a third time interval, an eighth preset number, a second preset duration, a fourth time interval, etc.), as well as the identification area, the upper edge of the fence, and the lower edge of the fence can be set by the user on the user terminal, and then the user terminal sends these contents to the electronic device, and finally the electronic device can perform climbing behavior identification and detect whether there are people entering the management area based on the content set by the user.

[0208] In some embodiments, before performing climbing behavior recognition, the electronic device receives the user's marking operation on the upper edge of the fence in the image of the video file, and then the electronic device determines the position of the upper edge of the fence in each frame of the video file; and / or, the electronic device receives the user's marking operation on the lower edge of the fence in the image of the video file, and then the electronic device determines the position of the lower edge of the fence in each frame of the video file; and / or, the electronic device receives the user's marking operation on the identification area in the image of the video file, and then the electronic device determines the position of the identification area in each frame of the video file; and / or, the electronic device receives the user's marking operation on the management area in the image of the video file, and then the electronic device determines the position of the management area in each frame of the video file.

[0209] For example, for a frame of image captured by a camera installed at a point, such as Fig.12 As shown, the user marks the management area, identification area, upper edge of the fence and lower edge of the fence. The electronic device can directly overlay the marked management area, identification area, upper edge of the fence and lower edge of the fence on each frame image after the frame image to obtain the positions of the identification area, upper edge of the fence and lower edge of the fence on each frame image.

[0210] In this way, the inspectors mark the management area, identification area, upper edge of the fence and lower edge of the fence by themselves. The electronic equipment can identify climbing behavior and detect whether there are people entering the management area based on the content marked by the user, thereby meeting the user's detection requirements.

[0211] The following describes the climbing behavior identification process and the detection process of people entering the management area provided by this application from the perspective of the overall process:

[0212] The electronic device reads the video stream and the climbing behavior detection rule (S1), wherein the climbing behavior detection rule includes relevant parameters for climbing behavior detection, as well as a management area, an upper edge of a fence, and a lower edge of a fence.

[0213] The image is input into the human key point detection model to obtain the human key points of the people in the image.

[0214] like Fig.13 As shown in (a) in the climbing behavior recognition process:

[0215] First, a preliminary target screening is performed on the human body in the target image. The human body that passes the preliminary target screening satisfies: the confidence of at least one hand key point among the human body key points is greater than a first preset confidence, and the pixel distance between the hand key point greater than the first preset confidence and the upper edge of the fence is less than a first pixel distance, and the confidence of at least one foot key point among the human body key points is greater than a second preset confidence, and the pixel distance between the foot key point greater than the second preset confidence and the lower edge of the fence is less than the second pixel distance.

[0216] The human bodies that pass the initial target screening are classified into two categories. If the climbing behavior score is greater than 0.9, the human body is marked as performing climbing behavior in the target image. If the climbing behavior score is less than 0.3, the human body is marked as not performing climbing behavior in the target image. If the climbing behavior score is in the interval [0.3, 0.9], it is judged whether the human body performs climbing behavior in the target image based on the action characteristics of the human body.

[0217] Determine whether the target image is in congestion detection mode.

[0218] In the non-congestion detection mode, each human body in the target image is continuously detected. If the human body performs climbing behavior in the target image, the state of the human body is updated to: T2 = T1 [1:] + [Y]. If the human body does not perform climbing behavior in the target image, the state of the human body is updated to: T2 = T1 [1:] + [N]. If all the state arrays of a human body are Y, it is determined whether the target outputs climbing behavior alarm information for the first time, or the time interval from the last output of climbing behavior alarm information is greater than the preset time interval. If so, the climbing behavior alarm information is output to indicate that the human body performs climbing behavior in the target image. If not, the above step S1 is continued.

[0219] In the congestion detection mode, if there is a human body performing climbing behavior in the target image, the state of the target image is updated to: A2 = A1[1:] + [Y]; if there is no human body performing climbing behavior in the target image, the state of the target image is updated to A2 = A1[1:] + [N]. If all the state arrays corresponding to the target image are Y, it is determined whether the target outputs climbing behavior alarm information for the first time, or the time interval from the last output of climbing behavior alarm information is greater than the preset time interval. If so, the climbing behavior alarm information is output, indicating that there is a human body performing climbing behavior in the target image; if not, the above step S1 is continued.

[0220] like Fig.13 As shown in (b) in the process of detecting people entering the management area:

[0221] First, it is determined whether there is a human body whose target key point falls into the management area in the target image. If so, it is determined that the human body is a human body entering the management area. If not, the above step S1 is continued.

[0222] Determine whether the target image is in congestion detection mode.

[0223] In the non-congestion detection mode, each human body in the target image is continuously detected. If a human body enters the management area in the target image, the state of the human body is updated to: T2 = T1 [1:] + [Y]. If the human body does not enter the management area in the target image, the state of the human body is updated to: T2 = T1 [1:] + [N]. If the state array of a human body is all Y, it is determined whether an alarm message indicating the climbing behavior of the human body has appeared within the preset time period before. If it has appeared, it is determined whether the target outputs the alarm message of entering the management area for the first time, or the time interval from the last output of the alarm message of entering the management area is greater than the preset time interval. If so, the alarm message of entering the management area is output to indicate that the human body has entered the management area in the target image. If not, the above step S1 is continued. If the alarm message indicating the climbing behavior of the human body has not appeared within the preset time period before, the above step S1 is continued.

[0224] In the congestion detection mode, if there is a human body in the target image that enters the management area, the state of the target image is updated to: A2 = A1 [1:] + [Y]; if there is no human body in the target image that enters the management area, the state of the target image is updated to A2 = A1 [1:] + [N]. If the state array corresponding to the target image is all Y, it is determined whether there has been a climbing behavior alarm message within the preset time period before. If it has, it is determined whether the target outputs the alarm message of entering the management area for the first time, or the time interval from the last output of the alarm message of entering the management area is greater than the preset time interval. If so, the alarm message of entering the management area is output, indicating that there is a human body in the target image that enters the management area. If not, the above step S1 is continued. If there has been no climbing behavior alarm message within the preset time period before, the above step S1 is continued.

[0225] The following describes the process of reporting climbing behavior identification events and the process of reporting detection events of people entering the management area in this application from the perspective of the overall process:

[0226] like Fig.14As shown, first receive the real-time video stream sent by the front-end camera, analyze the real-time video stream, and determine whether the climbing behavior alarm conditions and / or the alarm conditions for entering the management area are met. If they are met, report the alarm information and pull the event video to manually verify the alarm information. If the alarm information indicates the behavior of climbing the fence, manually verify the follow-up of the climbing behavior. If there is a risk of climbing into the management area in the future, it will be pushed to the staff at the image acquisition location immediately to notify the staff to handle it in time. If the alarm information indicates the behavior of entering the management area after climbing, notify the staff at the image acquisition location to handle it in time. If not satisfied, continue to read the next frame of the image.

[0227] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to achieve the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy to realize that the technical goals in this field are combined with the units and algorithm steps of each example described in the embodiments disclosed in this article, and the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical goals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0228] like Fig.15 As shown, the present application embodiment further provides a climbing behavior recognition device, which is used to execute the climbing behavior recognition method shown in the above method embodiment. The climbing behavior recognition device 400 includes: a processing module 401 and a transmission module 402.

[0229] The processing module 401 is used to identify a frame of a target image in a video file to obtain the key points of the first human body in the target image; the processing module 401 is also used to obtain a climbing behavior score of the first human body based on the key points of the first human body, and the climbing behavior score is used to characterize the possibility of the first human body performing a climbing behavior; the processing module 401 is also used to determine the action characteristics of the first human body based on the key points of the first human body when the climbing behavior score of the first human body is within a preset interval, and the action characteristics include one or more of the elbow-shoulder-hip angle, the shoulder-hip-knee angle, and the knee height difference; the elbow-shoulder-hip angle The angle is an angle formed by the shoulder key point on the same side of the first human body as the vertex, and the line segment formed by the shoulder key point and the elbow key point, and the line segment formed by the shoulder key point and the hip key point as the sides; the shoulder-hip-knee angle is an angle formed by the hip key point on the same side of the first human body as the vertex, and the line segment formed by the hip key point and the shoulder key point, and the line segment formed by the hip key point and the knee key point as the sides; the knee height difference is the pixel height difference between the two knee key points of the first human body; the processing module 401 is also used to identify whether the first human body in the target image performs a climbing behavior based on the action characteristics of the first human body.

[0230] In one possible implementation, the processing module 401 is specifically used to: determine that the first human body performs climbing behavior when the action characteristics include elbow, shoulder and hip angles, and at least one elbow, shoulder and hip angle of the first human body is greater than a first angle threshold; determine that the first human body performs climbing behavior when the action characteristics include shoulder, hip and knee angles, and at least one shoulder, hip and knee angle of the first human body is less than a second angle threshold; determine that the first human body performs climbing behavior when the action characteristics include knee height difference, and the knee height difference of the first human body is greater than a target pixel height; wherein the target pixel height is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points of the first human body and a first preset coefficient.

[0231] In another possible implementation, the processing module 401 is specifically used to: input the human body key points of the first human body into a climbing behavior scoring model constructed based on a binary classification algorithm, and obtain a climbing behavior score of the first human body output by the climbing behavior scoring model, wherein the climbing behavior scoring model is trained by a preset training set, and the preset training set includes preset sample images, and the preset sample images contain at least one of the first human body with an elbow, shoulder and hip angle less than or equal to a first angle threshold, the first human body with a shoulder, hip and knee angle greater than a second angle threshold, and the first human body with a knee height difference less than or equal to a target pixel height.

[0232] In another possible implementation, the processing module 401 is further used to: determine that the first person has performed a climbing behavior when the climbing behavior score of the first person is greater than an upper limit value of a preset interval; and determine that the first person has not performed a climbing behavior when the climbing behavior score of the first person is less than a lower limit value of the preset interval.

[0233] In another possible implementation, the processing module 401 is specifically used to: identify the target image, obtain the human body key points of the human body in the target image, and the confidence of the human body key points of the human body; based on the confidence of the human body key points of the human body, determine the human body that meets the screening conditions as the first human body; delete the human body key points of the human body other than the first human body in the target image, and retain the human body key points of the first human body in the target image; wherein the screening conditions include at least one of the following: the confidence of at least one hand key point among the human body key points is greater than the first preset confidence, and is greater than the first preset confidence. The pixel distance between the hand key point with the confidence level and the upper edge of the fence is less than the first pixel distance; the first pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points and the second preset coefficient; the confidence level of at least one foot key point among the human body key points is greater than the second preset confidence level, and the pixel distance between the foot key point greater than the second preset confidence level and the lower edge of the fence is less than the second pixel distance; the second pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the human body key points and the third preset coefficient.

[0234] In another possible implementation, the processing module 401 is further used for: in a congestion detection mode, if there is any first human body in the target image performing a climbing behavior, recording the target image as a first state; in a non-congestion detection mode, if the first human body in the target image performs a climbing behavior, recording the first human body as a second state; the transmission module 402 is used for: in a congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, outputting a first alarm message, the first alarm message being used to prompt that there is a human body performing a climbing behavior; or, in a congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, and the first alarm message is the same as the last time the first alarm was issued. When the time interval of the warning information is greater than the first time interval, the first warning information is output; the transmission module 402 is also used for: in the non-congestion detection mode, if the same first human body exists in the target image and the second preset number of frame images that are continuous with the target image before the target image, and the first human body is in the second state, the second warning information is output, and the second warning information is used to prompt the first human body to perform climbing behavior; or, in the non-congestion detection mode, if the same first human body exists in the target image and the second preset number of frame images that are continuous with the target image before the target image, and the first human body is in the second state, and the time interval with the last second warning information indicating the first human body is greater than the second time interval, the second warning information is output.

[0235] In another possible implementation, the processing module 401 is further used for: when the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a third preset number of frame images that are continuous with the target image before the target image is less than or equal to the first value, and there is no human body performing climbing behavior in the target image, switching the congestion detection mode to the non-congestion detection mode; or, when the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a fourth preset number of frame images that are continuous with the target image before the target image is less than or equal to the second value, and there is no human body performing climbing behavior in the identification area, switching the congestion detection mode. to the non-congestion detection mode; the processing module 401 is also used for: when the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the fifth preset number of frame images that are continuous with the target image before the target image is greater than the first value, and there is no human body in a climbing state in the target image, switching the non-congestion detection mode to the congestion detection mode; or, when the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the sixth preset number of frame images that are continuous with the target image before the target image is greater than the second value, and there is no human body in a climbing state in the target image, switching the non-congestion detection mode to the congestion detection mode.

[0236] In another possible implementation, the processing module 401 is also used for: in a congestion detection mode, if a target key point among the human key points of any human body in the target image falls into the management area, the target image is recorded as the third state; in a non-congestion detection mode, if a target key point among the human key points of a second human body in the target image falls into the management area, the second human body is recorded as the fourth state, and the second human body is any human body in the target image; the transmission module 402 is also used for: in a congestion detection mode, if the target image and the seventh preset number of frame images continuous with the target image before the target image are all in the third state, and the first alarm information has been output within the first preset time length before the target image, the third alarm information is output, and the third alarm information is used to prompt that there is a human body in the target image that illegally enters the management area; or, in a congestion detection mode, if the target image and the seventh preset number of frame images continuous with the target image before the target image are all in the third state, within the first preset time length before the target image The first alarm information is output, and the time interval with the last issuance of the third alarm information is greater than the third time interval, the third alarm information is output; the transmission module 402 is also used for: in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image, and the second human body is in the fourth state, and the second alarm information indicating the second human body has been output within the second preset time length before the target image, the fourth alarm information is output, and the fourth alarm information is used to prompt that the second human body has illegally entered the management area; or, in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image, and the second human body is in the fourth state, the second alarm information indicating the second human body has been output within the second preset time length before the target image, and the time interval with the last issuance of the fourth alarm information indicating the second human body is greater than the fourth time interval, the fourth alarm information is output.

[0237] It should be noted that Fig.15 The division of modules in the example is schematic and is only a logical function division. There may be other division methods in actual implementation. For example, two or more functions may be integrated into one processing module. The above integrated modules may be implemented in the form of hardware or software function modules.

[0238] Another embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the climbing behavior identification method shown in the above embodiment.

[0239] In actual implementation, the processing module 401 and the transmission module 402 can be implemented by the processor of the electronic device calling the computer program code in the memory. The specific execution process can refer to the description of the climbing behavior identification method above, which will not be repeated here.

[0240] Another embodiment of the present application further provides a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the steps of the climbing behavior identification method shown in the above embodiment are implemented.

[0241] In another embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions. When the computer instructions are executed by a processor, the steps of the climbing behavior identification method shown in the above embodiment are implemented.

[0242] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instruction is loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that can be accessed by the computer or a data storage device such as a server, data center, etc. that contains one or more servers that can be integrated with the medium. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), etc.

[0243] The above is only a specific implementation of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific implementation provided by the present application, which should all be included in the protection scope of the present application.

Claims

1. A climbing behavior recognition method, characterized in that: include: Based on the human body key points of the first human body in a frame of target image of the video file, the action features of the first human body are determined, and the action features include one or more of elbow-shoulder-hip angle, shoulder-hip-knee angle, and knee height difference; the elbow-shoulder-hip angle is an angle formed by the shoulder key point on the same side of the first human body as a vertex, and a line segment formed by the shoulder key point and the elbow key point, and a line segment formed by the shoulder key point and the hip key point as sides; the shoulder-hip-knee angle is an angle formed by the hip key point on the same side of the first human body as a vertex, and a line segment formed by the hip key point and the shoulder key point, and a line segment formed by the hip key point and the knee key point as sides; the knee height difference is the pixel height difference between two knee key points of the first human body; Based on the motion feature of the first human body, it is identified whether the first human body in the target image performs a climbing behavior.

2. The method according to claim 1, characterized in that The identifying, based on the motion feature of the first human body, whether the first human body in the target image performs a climbing behavior includes: In a case where the action feature includes the elbow, shoulder and hip angles, and at least one elbow, shoulder and hip angle of the first human body is greater than a first angle threshold, determining that the first human body performs a climbing behavior; In a case where the action feature includes the shoulder, hip and knee angles, and at least one of the shoulder, hip and knee angles of the first human body is less than a second angle threshold, determining that the first human body performs a climbing behavior; When the action feature includes the knee height difference, and the knee height difference of the first human body is greater than the target pixel height, determining that the first human body performs a climbing behavior; The target pixel height is the product of the pixel distance between the shoulder key point and the hip key point among the key points of the first human body and a first preset coefficient.

3. The method according to claim 1, characterized in that The human body in the target image that meets the screening condition is the first human body; wherein the screening condition includes at least one of the following: The pixel distance between at least one of the human body key points and the upper edge of the fence is less than a first pixel distance; the first pixel distance is used to represent a pixel distance threshold at which the human body hand is connected to the upper edge of the fence; The pixel distance between at least one foot key point among the human body key points and the lower edge of the fence is less than a second pixel distance; the second pixel distance is used to represent the pixel distance threshold where the human foot is connected to the lower edge of the fence.

4. The method according to claim 3, characterized in that The confidence of the at least one hand key point is greater than a first preset confidence; the confidence of the at least one foot key point is greater than a second preset confidence; The first pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the key points of the human body and the second preset coefficient; The second pixel distance is the product of the pixel distance between the shoulder key point and the hip key point among the key points of the human body and the third preset coefficient.

5. The method according to claim 1, characterized in that: Before determining the motion features of the first human body based on the human body key points of the first human body, the method further includes: Obtaining a climbing behavior score of the first human body based on the human body key points of the first human body, wherein the climbing behavior score is used to characterize the possibility of the first human body performing a climbing behavior; The determining the motion features of the first human body based on the human body key points of the first human body includes: When the climbing behavior score of the first human body is within a preset range, the motion characteristics of the first human body are determined based on the human body key points of the first human body.

6. The method according to claim 1, characterized in that The method further comprises: In the congestion detection mode, if there is any first human body in the target image performing a climbing behavior, the target image is recorded as being in the first state; In the non-congestion detection mode, if the first human body in the target image performs a climbing behavior, the first human body is recorded as being in the second state; In the congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, a first alarm message is output, and the first alarm message is used to prompt that a human body is performing a climbing behavior; or, in the congestion detection mode, if the target image and a first preset number of frame images that are continuous with the target image before the target image are all in the first state, and the time interval from the last issuance of the first alarm message is greater than the first time interval, the first alarm message is output; In the non-congestion detection mode, if the target image and a second preset number of frame images that are continuous with the target image before the target image all contain the same first human body and the first human body is in the second state, a second alarm message is output, and the second alarm message is used to prompt the first human body to perform a climbing behavior; or, in the non-congestion detection mode, if the target image and a second preset number of frame images that are continuous with the target image before the target image all contain the same first human body and the first human body is in the second state, and the time interval between the last second alarm message indicating the first human body and the second time interval is greater than the second time interval, the second alarm message is output.

7. The method according to claim 6, characterized in that The method further comprises: In the case where the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a third preset number of frame images that are continuous with the target image before the target image is less than or equal to the first value, and there is no human body performing climbing behavior in the target image, the congestion detection mode is switched to the non-congestion detection mode; or, in the case where the target image corresponds to the congestion detection mode, if the number of human bodies in the target image and a fourth preset number of frame images that are continuous with the target image before the target image is less than or equal to the second value, and there is no human body performing climbing behavior in the identification area, the congestion detection mode is switched to the non-congestion detection mode; In the case where the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the fifth preset number of frame images that are continuous with the target image before the target image is greater than the first value, and there is no human body in a climbing state in the target image, the non-congestion detection mode is switched to the congestion detection mode; or, in the case where the target image corresponds to the non-congestion detection mode, if the number of human bodies in the target image and the sixth preset number of frame images that are continuous with the target image before the target image is greater than the second value, and there is no human body in a climbing state in the target image, the non-congestion detection mode is switched to the congestion detection mode.

8. The method according to claim 6, characterized in that The method further comprises: In the congestion detection mode, if a target key point among the key points of any human body in the target image falls into the management area, the target image is recorded as the third state; In the non-congestion detection mode, if a target key point among the human key points of a second human body in the target image falls into a management area, the second human body is recorded as being in a fourth state, and the second human body is any human body in the target image; In the congestion detection mode, if the target image and the seventh preset number of frame images that are continuous with the target image before the target image are all in the third state, and the first alarm information has been output within the first preset time length before the target image, the third alarm information is output, and the third alarm information is used to prompt that there is a human body in the target image that illegally enters the management area; or, in the congestion detection mode, if the target image and the seventh preset number of frame images that are continuous with the target image before the target image are all in the third state, the first alarm information has been output within the first preset time length before the target image, and the time interval between the last issuance of the third alarm information and the last issuance of the third alarm information is greater than the third time interval, the third alarm information is output; In the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image, and the second human body is in the fourth state, and the second alarm information indicating the second human body has been output within the second preset time period before the target image, the fourth alarm information is output, and the fourth alarm information is used to prompt the second human body to illegally enter the management area; or, in the non-congestion detection mode, if the same second human body exists in the target image and the eighth preset number of frame images that are continuous with the target image before the target image, and the second human body is in the fourth state, the second alarm information indicating the second human body has been output within the second preset time period before the target image, and the time interval between the last issuance of the fourth alarm information indicating the second human body and the fourth time interval is greater than the fourth time interval, the fourth alarm information is output.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the climbing behavior identification method according to any one of claims 1 to 8.

10. A climbing behavior recognition system, characterized in that: include: An image acquisition device, used for acquiring video files; An electronic device, configured to execute the method according to any one of claims 1 to 8; A prompting device is used to receive and display a first alarm message, a second alarm message, a third alarm message and a fourth alarm message sent by the electronic device, wherein the first alarm message is used to indicate that a human body performs climbing behavior in a frame of a target image of the video file, the second alarm message is used to indicate that a first human body performs climbing behavior, the third alarm message is used to indicate that a human body illegally enters a management area in a frame of a target image of the video file, and the fourth alarm message is used to indicate that a second human body illegally enters the management area.