Parabolic detection method and system based on human body local binary feature point alignment

By using a method based on the alignment of local binary feature points of the human body and the recognition of salient objects, the throwing action and the thrown object of pedestrians can be accurately detected, which solves the problem of accuracy in the detection of objects thrown in community monitoring and enables timely identification and warning of pedestrian throwing behavior.

CN115205783BActive Publication Date: 2025-11-04QINGDAO WINDAKA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210821453.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-11-04
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately detect pedestrians' throwing actions and objects in community surveillance, resulting in inaccurate object detection.

Method used

A method based on local binary feature point alignment of the human body is adopted, combined with salient object recognition. By detecting key points of the pedestrian's limbs, torso and head, the throwing action type is determined, and a frame-by-frame tracking salient object detection method is used to identify and track the thrown object.

Benefits of technology

It improves the accuracy of detecting pedestrian littering, enabling timely warnings to community security guards to stop inappropriate behavior and enhancing community safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205783B_ABST
    Figure CN115205783B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of target detection, and proposes a parabola detection method and system based on human local binary feature point alignment, which comprises the following steps: acquiring a continuous frame image to be detected, calibrating key feature points of a pedestrian in an image based on local binary features of pixel points, forming shape contours of each part of the pedestrian according to the key feature points, and judging whether the pedestrian makes a throwing action and the type of the throwing action according to the relative positions of the center points of each part of the human body; when a throwing action of the pedestrian in the image is recognized, taking the center point of the trunk of the pedestrian as a reference, and determining a possible area of a thrown object according to the type of the throwing action of the pedestrian; and detecting the throwing of the object by the pedestrian by using a salient object detection method according to pixel context information, and determining that the pedestrian throws an object when making a throwing action if the object is detected. The present disclosure simultaneously uses the method of salient object recognition to detect and track the thrown object based on the alignment of human local binary feature points, thereby improving the accuracy of parabola detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of object detection, and in particular, to a paring detection method and system based on local binary feature point alignment of human body. BACKGROUND

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute the prior art.

[0003] Object detection is an important research direction in the field of computer vision, and the applicability requirement for the object is relatively wide. As long as the object is visible to the naked eye and actually exists, it can be detected through related image retrieval technology, so as to further expand to other applications based on the detected object. For example, the most widely used vehicle tracking technology is based on vehicle detection and attribute matching method; the purpose of paring detection in the community is mainly to prevent residents in the community from making paring behaviors that may injure others or their own property in inappropriate places (such as community hall, parking lot, etc.), and the community monitoring will immediately give a warning to remind the community security personnel to stop the behavior of the resident.

[0004] For paring detection, the throwing action of the pedestrian and the object appearing in the throwing direction are detected at the same time, that is, it is considered that the pedestrian is making a throwing action, and the security personnel will go to the scene in time to stop it.

[0005] However, the paring detection has the following difficulties: how to accurately model the throwing action of the pedestrian to judge the current behavior of the pedestrian, and to judge the specific throwing object, which is the key problem faced by paring detection. SUMMARY

[0006] In order to solve the above problems, the present disclosure provides a paring detection method and system based on local binary feature point alignment of human body, which detects and tracks the throwing object by using the method of significant object recognition based on local binary feature point alignment of human body, and can meet the paring detection demand in complex community scenes.

[0007] In order to achieve the above purpose, the present disclosure adopts the following technical solutions:

[0008] One or more embodiments provide a paring detection method based on local binary feature point alignment of human body, comprising the following steps:

[0009] Obtaining a continuous frame image to be detected, labeling key feature points of a pedestrian in the image based on local binary features of pixel points, forming shape contours of each part of the pedestrian according to the key feature points, and judging whether the pedestrian makes a throwing action and the type of the throwing action according to the relative positions of the center points of each part of the human body.

[0010] When the throwing action of the pedestrian in the image is recognized, a region where the thrown object is likely to appear is determined according to the type of the throwing action of the pedestrian, with the center point of the trunk of the pedestrian as the reference;

[0011] Based on the determined region, the thrown object of the pedestrian is detected according to the pixel context information by using a salient object detection method, and if an object is detected in the determined region, it is determined that the object is thrown by the pedestrian when the throwing action is made.

[0012] One or more embodiments provide a thrown object detection system based on alignment of local binary feature points of a human body, comprising:

[0013] A throwing action recognition module is configured to acquire a continuous frame image to be detected, label key feature points of a pedestrian in the image based on local binary features of pixels, form shape contours of limbs, a trunk and a head of the pedestrian according to the key feature points, and determine whether the pedestrian makes a throwing action and the type of the throwing action according to the relative positions of the center points of the parts of the human body.

[0014] A thrown object region determination module is configured to determine a region where a thrown object is likely to appear when the throwing action of the pedestrian in the image is recognized, with the center point of the trunk of the pedestrian as the reference.

[0015] A detection determination module is configured to detect the thrown object of the pedestrian based on the determined region according to the pixel context information by using a salient object detection method, and if an object is detected in the determined region, it is determined that the object is thrown by the pedestrian when the throwing action is made.

[0016] An electronic device includes a memory and a processor, and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the steps of the above method are completed.

[0017] A computer readable storage medium for storing computer instructions, when the computer instructions are executed by a processor, the steps of the above method are completed.

[0018] Compared with the prior art, the beneficial effects of the present disclosure are:

[0019] The present disclosure detects the throwing action of the pedestrian based on alignment of key points of parts of the human body based on local binary features, simultaneously detects and tracks the thrown object by using a salient object recognition method, integrates the analysis results from the pedestrian and the object, and determines whether the pedestrian in the monitoring range truly makes a throwing behavior, thereby improving the accuracy of detection.

[0020] The advantages of the present disclosure and the advantages of additional aspects will be described in detail in the specific embodiments below. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings, which form a part of this description, are included to provide further understanding of the disclosure, and are incorporated in and constitute a part of this description. The illustrative embodiments of the disclosure and their

[0022] Fig. 1 is a flow chart of the method of embodiment 1 of the disclosure;

[0023] Fig. 2 is a structural schematic diagram of the Body-LBF feature point alignment model of embodiment 1 of the disclosure. DETAILED DESCRIPTION

[0024] The disclosure will be further described below in conjunction with the drawings and embodiments.

[0025] It should be noted that the following detailed description is merely exemplary in nature and is intended to provide further description of the disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0026] It should be noted that the terms used herein are merely for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the disclosure. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of the features, steps, operations, devices, components and / or combinations thereof. It should be noted that the various embodiments in the disclosure and the features in the embodiments can be combined with each other without conflict. The embodiments will be described in detail below with reference to the accompanying drawings.

[0027] Currently, the key of the parabolic detection mainly lies in two points, the first is the detection of the throwing action of the pedestrian, and the second is the detection of the thrown object. In general, since the action range of the pedestrian is relatively small, the modeling of the throwing action must be based on the fine-grained features, accurate to the actions of the limbs, torso and head. The objects that can be thrown by the pedestrian generally belong to small targets moving quickly, which are not easy to capture and extract features in the same image.

[0028] Therefore, the disclosure proposes a parabola detection method based on human key point alignment and salient object detection for the community monitoring application scenario in view of the above-mentioned difficulties in parabola detection. The method proposes a local binary feature point alignment model based on the body (Body-LBF) for the detection of pedestrian throwing action. The Body-LBF establishes a feature point alignment model for four kinds of pedestrian throwing actions, and uses the action model as a benchmark to determine the throwing action of the pedestrian in the image. The accuracy of the action determination mainly goes into the torso, limbs, and head of the human body. In addition, in order to solve the problem of too low signal-to-noise ratio and difficult to capture in the detection of the thrown object, a salient object detection method based on pixel context is used to calibrate and track the thrown object frame by frame. By establishing tracking for the pedestrian throwing action and the thrown object, the parabola behavior of the pedestrian can be captured in real time and accurately. The following will be described with specific examples.

[0029] Embodiment 1

[0030] In the technical solution disclosed in one or more embodiments, as shown in Figs. 1-2 The parabola detection method based on local binary feature point alignment of the human body includes the following steps:

[0031] Step 1, acquire a continuous frame image to be detected, label the key feature points of the pedestrian in the image based on the local binary feature of the pixel points, form the shape contour of the limbs, torso, and head of the pedestrian according to the key feature points, and determine whether the pedestrian makes a throwing action and the type of the throwing action according to the relative position of the center points of each part of the human body;

[0032] Step 2, when the pedestrian in the image is identified to make a throwing action, determine the area where the thrown object is likely to appear according to the type of the throwing action of the pedestrian with the center point of the torso of the pedestrian as a benchmark;

[0033] Step 3, based on the determined area, detect the parabola of the pedestrian by using a salient object detection method according to the pixel context information; if an object is detected in the determined area, it is determined that the pedestrian has thrown an object when making a throwing action.

[0034] In this embodiment, the key points of the human body parts are aligned based on the local binary feature to detect the throwing action of the pedestrian, and a salient object recognition method is used to detect and track the thrown object. The analysis results from the pedestrian and the object are integrated, and it is determined whether the pedestrian in the monitoring range has truly made a parabola behavior, thereby improving the accuracy of the detection.

[0035] Further, after determining that the pedestrian throws the object when making the throwing action, the method further comprises the steps of tracking the pedestrian and the thrown object: if the last frame contains the throwing action labeling result of the pedestrian, the labeling result of the last frame and the labeling result of the current frame are matched by using the Hungarian algorithm to track the throwing pedestrian and the object.

[0036] In step 1, a continuous frame image to be detected is obtained, key feature points of a pedestrian in the image are labeled, and a shape contour of limbs, a trunk and a head of the pedestrian is formed according to the key feature points, and a method for judging whether the pedestrian makes a throwing action and a throwing action type according to relative positions of center points of each part of the human body comprises the following steps:

[0037] In step 11, a community pedestrian throwing action data set is obtained, and key feature points of a human body are labeled, and the feature points can form a shape contour of limbs, a trunk and a head of the pedestrian.

[0038] The detection of the embodiment can be applied to a community monitoring scene, and the collected data set can mainly come from community residents, and the data set mainly includes four types of throwing actions: forward throwing, upward throwing, downward throwing and backward throwing.

[0039] For each pedestrian making a throwing action, a plurality of necessary key points are used to label each part of the human body including a head, limbs and a trunk part, and the key feature points can form a contour for each part of the human body.

[0040] In step 12, the processed data set in step 11 is taken as a training set, and a Body-LBF feature point alignment model is established and trained, and the model is used to detect and automatically label limbs, a trunk and a head of the pedestrian.

[0041] A local binary feature point alignment model (Body-LBF) based on a human body is used to learn a corresponding binary representation vector of the key feature points forming the contour of the human body part.

[0042] The structure of the Body-LBF is shown in Fig. 2 and includes a plurality of cascaded convolution layers, the convolution kernel size of the convolution layers is sequentially reduced, and a fully connected layer is connected after the last convolution layer, and the fully connected layer outputs a feature binary vector corresponding to a set key point.

[0043] Further, the training method of the Body-LBF feature point alignment model comprises the following steps:

[0044] In step 121, an image of a training data set is obtained, key feature points of a pedestrian in the image are labeled, and a local binary feature corresponding to the key feature points is obtained.

[0045] Specifically, in each image in the data set, the key feature points to be calibrated are learned by randomly selecting pixel points in the vicinity of the key feature points to obtain the local binary features corresponding to the key feature points.

[0046] Optionally, the vicinity can be a square candidate region centered on the key feature point, and the pixel points in the square candidate region centered on the key feature point are selected for residual learning.

[0047] In some embodiments, the extraction method of the local binary feature is as follows: the pixel value of the key feature point is compared with the pixel points in the candidate region, if the pixel value of the feature point is lower than the selected pixel point, the vector value is 0; if the pixel value of the feature point is higher than the selected pixel point, the vector value is 1; the binary vectors of all key feature points are integrated in sequence to form a vector composed of binary values, and the binary vectors of all key feature points are connected according to the front and back positions on the image to form a multi-dimensional binary vector for the human body part contour, i.e., a global feature representation of the human body part contour.

[0048] Step 122: taking all multi-dimensional binary vectors representing the human body part contour in each image to form a global feature representation of the human body part contour as the output, taking the original images in the training set as the input, and outputting to the Body-LBF feature point alignment model for training.

[0049] After calculating all multi-dimensional binary vectors representing the human body part contour in each image, the Body-LBF feature point alignment model can be trained by taking the throwing action images in the data set as the input and taking the global feature representation vector calibrated by the image as the output.

[0050] Step 123: calculating the positioning loss according to the actual output of the Body-LBF feature point alignment model and the global feature representation vector, adjusting the parameters of the Body-LBF feature point alignment model according to the loss value, and iteratively training until a set number of iterations is reached to obtain the trained Body-LBF feature point alignment model.

[0051] Through the above trained Body-LBF model, the head, limbs, and torso of the pedestrian are correctly calibrated by a plurality of necessary key points, and each key point can form a contour around each human body part.

[0052] Step 13: inputting the continuous frame images to be detected, calibrating the limbs, torso, and head of the human body in the image, and judging whether the pedestrian makes a throwing action according to the relative positions of the center points of each part of the human body.

[0053] The limbs, torso and head of the human body in the calibration image are calibrated, specifically: the trained Body-LBF model is input with the image to be detected, and outputs the integrated vector of the local binary feature corresponding to each key feature point, and each key feature point of the human body part in the output vector is accurately located according to the local binary feature of the key feature point.

[0054] The relative position of the center point of each part of the human body is determined to determine whether the pedestrian makes a throwing action, specifically, the relative position of the center point of the closed shape surrounded by the key points of each part of the human body is determined, and the relative position of the arm and the head and the torso is determined. If the relative position of each part of the human body satisfies the relative position of the throwing type, the throwing type can be corresponded, for example, the arm of the pedestrian is straight and located in front of the head and the torso, that is, the person makes a forward throwing action, so it can be determined that the thrown object is likely to be located in the front area of the pedestrian, and the thrown object in the front area of the pedestrian is detected.

[0055] In step 3, if the pedestrian makes a throwing action, the center point of the torso of the pedestrian is taken as the reference, the possible area of the object is determined according to the throwing action type of the pedestrian, and the significant object detection method is used to detect the thrown object according to the pixel context information.

[0056] The throwing action of the pedestrian and the thrown object are detected and tracked. Since the thrown object is generally a small target moving quickly, it is difficult to accurately capture in a single frame image. Optionally, the frame-by-frame matching method is used to track the pedestrian making the action and the thrown object moving quickly in step 3. The frame-by-frame matching method in this embodiment uses the Hungarian algorithm as the matching algorithm.

[0057] In order to make the recognition of the thrown object higher and easy to establish data association in the tracking process, the significant object detection method is used in this embodiment, specifically: the object thrown by the pedestrian is accurately detected through the color saliency of the object edge and the analysis of the pixel context information, and the object is tracked in real time for a short time.

[0058] According to the analysis result in step 3, if an object is detected in the determined area, it is considered that the pedestrian has thrown the object when making the throwing action, and the detection results of each part of the human body of the throwing pedestrian and the corresponding object are calibrated on the image.

[0059] Further, if the last frame contains the calibration result of the throwing action of the pedestrian, the Hungarian algorithm is used to match the calibration results of the current frame to track the throwing pedestrian and the object.

[0060] Hungarian assignment algorithm converts the matching problem of objects between frames into the shortest distance assignment problem. Since the target cannot perform long distance movement within a frame, the object closest to the target in the previous frame and having similar human body part pixel features in the next frame is the correct matching object of the target in the next frame. In addition, if the object does not change position after the set maximum stay time, it indicates that the object thrown by the pedestrian has landed, and the short-time tracking of the pedestrian and the thrown object ends.

[0061] In this embodiment, the method of human body part key point alignment and salient object recognition is used to detect and track the throwing action of the pedestrian and the thrown object respectively, and the analysis results from the pedestrian and the object are integrated to determine whether the pedestrian in the monitoring range actually makes a throwing behavior, thereby improving the detection accuracy. When applied to a community, if a pedestrian makes a throwing behavior in an inappropriate place, the corresponding alarm in the monitoring camera will immediately warn to remind the community security to stop the behavior of the resident.

[0062] Embodiment 2

[0063] Based on embodiment 1, this embodiment provides a throwing detection system based on human body local binary feature point alignment, comprising:

[0064] The throwing action recognition module is configured to acquire a continuous frame image to be detected, label the key feature points of the pedestrian in the image based on the local binary features of the pixel points, form the shape contours of the limbs, torso and head of the pedestrian according to the key feature points, and determine whether the pedestrian makes a throwing action and the type of the throwing action according to the relative positions of the center points of the parts of the human body;

[0065] The throwing area determination module is configured to determine the area where the thrown object is likely to appear based on the center point of the torso of the pedestrian as the reference when the throwing action of the pedestrian in the image is recognized according to the type of the throwing action of the pedestrian;

[0066] The detection determination module is configured to detect the throwing of the pedestrian by the salient object detection method based on the determined area according to the pixel context information, and determine that the pedestrian throws an object when making a throwing action if an object is detected in the determined area.

[0067] It should be noted that each module in this embodiment corresponds to each step in embodiment 1, and the specific implementation process is the same, which will not be repeated here.

[0068] This embodiment provides an electronic device, which includes a memory and a processor, and computer instructions stored in the memory and running on the processor. When the computer instructions are run by the processor, the steps of the method of embodiment 1 are completed.

[0069] Embodiment 4

[0070] The embodiment provides a computer readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the method in embodiment 1.

[0071] The electronic device proposed in the present disclosure can be a mobile terminal and a non-mobile terminal, the non-mobile terminal including a desktop computer, and the mobile terminal including a smart phone (such as an Android phone, an IOS phone, etc.), smart glasses, a smart watch, a smart bracelet, a tablet computer, a notebook computer, a personal digital assistant, and the like, which can perform wireless communication.

[0072] It should be understood that in the present disclosure, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSPs, application-specific integrated circuits ASICs, ready-to-program gate arrays FPGA or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0073] The memory can include read-only memory and random access memory, and provide instructions and data to the processor, and a portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.

[0074] In the implementation process, each step of the above method can be completed by integrated logic circuits of hardware in the processor or instructions in the form of software. The steps of the method disclosed in the present disclosure can be directly embodied as hardware processor execution completion, or executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, and the like. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here. Those skilled in the art can realize that the units of each example described in combination with the embodiments disclosed herein, i.e. the algorithm steps, can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software mode depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.

[0075] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0076] In several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, and the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0077] The functions if realized in the form of software function units and sold or used as independent products can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure essentially or say the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of software products, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present disclosure. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program code storage media.

[0078] The above only describes the preferred embodiments of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

[0079] The above describes the specific embodiments of the present disclosure in conjunction with the accompanying drawings, but is not intended to limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present disclosure without creative labor are still within the protection scope of the present disclosure.

Claims

1. A parabolic detection method based on the alignment of local binary feature points of the human body, characterized in that, Includes the following steps: The method involves acquiring consecutive frames of images to be detected, identifying key feature points of the pedestrian in the image based on local binary features of pixels, forming the shape contours of the pedestrian's limbs, torso, and head based on the key feature points, and determining whether the pedestrian has made a throwing motion and the type of throwing motion based on the relative positions of the center points of each part of the body. The method includes the following steps: Step 11: Obtain a dataset of community pedestrian throwing actions and mark the key feature points of the human body. The feature points can form the shape outline of the pedestrian's limbs, torso, and head. Step 12: Using the processed dataset from Step 11 as the training set, establish and train the Body-LBF feature point alignment model; The training method for the Body-LBF feature point alignment model includes the following steps: Acquire images from the training dataset, label the key feature points of pedestrians in the images, and obtain the local binary features corresponding to the key feature points; connect the binary vectors of all key feature points according to their positions in the image to form a multi-dimensional binary vector for the contour of the human body parts, that is, the global feature representation of the contour of the human body parts. The global feature representation of the human body part contour, which is composed of all multi-dimensional binary vectors representing the contour of human body parts in each image, is used as the output. The original images in the training set are used as the input, and the output is fed into the Body-LBF feature point alignment model for training. The localization loss is calculated based on the actual output of the Body-LBF feature point alignment model and the global feature representation vector. The parameters of the Body-LBF feature point alignment model are adjusted according to the loss value. The training is iterated until the set number of iterations is reached to obtain the trained Body-LBF feature point alignment model. Step 13: Input the continuous frame images to be detected, mark the limbs, torso and head of the human body in the images, and determine whether the pedestrian has made a throwing action based on the relative position of the center points of each part of the human body. The process of calibrating the limbs, torso, and head of the human body in an image is as follows: the trained Body-LBF model is input to the image to be detected, and the integrated vector of the local binary features corresponding to each key feature point is output. The points are then accurately located based on the local binary features of each key feature point of the human body in the output vector. When a pedestrian in an image is detected to be making a throwing motion, the area where the thrown object may appear is determined based on the center point of the pedestrian's torso and the type of throwing motion. Based on a defined region, a salient object detection method is used to detect objects thrown by pedestrians according to pixel context information. If an object is detected in the defined region, it is determined that the pedestrian threw the object when making the throwing action.

2. The parabolic detection method based on local binary feature point alignment of the human body as described in claim 1, characterized in that: After determining that the pedestrian threw an object when making the throwing motion, the process also includes the step of tracking the pedestrian and the thrown object: if the previous frame contains the calibration result of the pedestrian's throwing motion, the calibration result of the previous frame is matched with the calibration result of the current frame using the Hungarian algorithm to track the pedestrian and the thrown object.

3. The parabolic detection method based on local binary feature point alignment of the human body as described in claim 1, characterized in that: Throwing actions include forward throwing, upward throwing, downward throwing, and backward throwing; The various parts of the human body, including the head, limbs, and trunk, are marked.

4. The parabolic detection method based on local binary feature point alignment of the human body as described in claim 1, characterized in that: The method for obtaining a dataset of throwing actions by pedestrians in the community and identifying key feature points of the human body is as follows: By randomly selecting pixels in the vicinity of key feature points for residual learning, local binary features corresponding to key feature points are obtained.

5. The parabolic detection method based on local binary feature point alignment of the human body as described in claim 4, characterized in that: The area near a key feature point is a square candidate region of a set size centered on the key feature point.

6. The parabolic detection method based on local binary feature point alignment of the human body as described in claim 4, characterized in that, The method for extracting local binary features corresponding to key feature points is as follows: The pixel values ​​of key feature points are compared with the pixels in the candidate region. If the pixel value of a feature point is lower than that of the selected pixel, the vector value is 0; if the pixel value of a feature point is higher than that of the selected pixel, the vector value is 1. The obtained vector values ​​are integrated sequentially into a single binary vector. The binary vectors of all key feature points are then connected according to their positions in the image to form a multidimensional binary vector representing the contour of the human body, which is a global feature representation of the human body contour.

7. The parabolic detection method based on local binary feature point alignment of the human body as described in claim 1, characterized in that: The method of detecting objects thrown by pedestrians is adopted based on pixel context information, specifically by analyzing the color salience of the object edge and the pixel context information.

8. A parabolic detection system based on the human body local binary feature point alignment method according to claim 1, characterized in that, include: Throwing action recognition module: It is configured to acquire continuous frame images to be detected, identify key feature points of pedestrians in the image based on local binary features of pixels, form the shape outline of the pedestrian's limbs, torso and head based on key feature points, and determine whether the pedestrian has made a throwing action and the type of throwing action based on the relative position of the center points of each part of the human body. Throwing area determination module: It is configured to determine the possible area where the thrown object may appear based on the center point of the pedestrian's torso and the type of the pedestrian's throwing action when a pedestrian in the image is detected to make a throwing action. Detection and Judgment Module: Configured to detect objects thrown by pedestrians based on a defined region and pixel context information using a salient object detection method. If an object is detected in the defined region, it is determined that the pedestrian threw the object when making the throwing action.

9. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the steps of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the steps of any one of claims 1-7.

Citation Information

Patent Citations

  • Face alignment method

    CN106096560A

  • Detection method, system and device for kitchen garbage standard actions and medium

    CN112707058A