A human body tracking method and system
By using human skeleton detection and multi-dimensional information fusion technology in the recording and broadcasting system, the accuracy and robustness of human tracking and positioning in smart classrooms are solved, efficient and accurate tracking and shooting are achieved, and the use of equipment resources is reduced.
Patent Information
- Application Number
- CN202210009763.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-01-05
AI Technical Summary
In a flexible and free teaching environment, the existing recording and broadcasting system lacks the accuracy and robustness of human body tracking and positioning, making it difficult to meet the needs of smart classrooms.
By obtaining video frame images, using human skeleton detection algorithm to detect human skeleton data, combining bone characteristics, face characteristics, human behavior and color characteristics information, multi-dimensional information fusion is carried out to build a human body tracking model to achieve efficient tracking and shooting.
It improves the accuracy and robustness of the human tracking algorithm, reduces dependence on auxiliary cameras, and simplifies equipment resources and environment construction.
Smart Images

Figure CN114511589B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer vision technology, and particularly to a human body tracking method and system. Background Art
[0002] With the development of artificial intelligence technology and Internet technology, and with the requirements of "empowering teachers with artificial intelligence" and "empowering teaching with artificial intelligence" put forward, the "intelligent classroom" has changed the traditional teaching mode and learning environment, and also poses greater challenges to the tracking technology in the recording and broadcasting system.
[0003] The previous recording and broadcasting environment was based on the traditional teaching environment, and multiple auxiliary cameras were used to collect and locate video images for teacher location analysis and blackboard writing location analysis. And the image analysis and location can only simply rely on the motion characteristics of the target. For a more flexible and free teaching environment, there are great challenges in the accuracy and robustness of tracking and location. Summary of the Invention
[0004] Therefore, the embodiments of the present application provide a human body tracking method and system, which perform tracking based on different dimensional information such as human body bones, face features, human body behaviors, and color feature information, and efficiently assist in tracking and shooting.
[0005] In order to achieve the above object, the embodiments of the present application provide the following technical solutions:
[0006] According to the first aspect of the embodiments of the present application, a human body tracking method is provided, and the method includes:
[0007] Obtain a video frame image;
[0008] Use a human body bone detection algorithm to detect whether there is recognizable human body bone data in the video frame image;
[0009] If so, select a target tracking mode based on the video frame image to determine a human body target to be tracked according to a preset target start mode;
[0010] Perform bone key point detection based on the human body bone data to obtain bone feature information; when it is detected that the face faces the camera based on the bone feature information, perform face detection and face feature extraction to obtain face feature information; use a target behavior recognition model to analyze human body posture information to obtain target behavior information; use a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; perform motion information detection according to the bone position information and posture information in the bone feature information to obtain target motion information; and select a tracking number mode according to the recognizable human body bone data and its distance relationship with the tracking target to obtain a target tracking method;
[0011] Input the target tracking mode, skeletal feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method into the human body tracking model, and output the position information and behavior information of the tracked target human body.
[0012] Optionally, the method for detecting skeletal key points based on human body skeletal data to obtain skeletal feature information includes:
[0013] Based on the skeletal data of the head, neck, shoulders, arms, and waist in the human body skeletal data, as well as the head and neck direction data, shoulder width data, upper body height data, and the position relationship data between the head, shoulders, elbows, and hands, collect the data of the set key point regions in the human body skeleton.
[0014] Perform skeletal feature processing on the collected human body skeletal key point data to obtain skeletal feature information; the skeletal feature information includes human body motion direction information, human body position information, human body size information, and human body posture information.
[0015] Optionally, the method further includes:
[0016] If the detected human body posture information is raising the hand, add or switch to track the human body target.
[0017] Optionally, when it is detected based on the skeletal feature information that the face is facing the camera, perform face detection and face feature extraction to obtain face feature information, including:
[0018] Determine the face detection area according to the skeletal feature information, use the face detection model to perform face detection, and when it is detected that the face is facing the camera, use the face feature extraction model to perform face feature extraction to obtain face feature information.
[0019] Optionally, use the target behavior recognition model to analyze the human body posture information to obtain target behavior information, including:
[0020] Obtain the sample data of the key nodes and lengths and angles of the head, shoulders, neck, elbows, hands, and waist in the human body skeletal data;
[0021] Use the target behavior recognition model to detect the target behavior to obtain target behavior information; the target behavior information includes writing on the blackboard and explaining PPT; wherein, the target behavior recognition model is built by modeling the two target behaviors of writing on the blackboard and explaining PPT, distinguishing the positive and negative samples of the two target behaviors, and training the positive and negative samples based on the vector machine method to obtain the target behavior recognition model parameters.
[0022] Optionally, select the tracking number mode according to the recognizable human body skeletal data and its distance relationship with the tracked target to obtain the target tracking method, including:
[0023] Obtain the quantity of recognizable human bone data within a set tracking range around the tracked human target and its distance from the tracked human target;
[0024] Add tracking status information according to the quantity of recognizable human bone data and its distance relationship with the tracked human target, and determine the target tracking method; the target tracking method includes single human target tracking method, double human target tracking method, and multi-human target tracking method.
[0025] Optionally, the training process of the human tracking model includes:
[0026] Model the target tracking mode, bone feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method obtained from the video frame image. Classify the video frame image according to the model corresponding to each piece of information, respectively collect the sample data of different models, and use the support vector machine method to train the video sample data corresponding to each model to obtain the human tracking model parameters, so as to construct the human tracking model.
[0027] According to the second aspect of the embodiments of the present application, a human tracking system is provided, and the system includes:
[0028] A video image acquisition module, configured to acquire video frame images;
[0029] A bone detection module, configured to use a human bone detection algorithm to detect whether there is recognizable human bone data in the video frame image;
[0030] A tracking target determination module, configured to select a target tracking mode based on the video frame image to determine the tracked human target according to a preset target start mode;
[0031] A model call module, configured to perform bone key point detection based on human bone data to obtain bone feature information; when it detects that the face faces the camera based on the bone feature information, perform face detection and face feature extraction to obtain face feature information; use a target behavior recognition model to analyze human body posture information to obtain target behavior information; use a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; perform motion information detection according to the bone position information and posture information in the bone feature information to obtain target motion information; and select a tracking number mode according to the recognizable human bone data and its distance relationship with the tracking target to obtain the target tracking method;
[0032] A tracking target information acquisition module is configured to input the target tracking mode, skeletal feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method into a human body tracking model, and output the position information and behavior information of the tracked target human body.
[0033] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor runs the computer program, it is configured to implement the method described in the first aspect above.
[0034] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to implement the method described in the first aspect above.
[0035] In summary, the embodiments of the present application provide a human body tracking method and system. By acquiring video frame images; using a human body skeleton detection algorithm to detect whether there is recognizable human body skeleton data in the video frame images; if so, selecting a target tracking mode based on the video frame images to determine a tracked human body target according to a preset target start mode; performing skeleton key point detection based on the human body skeleton data to obtain skeletal feature information; and when a face is detected facing the camera based on the skeletal feature information, performing face detection and face feature extraction to obtain face feature information; and using a target behavior recognition model to analyze human body posture information to obtain target behavior information; and using a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; and detecting motion information according to the skeletal position information and posture information in the skeletal feature information to obtain target motion information; and selecting a tracking number mode according to the recognizable human body skeleton data and its distance relationship with the tracked target to obtain a target tracking method; inputting the target tracking mode, skeletal feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method into a human body tracking model, and outputting the position information and behavior information of the tracked target human body. Tracking is performed based on different dimensional information such as human body skeleton, face features, human body behavior, and color feature information, which efficiently assists tracking and shooting. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0037] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the implementation conditions of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the efficacy that the present invention can produce and the purpose that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0038] Figure 1 It is a schematic flowchart of a human body tracking method provided by an embodiment of the present application;
[0039] Figure 2 It is a human body bone key point diagram provided by an embodiment of the present application;
[0040] Figure 3 It is a detailed flowchart of target tracking provided by an embodiment of the present application;
[0041] Figure 4 It is a block diagram of a human body tracking system provided by an embodiment of the present application;
[0042] Figure 5 It shows a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0043] Figure 6 It shows a schematic diagram of a computer-readable storage medium provided by an embodiment of the present application. Specific embodiments
[0044] The following specific embodiments illustrate the implementation manners of the present invention. Those familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope protected by the present invention.
[0045] The present invention collects video data to analyze multi-dimensional information such as the bone features, color features, motion features, face features, and posture features of the target, and fuses the multi-dimensional information data to create a tracking and positioning model to achieve the tracking and positioning of the target in a complex environment. The multi-dimensional information fusion technology greatly improves the accuracy and robustness of the target tracking algorithm. Moreover, the target tracking algorithm of the present invention directly analyzes the image data collected by the tracking and shooting camera, without the need to add two auxiliary cameras for classroom positioning and blackboard writing positioning, which greatly saves device resources and simplifies the construction and initialization settings of the environment.
[0046] Figure 1 It shows a human body tracking method provided by an embodiment of the present application, and the method includes:
[0047] Step 101: Obtain video frame images;
[0048] Step 102: Use a human skeleton detection algorithm to detect whether there is recognizable human skeleton data in the video frame images;
[0049] Step 103: If there is, select a target tracking mode based on the video frame images to determine the tracked human target according to a preset target start mode;
[0050] Step 104: Perform skeleton key point detection based on the human skeleton data to obtain skeleton feature information; when it is detected that the face is facing the camera based on the skeleton feature information, perform face detection and face feature extraction to obtain face feature information; use a target behavior recognition model to analyze the human body posture information to obtain target behavior information; use a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; perform motion information detection according to the skeleton position information and posture information in the skeleton feature information to obtain target motion information; select a tracking number mode according to the recognizable human skeleton data and its distance relationship with the tracked target to obtain a target tracking method;
[0051] Step 105: Input the target tracking mode, skeleton feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method into a human body tracking model, and output the position information and behavior information of the tracked target human body.
[0052] In a possible implementation manner, in step 104, the performing skeleton key point detection based on the human skeleton data to obtain skeleton feature information includes:
[0053] Collect data of a set key point area in the human skeleton based on the skeleton data of the head, neck, shoulders, arms, and waist parts in the human skeleton data, as well as the head and neck direction data, shoulder width data, upper body height data, and position relationship data between the head, shoulders, elbows, and hands; perform skeleton feature processing on the collected human skeleton key point data to obtain skeleton feature information; the skeleton feature information includes human body motion direction information, human body position information, human body size information, and human body posture information.
[0054] In a possible implementation manner, in step 104, the method further includes: if the detected human body posture information is raising a hand, add or switch the tracked human target.
[0055] In a possible implementation manner, in step 104, when it is detected that the face is facing the camera based on the skeleton feature information, performing face detection and face feature extraction to obtain face feature information includes:
[0056] Determine the face detection area based on the bone feature information, perform face detection using a face detection model. When it is detected that the face is facing the camera directly, use a face feature extraction model to extract face feature information, obtaining the face feature information.
[0057] In a possible implementation manner, in step 104, use a target behavior recognition model to analyze the human body posture information, obtaining the target behavior information, including:
[0058] Obtain the sample data of the key nodes (head, shoulders, neck, elbows, hands, waist) and lengths and angles in the human body bone data; use the target behavior recognition model to detect the target behavior, obtaining the target behavior information; the target behavior information includes writing on the blackboard and explaining PPT; wherein, the target behavior recognition model is built by modeling the two target behaviors of writing on the blackboard and explaining PPT, differentiating the positive and negative samples of the two target behaviors, and training the positive and negative samples based on the support vector machine method to obtain the target behavior recognition model parameters, so as to construct the target behavior recognition model.
[0059] In a possible implementation manner, in step 104, select the tracking number mode according to the recognizable human body bone data and its distance relationship with the tracking target, obtaining the target tracking method, including:
[0060] Obtain the number of recognizable human body bone data within the set tracking range around the tracked human body target and its distance from the tracked human body target; according to the number of recognizable human body bone data and its distance relationship with the tracked human body target, add tracking status information to determine the target tracking method; the target tracking methods include single human body target tracking method, double human body target tracking method, and multi-human body target tracking method.
[0061] In a possible implementation manner, in step 105, the training process of the human body tracking model includes:
[0062] Model the target tracking mode, bone feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method obtained from the video frame images, classify the video frame images according to the models corresponding to each information, respectively collect the sample data of different models, and use the support vector machine method to train the video sample data corresponding to each model to obtain the human body tracking model parameters, so as to construct the human body tracking model.
[0063] The following further explains the human body tracking method and system disclosed in the embodiments of the present application in conjunction with the accompanying drawings, including:
[0064] The human body tracking system in the embodiments of the present application further includes a preprocessing module, which is used to complete the initialization of the bone key point detection classifier, face detection classifier, face feature extraction model, and the initialization of the relevant data flow.
[0065] A data acquisition module, which is used to acquire human bone data; considering the code efficiency issue, only the necessary bone information, namely the bone data of the head, neck, shoulders, arms, and waist (see Figure 2 the key points 1, 2, 3, 4, 5, 6, 7, 8, 9, 12 in the 18-joint point diagram of the human body), as well as the head and neck direction data, shoulder width data, upper body height data, position relationship data between the shoulders, elbows, and hands, etc.
[0066] Furthermore, a bone detection model is used to detect bone key points. Among the human bone key points, the key point data of 1 human head, 2 necks, 3 right shoulders, 4 right elbows, 5 right hands, 6 left shoulders, 7 left elbows, 8 left hands, 9 right waists, and 12 left waists are selected for acquisition; and the acquired key point data is further processed to obtain bone feature data information: human movement direction data, human target position data, human size data, and human posture data.
[0067] At the same time, face feature data is also acquired: Similarly considering the code execution efficiency issue, the face feature acquisition time is only selected when the target human body faces the camera directly during the process of tracking the target. The face detection area only detects the area above the shoulders and neck of the human body. A face detection model is used to detect the face, and a face feature extraction model is used to extract face feature data; the face detection area is determined by the position data and posture data in the bone feature data, then face detection is performed, and finally a face feature extraction model is used to extract face feature data, or face feature data.
[0068] At the same time, color feature data is also acquired: Considering the code execution efficiency and the reliability of the color area, the target color range randomly acquires the color data of several points in the middle area of the upper body of the human target. The color acquisition area is determined by the position data and bone data in the bone feature data: that is, the core area of the upper body of the human body, and then several points' grayscale values are randomly acquired within the determined area.
[0069] After acquiring various data, the human body is tracked. Specifically, it can include the following:
[0070] First, it is judged whether to manually start tracking the target. If it is set to manually start tracking the target, the human body closest to the manually circled area is selected as the tracking target. Otherwise, the automatic target selection start mode is enabled, and at the same time, the gestures of each detected person are detected. If the agreed start target gesture is detected, the tracking target is restarted. During the tracking process, the human body gesture is detected at all times. When the agreed gesture is recognized, the tracking target can be switched at any time.
[0071] In the selection of the target tracking mode, the selection of the tracking target can be divided into three modes: automatic selection mode, manual selection mode, and semi-automatic selection mode. Specifically, it includes:
[0072] Automatic selection method: This method is the default target startup method. If there is only one person in the camera view, that person is started as the tracking target. If there are multiple people, the target that is closer to the camera and has its front facing the camera is selected as the default tracking target.
[0073] Manual selection method: This method requires manually selecting the tracking target by manually enclosing the target area.
[0074] Semi-automatic selection method: This method selects the tracking target by performing gesture recognition on the human body in the camera view.
[0075] In the selection of the target tracking method, the algorithm determines which tracking module to enter next based on the number of detected human skeletons. According to the number of human targets within the target tracking range, it is divided into a single human target tracking module, a double human target tracking module, and a multi-human target tracking module. Specifically, it includes:
[0076] Single human target tracking method: During the target tracking process, if there is only one human skeleton data within the tracking range, the tracking will be divided into two cases based on the tracking status information of the human target: whether the human target is in the overlapping process, and add the tracking status information, and then create specific tracking models according to the cases.
[0077] Double human target tracking method: During the target tracking process, if two human skeleton data appear within the target monitoring range, the tracking will add tracking status information based on the distance relationship between the two human skeleton information, and create specific tracking models.
[0078] Multi-human target tracking method: During the tracking process, if multiple human skeleton data appear within the target detection range (generally, at most three human targets appear within the detection range), the tracking will add tracking status information based on the distance relationship between the target and the multiple human targets, and create specific tracking models.
[0079] Create a human tracking model: In different tracking modules, determine the specific tracking model based on the tracking target human status information and the status information of relevant human targets. The tracking model combines the tracking status information of the tracking target and the human target, and fuses multi-dimensional information such as its face feature data, bone feature data, motion feature data, and color feature data to select the tracking target for the next frame, thereby realizing the tracking function.
[0080] In different tracking modules, the specific tracking model fuses multi-dimensional information such as the face feature information, bone feature information, motion feature information, and color feature information of the tracked target human body and the human body target in the video image for simple modeling. According to the characteristics of each tracking model, the video image is generally classified, and then the sample data of different models are manually collected respectively. The video sample data corresponding to each tracking model is trained by the method based on the support vector machine to obtain the parameters of each tracking model, and a large number of field tests are carried out to optimize the parameters of each model.
[0081] Create a target behavior recognition model: Collect the bone data collected by the data acquisition module, including sample data such as the lengths and angles of the key nodes of the shoulders, necks, elbows, hands, and waists and related key parts. Then, simply model the behaviors of writing on the blackboard and explaining PPT. Then, manually and precisely distinguish the positive and negative samples of the two behaviors, and train the relevant data of the positive and negative samples of the two behavior data by the method based on the support vector machine to obtain the model parameters for detecting the two behaviors. Continue to collect more samples for testing and optimizing the model parameters to complete the behavior recognition model.
[0082] Steps of target behavior recognition: Perform human body behavior recognition based on the bone data of the target human body (mainly using the key point data of the head, shoulders, neck, elbows, hands, and waist), which are mainly divided into two behaviors: writing on the blackboard and explaining PPT.
[0083] Figure 3 The detailed flowchart of the human body tracking method embodiment provided according to the embodiment of the present application is shown, including:
[0084] Establish a processing flow and initialize the bone detector, face detector, face feature extraction model, and startup mode of the tracking target, etc.; Use the human body bone detection algorithm to detect the human body bone information in the video image. If there is no human body bone data, take the next frame of the image. If there is, determine the tracking target according to the target startup mode and continue with the subsequent steps;
[0085] Analyze the detected human body bone data: Extract the bone key point information required by the algorithm for advanced data extraction to obtain the position information of 18 bone key points. And, according to the analysis of each bone data, analyze the gesture information of the human body. Select the gesture points 345678 among the key points. If it is detected that the hand is raised, switch the tracking target. The hand-raising model is used to detect the hand-raising information to determine whether to switch the tracking target again.
[0086] Based on the bone feature information, when the human body faces the camera frontally and face-on, perform face detection in the head and shoulder area, and then use the face feature extraction model to extract face features to obtain face feature information.
[0087] Analyze the posture information of the human body based on human bone data and input it into the behavior recognition model to obtain the target behavior information.
[0088] Determine the color information acquisition range according to the position information and size information of the human body analyzed from the human bone data, that is, randomly collect a number of points (such as 1000) and corresponding gray-scale values and other data within the key point areas 3, 6, 9, and 12 of the core part of the upper body of the human body as the color feature data of the human body.
[0089] Obtain the motion information of the human body, such as the motion direction and motion speed, according to the bone position information and posture information of each human body, such as moving left or right and the step length of the movement.
[0090] Determine the tracking method according to the number of bones extracted and their distance relationship with the tracking target: single human body tracking method, double human body tracking method, and multi-human body tracking method.
[0091] Furthermore, according to different tracking methods, select a pre-trained human body tracking model, fuse the information obtained above, such as multi-dimensional information such as human body color information, motion information, behavior information, and face information, input the multi-dimensional information into the corresponding tracking model, and output the position information and behavior information of the tracking target.
[0092] The human body tracking method provided by the embodiments of this application is a tracking algorithm based on the fusion of multi-dimensional information such as bones, colors, motions, and faces, which greatly improves the target tracking accuracy and robustness. The recognition of human behaviors enriches the later image-giving strategy, changes the previous method of using auxiliary cameras for analysis and positioning, saves resources, and greatly simplifies the construction and setting of the environment, providing strong tracking technical support for the new intelligent education recording and broadcasting technology.
[0093] In summary, the embodiment of the present application provides a human body tracking method, which includes obtaining a video frame image; using a human body bone detection algorithm to detect whether there is recognizable human body bone data in the video frame image; if so, selecting a target start mode preset in the target tracking mode based on the video frame image to determine the human body target to be tracked; performing bone key point detection based on the human body bone data to obtain bone feature information; when it is detected that the face faces the camera based on the bone feature information, performing face feature extraction to obtain face feature information; using a target behavior recognition model to analyze the human body posture information to obtain target behavior information; using a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; detecting motion information according to the bone position information and posture information in the bone feature information to obtain target motion information; selecting a tracking number mode according to the recognizable human body bone data and the distance to obtain a target tracking method; inputting the target tracking mode, bone feature information, face feature information, target behavior information, human body color feature information, target motion information and target tracking method into a human body tracking model, and outputting the position information and behavior information of the tracked target human body. Tracking is performed based on different dimensional information such as human body bones, face features, human body behaviors, and color feature information, which efficiently assists tracking and shooting.
[0094] Based on the same technical concept, the embodiment of the present application also provides a human body tracking system, as Figure 4 shown, the system includes:
[0095] A video image acquisition module 401, configured to acquire a video frame image;
[0096] A bone detection module 402, configured to use a human body bone detection algorithm to detect whether there is recognizable human body bone data in the video frame image;
[0097] A tracking target determination module 403, configured to select a target tracking mode based on the video frame image to determine the human body target to be tracked according to a preset target start mode;
[0098] A model calling module 404, configured to perform bone key point detection based on the human body bone data to obtain bone feature information; when it is detected that the face faces the camera based on the bone feature information, perform face detection and face feature extraction to obtain face feature information; use a target behavior recognition model to analyze the human body posture information to obtain target behavior information; use a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; detect motion information according to the bone position information and posture information in the bone feature information to obtain target motion information; select a tracking number mode according to the recognizable human body bone data and its distance relationship with the tracking target to obtain a target tracking method;
[0099] The target tracking information acquisition module 405 is configured to input the target tracking mode, skeletal feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method into the human body tracking model, and output the position information and behavior information of the tracked target human body.
[0100] The embodiment of the present application also provides an electronic device corresponding to the method provided in the foregoing embodiment. Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the method provided in any of the foregoing embodiments of the present application.
[0101] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one physical port 203 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0102] The bus 202 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store the program, and after the processor 200 receives the execution instruction, it executes the program. The method disclosed in any of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.
[0103] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.
[0104] The electronic device provided in the embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0105] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method provided in the foregoing embodiments. Please refer to Figure 6 , the computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method provided in any of the foregoing embodiments.
[0106] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.
[0107] The computer-readable storage medium provided in the above embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored in it.
[0108] It should be noted that:
[0109] The algorithms and displays provided herein are not inherently related to any particular computer, virtual apparatus, or other device. A variety of general-purpose apparatuses may also be used in conjunction with the teachings presented herein. The structure required to construct such apparatuses will be apparent from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present application.
[0110] In the specification provided herein, a number of specific details are set forth. However, it is understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0111] Similarly, it should be understood that in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0112] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except for the fact that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0113] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0114] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0115] It should be noted that the above embodiments illustrate rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0116] As described above, only the preferred specific embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A human body tracking method, characterized in that, The method includes: Obtain video frame images; Use a human skeleton detection algorithm to detect whether there is recognizable human skeleton data in the video frame images; If so, select a target tracking mode based on the video frame images to determine a tracked human target according to a preset target startup mode; Perform skeleton key point detection based on the human skeleton data to obtain skeleton feature information; when it is detected that the face is facing the camera based on the skeleton feature information, perform face detection and face feature extraction to obtain face feature information; use a target behavior recognition model to analyze human body pose information to obtain target behavior information; use a human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; perform motion information detection according to the skeleton position information and pose information in the skeleton feature information to obtain target motion information; select a tracking number mode according to the recognizable human skeleton data and its distance relationship with the tracking target to obtain a target tracking method; Input the target tracking mode, the skeleton feature information, the face feature information, the target behavior information, the human body color feature information, the target motion information, and the target tracking method into a human body tracking model, and output the position information and behavior information of the tracked target human body; Among them, selecting a tracking number mode according to the recognizable human skeleton data and its distance relationship with the tracking target to obtain the target tracking method includes: obtaining the number of recognizable human skeleton data and its distance from the tracked human target within a set tracking range around the tracked human target; adding tracking status information according to the number of recognizable human skeleton data and its distance relationship with the tracked human target to determine the target tracking method; the target tracking method includes a single human target tracking method, a dual human target tracking method, and a multi-human target tracking method.
2. The method according to claim 1, wherein Performing skeleton key point detection based on the human skeleton data to obtain skeleton feature information includes: Based on the skeleton data of the head, neck, shoulders, arms, and waist in the human skeleton data, as well as the head and neck direction data, shoulder width data, upper body height data, and position relationship data between the head, shoulders, elbows, and hands, collect data of a set key point area in the human skeleton; Perform skeleton feature processing on the collected human skeleton key point data to obtain the skeleton feature information; the skeleton feature information includes human body motion direction information, human body position information, human body size information, and human body pose information.
3. The method according to claim 2, characterized in that, The method further includes: If it is detected that the human body pose information is raising a hand, add or switch the tracked human target.
4. The method according to claim 1, wherein When it is detected that the face is facing the camera based on the skeleton feature information, performing face detection and face feature extraction to obtain face feature information includes: Determine a face detection area according to the skeleton feature information, use a face detection model to perform face detection, and when it is detected that the face is facing the camera, use a face feature extraction model to perform face feature extraction to obtain the face feature information.
5. The method according to claim 1, characterized in that Using a target behavior recognition model to analyze human body pose information to obtain target behavior information includes: Obtain the sample data of the key nodes of the head, shoulders, neck, elbows, hands, and waist, as well as the lengths and angles in the human body bone data; Use the target behavior recognition model to detect the target behavior and obtain target behavior information; the target behavior information includes writing on the blackboard and explaining PPT; wherein, the target behavior recognition model is constructed by modeling the two target behaviors of writing on the blackboard and explaining PPT, distinguishing the positive and negative samples of the two target behaviors, and training the positive and negative samples based on the support vector machine method to obtain the target behavior recognition model parameters.
6. The method according to claim 1, characterized in that, The training process of the human body tracking model includes: Model the target tracking mode, bone feature information, face feature information, target behavior information, human body color feature information, target motion information, and target tracking method obtained from the video frame image. Classify the video frame image according to the model corresponding to each information, respectively collect the sample data of different models, and use the support vector machine method to train the video sample data corresponding to each model to obtain the human body tracking model parameters, so as to construct the human body tracking model.
7. A human body tracking system, characterized in that, The system includes: A video image acquisition module for acquiring video frame images; A bone detection module for using a human body bone detection algorithm to detect whether there is recognizable human body bone data in the video frame image; A tracking target determination module for selecting a target tracking mode based on the video frame image to determine the tracking human body target according to a preset target start mode; A model calling module for performing bone key point detection based on the human body bone data to obtain bone feature information; and when the face is detected facing the camera based on the bone feature information, performing face detection and face feature extraction to obtain face feature information; and using the target behavior recognition model to analyze the human body posture information to obtain target behavior information; and using the human body color feature model to sample random points and gray values in the upper body area of the human body to obtain human body color feature information; and detecting motion information according to the bone position information and posture information in the bone feature information to obtain target motion information; and selecting a tracking number mode according to the recognizable human body bone data and its distance relationship with the tracking target to obtain a target tracking method; wherein, selecting a tracking number mode according to the recognizable human body bone data and its distance relationship with the tracking target to obtain a target tracking method includes: obtaining the number of recognizable human body bone data in a set tracking range around the tracking human body target and its distance from the tracking human body target; adding tracking status information according to the number of recognizable human body bone data and its distance relationship with the tracking human body target to determine the target tracking method; the target tracking method includes a single human body target tracking method, a double human body target tracking method, and a multi-human body target tracking method; A tracking target information acquisition module, configured to input the target tracking mode, the skeletal feature information, the face feature information, the target behavior information, the human body color feature information, the target motion information, and the target tracking method into a human body tracking model, and output the position information and behavior information of the tracked target human body.
8. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor runs the computer program, it is configured to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer-readable instruction is stored thereon, and the computer-readable instruction can be executed by a processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Classroom behavior detection method and electronic equipment
CN110781843A
Human body posture recognition method and device based on skeleton key points, storage medium and terminal
CN111680562A
Information associated analysis method and apparatus, and storage medium and electronic device
WO2020114138A1