Head movement detection method, device, equipment and storage medium

By using lightweight face detection and posture detection models on devices with limited computing power, effective detection of head movement is achieved, the problem of high computing power consumption in the existing technology is solved, and the application scenarios are expanded.

CN115497131BActive Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210948677.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-18
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

The existing head motion detection methods require a lot of computing power and cannot be effectively tested on devices with limited computing power.

Method used

The lightweight face detection model and face posture detection model are adopted to collect images through the camera device, and the lightweight neural network is used to perform face area detection and head deflection angle calculation, and the deflection angle is compared with the standard angle to determine the effectiveness of the motion.

Benefits of technology

Head motion detection is implemented on devices with limited computing power, reducing the computing power requirements required for detection and can be deployed in more application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497131B_ABST
    Figure CN115497131B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and discloses a head movement detection method, device, equipment and storage medium. The method includes: when a movement detection request is received, calling a pre-set camera device to collect the current person image of the target user; inputting the current person image of the target user into a pre-set face detection model to detect a face region image, the face detection model being a lightweight neural network model; inputting the face region image into a pre-set face pose detection model for regression calculation to obtain the target deflection angle of the head of the target user in the target direction, the face pose detection model being a lightweight neural network model; comparing the target deflection angle with a pre-set standard deflection angle, and if the target deflection angle reaches the standard deflection angle, determining that the current head movement is effective. The present invention performs head movement detection based on a lightweight detection model, and requires less computing power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method, device, equipment and storage medium for head movement detection. Background Art

[0002] Working in the same posture for a long time will cause cervical problems in the future. Therefore, office workers need to perform appropriate and reasonable head movements to reduce health problems such as cervical spondylosis. How to detect and monitor head movements is an important means to ensure the effectiveness of the movements.

[0003] In the existing technology, head movements are usually detected and monitored by a camera in cooperation with a server with relatively high computing power. However, head movement detection cannot be performed when the computing power of the device is limited. Summary of the Invention

[0004] The main object of the present invention is to solve the problem that the existing head movement detection method requires a large amount of computing power.

[0005] The first aspect of the present invention provides a method for head movement detection, including:

[0006] When a movement detection request sent by a terminal is received, a preset camera device in the terminal is called to collect the current person image of the target user;

[0007] The current person image of the target user is input into a preset face detection model for detection to obtain the current face region image of the target user, where the face detection model is a lightweight neural network model;

[0008] The current face region image of the target user is input into a preset face pose detection model for regression calculation to obtain the target deflection angle of the head of the target user in the target direction, where the face pose detection model is a lightweight neural network model;

[0009] The target deflection angle is compared with a preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective.

[0010] Optionally, in the first implementation manner of the first aspect of the present invention, after the target deflection angle is compared with the preset standard deflection angle, and if the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective, the method further includes:

[0011] Taking the intersection point between the horizontal axis and the vertical axis of the face region image as the origin of the coordinate system, taking the horizontal direction as the horizontal axis direction of the coordinate system, and taking the vertical direction as the vertical axis direction of the coordinate system, a plane rectangular coordinate system is constructed;

[0012] Calculate the length of the face region image in the target direction;

[0013] Perform trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the target coordinate position of the target user's head in the plane rectangular coordinate system.

[0014] Optionally, in the second implementation manner of the first aspect of the present invention, after performing trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the target coordinate position of the target user's head in the plane rectangular coordinate system, it further includes:

[0015] Determine the stable coordinate interval of the target user's head in the plane rectangular coordinate system based on the preset stable demand information;

[0016] If the target deflection angle is greater than the preset abnormal deflection threshold or the target coordinate position is not within the stable coordinate interval, it is determined that the current shooting environment is unstable, and stability correction is performed on the target coordinate position.

[0017] Optionally, in the third implementation manner of the first aspect of the present invention, the inputting the current person image of the target user into the preset face detection model for detection to obtain the current face region image of the target user includes:

[0018] Input the current person image of the target user into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network;

[0019] Call the face feature extraction network to extract face features of at least one feature size from the current person image of the target user;

[0020] Call the face feature recognition network to identify alternative face regions in the face features of each feature size according to the preset number of prior boxes corresponding to each feature size;

[0021] Call the face feature screening network to screen each of the alternative face regions to obtain the current face region image of the target user.

[0022] Optionally, in the fourth implementation manner of the first aspect of the present invention, before inputting the current face region image of the target user into the preset face pose detection model for regression calculation to obtain the target deflection angle of the target user's head in the target direction, it further includes:

[0023] Construct a face pose image set for model training, and label the deflection angles of each image in the face pose image set in the target direction;

[0024] According to the preset image division ratio, divide the face pose image set into a face pose training image set and a face pose verification image set;

[0025] Based on each image in the face pose training image set, perform regression training on the preset Mobile-Net convolutional neural network model for the deflection angle;

[0026] Based on each image in the face pose verification image set, perform model verification on the Mobile-Net convolutional neural network model after regression training to obtain the face pose detection model.

[0027] Optionally, in the fifth implementation manner of the first aspect of the present invention, the Mobile-Net convolutional neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The performing regression training on the preset Mobile-Net convolutional neural network model for the deflection angle based on each image in the face pose training image set includes:

[0028] Invoke the pose feature extraction network to extract the target face pose feature from the target training image in the face pose training image set;

[0029] Invoke the deflection angle classification network, and calculate the multi-class deflection interval probability distribution corresponding to the target face pose feature according to the target face pose feature and the preset deflection interval probability matrix;

[0030] Select the deflection interval with the largest probability value from the multi-class deflection interval probability distribution as the target deflection interval corresponding to the target face pose feature, and obtain the regression function corresponding to the target deflection interval;

[0031] Invoke the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose feature and the regression function, and obtain the deflection angle of the target training image in the target direction;

[0032] Based on the preset loss function and the annotation information of the deflection angle of the target training image in the target direction, calculate the loss value corresponding to the deflection angle of the target training image in the target direction;

[0033] Perform stochastic gradient descent on the network parameters of the Mobile-Net convolutional neural network model according to the loss value, and calculate the loss value again until the loss value is less than a preset threshold, then determine that the Mobile-Net convolutional neural network model converges and end the regression training.

[0034] Optionally, in the sixth implementation manner of the first aspect of the present invention, after comparing the target deflection angle with a preset standard deflection angle, if the target deflection angle reaches the standard deflection angle and it is determined that the current head movement is effective, the following steps are further included:

[0035] Count the number of effective movements of the target user within a preset time period;

[0036] If the number of effective movements of the target user within the preset time period is greater than or equal to a preset threshold, determine that the current movement of the target user is qualified and execute a preset positive feedback strategy for movement;

[0037] If the number of effective movements of the target user within the preset time period is less than the preset threshold, determine that the current movement of the target user is unqualified and execute a preset negative feedback strategy for movement.

[0038] The second aspect of the present invention provides a head movement detection device, including:

[0039] An image acquisition module, configured to call a preset camera device in the terminal to acquire a current person image of the target user when receiving a movement detection request sent by the terminal;

[0040] A face detection module, configured to input the current person image of the target user into a preset face detection model for detection to obtain a current face region image of the target user, where the face detection model is a lightweight neural network model;

[0041] A deflection calculation module, configured to input the current face region image of the target user into a preset face pose detection model for regression calculation to obtain a target deflection angle of the head of the target user in the target direction, where the face pose detection model is a lightweight neural network model;

[0042] An effective detection module, configured to compare the target deflection angle with a preset standard deflection angle, and if the target deflection angle reaches the standard deflection angle, determine that the current head movement is effective.

[0043] Optionally, in the first implementation manner of the second aspect of the present invention, the head movement detection device further includes a head position calculation module, and the head position calculation module specifically includes:

[0044] A coordinate system construction unit for constructing a plane rectangular coordinate system with the intersection point between the horizontal central axis and the vertical central axis of the face region image as the origin of the coordinate system, the horizontal direction as the horizontal axis direction of the coordinate system, and the vertical direction as the vertical axis direction of the coordinate system;

[0045] A length calculation unit for calculating the length of the face region image in the target direction;

[0046] A coordinate calculation unit for performing trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the target coordinate position of the target user's head in the plane rectangular coordinate system.

[0047] Optionally, in the second implementation manner of the second aspect of the present invention, the head position calculation module specifically includes:

[0048] A coordinate system construction unit for constructing a plane rectangular coordinate system with the intersection point between the horizontal central axis and the vertical central axis of the face region image as the origin of the coordinate system, the horizontal direction as the horizontal axis direction of the coordinate system, and the vertical direction as the vertical axis direction of the coordinate system;

[0049] A length calculation unit for calculating the length of the face region image in the target direction;

[0050] A coordinate calculation unit for performing trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the target coordinate position of the target user's head in the plane rectangular coordinate system.

[0051] A stability correction unit for determining the stable coordinate interval of the target user's head in the plane rectangular coordinate system based on the preset stability requirement information; if the target deflection angle is greater than the preset abnormal deflection threshold or the target coordinate position is not within the stable coordinate interval, it is determined that the current shooting environment is unstable, and stability correction is performed on the target coordinate position.

[0052] Optionally, in the third implementation manner of the second aspect of the present invention, the face detection module specifically includes:

[0053] An image input unit for inputting the current person image of the target user into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network;

[0054] A face feature extraction unit, configured to call the face feature extraction network to extract face features of at least one feature size from the current person image of the target user;

[0055] An alternative face recognition unit, configured to call the face feature recognition network to recognize alternative face regions in the face features of each feature size according to a preset number of prior boxes corresponding to each feature size;

[0056] A face screening unit, configured to call the face feature screening network to screen each of the alternative face regions to obtain the current face region image of the target user.

[0057] Optionally, in the fourth implementation manner of the second aspect of the present invention, the head movement detection device further includes a face pose detection model construction module, and the face pose detection model construction module specifically includes:

[0058] An image set construction unit, configured to construct a face pose image set for model training and label the deflection angles of each image in the face pose image set in the target direction;

[0059] An image set division unit, configured to divide the face pose image set into a face pose training image set and a face pose verification image set according to a preset image division ratio;

[0060] A regression training unit, configured to perform regression training of the deflection angle on a preset Mobile-Net convolutional neural network model based on each image in the face pose training image set;

[0061] A model verification unit, configured to perform model verification on the Mobile-Net convolutional neural network model after regression training based on each image in the face pose verification image set to obtain the face pose detection model.

[0062] Optionally, in the fifth implementation manner of the second aspect of the present invention, the Mobile-Net convolutional neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network, and the regression training unit is specifically configured to:

[0063] Invoke the pose feature extraction network to extract the target face pose feature from the target training image in the face pose training image set; invoke the yaw angle classification network, and calculate the multi-class yaw angle interval probability distribution corresponding to the target face pose feature according to the target face pose feature and the preset yaw angle interval probability matrix; select the yaw angle interval with the largest probability value from the multi-class yaw angle interval probability distribution as the target yaw angle interval corresponding to the target face pose feature, and obtain the regression function corresponding to the target yaw angle interval; invoke the yaw angle regression network, and perform yaw angle regression calculation according to the target face pose feature and the regression function to obtain the yaw angle of the target training image in the target direction; based on the preset loss function and the annotation information of the yaw angle of the target training image in the target direction, calculate the loss value corresponding to the yaw angle of the target training image in the target direction; perform stochastic gradient descent on the network parameters of the Mobile-Net convolutional neural network model according to the loss value, and calculate the loss value again until the loss value is less than the preset threshold, determine that the Mobile-Net convolutional neural network model converges, and end the regression training.

[0064] Optionally, in the sixth implementation manner of the second aspect of the present invention, the head movement detection device further includes a movement qualification detection module, and the movement qualification detection module is specifically configured to:

[0065] Count the number of valid movements of the target user within a preset time period; if the number of valid movements of the target user within the preset time period is greater than or equal to the preset threshold, determine that the target user's current movement is qualified, and execute the preset movement positive feedback strategy; if the number of valid movements of the target user within the preset time period is less than the preset threshold, determine that the target user's current movement is unqualified, and execute the preset movement negative feedback strategy.

[0066] The third aspect of the present invention provides a head movement detection device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to enable the head movement detection device to execute the steps of the above head movement detection method.

[0067] The fourth aspect of the present invention provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is enabled to execute the steps of the above head movement detection method.

[0068] In the technical solution provided by the present invention, when a motion detection request sent by a terminal is received, a pre-installed camera device in the terminal is called to collect the current person image of the target user; the current person image of the target user is input into a pre-installed lightweight face detection model for detection to obtain the current face region image of the target user; the current face region image of the target user is input into a pre-installed lightweight face pose detection model for regression calculation to obtain the target deflection angle of the head of the target user in the target direction. Finally, the target deflection angle is compared with a pre-set standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective. Based on the lightweight face detection model and face pose detection model, the present invention performs head movement detection on terminal devices with limited computing power, reducing the computing power requirements for detection, and thus can be deployed in more application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 Schematic diagram of the first embodiment of the head movement detection method in the embodiment of the present invention;

[0070] Figure 2 Schematic diagram of the second embodiment of the head movement detection method in the embodiment of the present invention;

[0071] Figure 3 Schematic diagram of the third embodiment of the head movement detection method in the embodiment of the present invention;

[0072] Figure 4 Schematic diagram of the fourth embodiment of the head movement detection method in the embodiment of the present invention;

[0073] Figure 5 Schematic diagram of an embodiment of the head movement detection device in the embodiment of the present invention;

[0074] Figure 6 Schematic diagram of another embodiment of the head movement detection device in the embodiment of the present invention;

[0075] Figure 7 Schematic diagram of an embodiment of the head movement detection device in the embodiment of the present invention;

[0076] Figure 8 Schematic diagram of an embodiment of establishing a two-dimensional plane rectangular coordinate system in the embodiment of the present invention;

[0077] Figure 9 Schematic diagram of the first triangular geometric relationship in the face region image in the embodiment of the present invention;

[0078] Figure 10 Schematic diagram of the second triangular geometric relationship in the face region image in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0079] An embodiment of the present invention provides a head movement detection method, device, equipment and storage medium, which consume less computing power.

[0080] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0081] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, sense the environment, acquire knowledge and use knowledge to obtain the best results of theory, method, technology and application system.

[0082] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0083] It can be understood that the execution subject of the present invention can be a head movement detection device, or a terminal or a server, and specific limitations are not made here. An embodiment of the present invention is described by taking the server as the execution subject as an example.

[0084] For ease of understanding, the specific process of the embodiment of the present invention is described below. Please refer to Figure 1 The first embodiment of the head movement detection method in the embodiment of the present invention includes:

[0085] 101. When receiving a movement detection request sent by a terminal, call a preset camera device in the terminal to collect a current person image of a target user;

[0086] It can be understood that when the user is about to perform a head movement, a prompt instruction for the start of the head movement can be sent to the terminal by clicking a virtual button or inputting a voice command, etc. When the terminal receives this prompt instruction, it sends a motion detection request for the head to the server. When the server receives this request, it sends a scheduling instruction to the terminal to call the camera device in the terminal to capture a person image. The terminal includes but is not limited to devices such as mobile phones and sports watches.

[0087] It should be noted that the user should perform the head movement within the shootable field of view of the camera device to record the entire process of the movement. The target user is the user presented within the shootable field of view of the camera device. The shooting mode is continuous shooting, and the focus mode is continuous focus to capture the head movement.

[0088] Optionally, the server can ensure the stability of the imaging of the camera device based on the optical image stabilization algorithm (Optical Image Stabilizer, OIS), the electronic image stabilization algorithm (Electronic Image Stabilizer, EIS), or even physical anti-shake methods (such as a gimbal, gyroscope).

[0089] Optionally, when multiple person images are captured within the lens of the camera device, the server can determine the person who is currently performing a head movement as the target user based on the motion capture algorithm, so as to accurately and quickly focus on the target user; at the same time, when the motion of the person is captured by the motion capture algorithm, the person image acquisition of the target user is triggered, thus realizing unsupervised head movement detection with higher efficiency;

[0090] Optionally, to improve the accuracy of capturing the target user, a certain tolerance value is set in the motion capture algorithm, that is, the person who performs a specific posture or action amplitude in front of the capture lens reaches the preset threshold is determined as the target user, so as to avoid misidentifying a person with normal actions as the target user.

[0091] Optionally, the user can use the terminal to upload their own person image to the server in advance, that is, the server records the person image of the target user, so as to quickly focus on the recorded target user when calling the camera device to take a picture, and then take pictures and record the images of their head movement.

[0092] 102. Input the current person image of the target user into a preset face detection model for detection to obtain the current face region image of the target user, where the face detection model is a lightweight neural network model;

[0093] It can be understood that after training an object detection network with a face dataset for face detection, a face detection model for detecting face regions from images can be obtained. The face dataset can use publicly available network datasets, such as Wider Face, FDDB, Database of Face Attribute Classification, etc. This embodiment does not limit it.

[0094] It should be noted that the face detection model adopted in this embodiment is a lightweight model, so the object detection network trained by it needs to ensure lightweight. In this embodiment, lightweight detection networks such as LFFD, Mobile-Net, and Slim-320 can be used. After training, it can be deployed on devices with limited memory and low computing power for edge computing. Preferably, Slim-320 with a size of only 1MB can be used to train the face detection model.

[0095] 103. Input the current face region image of the target user into a preset face pose detection model for regression calculation to obtain the target deflection angle of the target user's head in the target direction. Here, the face pose detection model is a lightweight neural network model.

[0096] It can be understood that this face pose detection model includes a backbone network and a regression network. The backbone network can adopt lightweight networks such as Mobile-Net, ShuffleNet, and SqueezeNet to extract a high-dimensional non-linear feature from the input face region image, and then input this non-linear feature into the regression network for regression calculation to obtain the corresponding regression value, that is, the target deflection angle of the target user's head in the target direction.

[0097] The target direction includes the horizontal direction (Yaw), or the vertical direction (Pitch), or the horizontal plane direction (Roll), or a combination of the three. When the target direction is the horizontal direction, it means that the target user's head can shake left and right; when the target direction is the vertical direction, it means that the target user's head can shake up and down; when the target direction is the horizontal plane direction, it means that the target user's head can rotate within the plane; when the target direction is a combination of the three, it means that the target user's head can shake in any direction.

[0098] It should be noted that in this embodiment, the specific regression calculation method of the regression network is not specifically limited, and it can be direct regression or branch regression. The former tends to be associated, and the latter is more inclined to independent and identically distributed.

[0099] 104. Compare the target deflection angle with a preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective.

[0100] It can be understood that in the definition of a valid head movement, the deflection angle that the head should reach is specified, that is, the standard deflection angle. For example, the standard deflection angle corresponding to an action definition is 60 degrees. When the actual target deflection angle of the user in a head movement is greater than or equal to 60 degrees, it is determined that this head movement is valid; if the target deflection angle is less than 60 degrees, it is determined that this head movement is invalid.

[0101] Optionally, after receiving the motion detection request sent by the terminal, the server also counts the number of valid movements of the target user within a preset time period; if the number of valid movements of the target user within the preset time period is greater than or equal to a preset threshold, it is determined that the target user's current movement is qualified, and a preset positive motion feedback strategy is executed; if the number of valid movements of the target user within the preset time period is less than the preset threshold, it is determined that the target user's current movement is unqualified, and a preset negative motion feedback strategy is executed. The positive motion feedback strategy and the negative motion feedback strategy are motion incentive behavior measures for rewarding and punishing the performance results of the target user's current movement, and this embodiment does not make specific limitations on them.

[0102] Optionally, when the target direction is a combination of multiple directions, different standard deflection angles can be set according to different directions. For example, the horizontal direction can be set to 60 degrees, and the vertical direction can be set to 30 degrees.

[0103] In the embodiment of the present invention, head movement detection is performed on a terminal device with limited computing power based on a lightweight face detection model and a face pose detection model, reducing the computing power requirements for detection, so that it can be deployed in more application scenarios.

[0104] Please refer to Figure 2 , the second embodiment of the head movement detection method in the embodiment of the present invention includes:

[0105] 201. When receiving the motion detection request sent by the terminal, call the preset camera device in the terminal to collect the current person image of the target user;

[0106] 202. Input the current person image of the target user into the preset face detection model for detection to obtain the current face region image of the target user, where the face detection model is a lightweight neural network model;

[0107] 203. Input the current face region image of the target user into the preset face pose detection model for regression calculation to obtain the target deflection angle of the target user's head in the target direction, where the face pose detection model is a lightweight neural network model;

[0108] 204. Compare the target deflection angle with the preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective.

[0109] Among them, the execution steps of steps 201 - 204 are similar to those of the above steps 101 - 104, and will not be elaborated here specifically.

[0110] 205. Take the intersection point between the horizontal axis and the vertical axis of the face region image as the origin of the coordinate system, take the horizontal direction as the horizontal axis direction of the coordinate system, and take the vertical direction as the vertical axis direction of the coordinate system to construct a plane rectangular coordinate system.

[0111] It can be understood that the establishment of the coordinate system is used to quantify the specific position relationship. In this embodiment, a two - dimensional plane rectangular coordinate system is established. For details, please refer to Figure 8 , where A and B are the mid - points of the two vertical sides of the image respectively, and the straight line AB is the vertical axis of the face region image. Similarly, C and D are the mid - points of the two horizontal sides of the image respectively, and the straight line CD is the horizontal axis of the face region image. The intersection point O of the two axes is the origin of the coordinate system, the straight line AB is the horizontal axis X direction, the straight line CD is the vertical axis Y direction, and the P point represents the current position of the face head.

[0112] Optionally, when the target direction includes Yaw, Pitch, and Roll directions, a three - dimensional coordinate system can be established accordingly.

[0113] 207. Calculate the length of the face region image in the target direction.

[0114] It can be understood that when the target direction is the horizontal direction (Yaw), calculate the length of the face region image in the horizontal direction, that is, Figure 8 the length of the line segment AB in Figure 8 ; when the target direction is the vertical direction (Pitch), calculate the length of the face region image in the vertical direction, that is,

[0115] the length of the line segment CD in

[0116] It can be understood that referring to Figure 9 , the length of the face region image in the horizontal direction is denoted as L AB, the standard deflection angle is denoted as α, the target deflection angle of the current target user's head in the horizontal direction is denoted as β, E represents the target user himself, O is the coordinate origin, and according to its trigonometric geometric relationship, the abscissa X corresponding to the current position of the face head can be calculated P , please refer to Formula 1:

[0117]

[0118] Similarly, referring to Figure 10 , the length of the face area image in the vertical direction is denoted as L CD , the standard deflection angle is denoted as α, the target deflection angle of the current target user's head in the vertical direction is denoted as β, E represents the target user himself, O is the coordinate origin, and according to its trigonometric geometric relationship, the ordinate Y corresponding to the current position of the face head can be calculated P , please refer to Formula 2:

[0119]

[0120] Optionally, after determining and outputting the target coordinate position, the server also determines the stable coordinate interval of the target user's head in the plane rectangular coordinate system based on the preset stable demand information; if the target deflection angle is greater than the preset abnormal deflection threshold or the target coordinate position is not within this stable coordinate interval, it is determined that the current shooting environment is unstable, and stability correction is performed on the target coordinate position, so as to avoid instability caused by terminal instability or natural head shaking. Specifically, the server can preset the stable interval, please refer to Figure 8 , the middle area between the left and right dashed lines is the stable interval in the horizontal direction, and the middle area between the upper and lower dashed lines is the stable interval in the vertical direction. When the target coordinate position (i.e., point P) exceeds the stable interval in a certain direction, the server performs stability correction in that direction. For example, when the abscissa X P of the target coordinate position exceeds the stable interval in the horizontal direction, directly assign X P to L AB / 2, and when the ordinate Y P of the target coordinate position exceeds the stable interval in the vertical direction, directly assign Y P to L CD / 2.

[0121] In the embodiments of the present invention, the process of calculating the specific deflection position of the target user's head is described in detail. By establishing a coordinate system and performing trigonometric geometric operations, the accurate position where the target user's head is currently located is determined, and then whether the movement meets the standard can be accurately judged based on this position or the user can be accurately instructed to perform action correction.

[0122] Please refer toFigure 3 , the third embodiment of the head movement detection method in the embodiments of the present invention includes:

[0123] 301. When receiving a motion detection request sent by the terminal, call the preset camera device in the terminal to collect the current person image of the target user;

[0124] Among them, the execution steps of step 301 are similar to those of the above step 101, and will not be elaborated here specifically.

[0125] 302. Input the current person image of the target user into the preset face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network;

[0126] It can be understood that before inputting the current person image of the target user into the preset face detection model, the server also trains and generates the face detection model, specifically including: inputting the sample image into the initial model; extracting face features of at least one feature size through the initial model; identifying candidate face regions in the face features of each feature size according to the number of prior boxes corresponding to each feature size through the initial model; screening the candidate face regions through the initial model to determine the target face region in the sample image; training the initial model according to the standard face region of the sample image to obtain the face detection model.

[0127] Optionally, after training the initial model and before obtaining the face detection model, it further includes: obtaining the currently trained initial model, and selecting at least one size parameter of the currently trained initial model; adjusting the values of each size parameter to obtain at least one alternative model; inputting multiple sample images into each alternative model to obtain the performance data of each alternative model; screening to obtain the target model according to the performance data of each alternative model, and updating the currently trained initial model.

[0128] Optionally, the size parameter includes the kernel size of at least one convolution kernel and / or the number of output channels of at least one convolution layer.

[0129] 303. Call the face feature extraction network to extract face features of at least one feature size from the current person image of the target user;

[0130] It can be understood that facial features are used to describe the image features of a face. Exemplarily, the image features may include color features, texture features, shape features, spatial relationship features, etc. of the image. The face detection model is used to extract features from the input image to obtain facial features. The feature size is the size of the facial features, and the facial features may be a feature map. The feature size may include parameters such as the width and height of the map. Exemplarily, the face detection model may extract facial features with a feature size of 32*32, facial features with a feature size of 16*16, or facial features with a feature size of 8*8. The face detection model may extract facial features through a feature extraction layer. Exemplarily, the feature extraction layer includes at least one convolutional layer. Optionally, the convolutional layer includes a DW convolution (Depthwise separable convolution), where using the DW convolution can reduce the number of parameters and computational costs, thereby achieving lightweight operations.

[0131] 304. Invoke the facial feature recognition network, and identify candidate face regions in the facial features of each feature size according to the preset number of prior boxes corresponding to each feature size;

[0132] The prior box (Anchor) may refer to boxes with different sizes and / or different aspect ratios preset on the image in advance. The prior box is used to be adjusted to continuously approach the true box containing the object to be detected. In fact, the prior box can be understood as having predefined the width and height of the face region to be detected. During the model prediction process, the width and height are used to process the facial features to predict the face region, where the prediction process is to adjust the (center) position and size of the prior box to obtain a candidate face region that is closest to the true face region. The sizes and / or aspect ratios of different prior boxes are different. The number of prior boxes with the same size and aspect ratio is one. Among them, the face detection model can identify candidate face regions through a classification layer. Exemplarily, the classification layer may include a cascaded Feature Pyramid Networks (FPN).

[0133] It should be noted that the feature size is inversely proportional to the number of prior boxes. The fact that the feature size is inversely proportional to the number of prior boxes indicates that the feature size corresponds to the number of types of prior boxes. The number of types of prior boxes corresponding to different feature sizes is different. The sizes and / or aspect ratios of the prior boxes corresponding to different feature sizes may be partially the same or completely different. Exemplarily, for the 32*32 feature size, there is 1 prior box with an aspect ratio of 1:1. Another example is that for the 16*16 feature size, there are 2 prior boxes with aspect ratios of 1:1 and 2:1 respectively.

[0134] In fact, when traversing prior boxes on face features, the amount of data to be calculated is related to the number, size, and aspect ratio of the prior boxes. Thus, the feature size corresponds to the number of prior boxes, and different numbers of prior boxes can be flexibly selected to be set on different feature sizes for face region prediction. In this embodiment, different numbers of prior boxes are used for detection in face features of different feature sizes, so that the matching calculation amount of the prior boxes corresponding to the face features of different feature sizes is different.

[0135] 305. Invoke the face feature screening network to screen each alternative face region to obtain the current face region image of the target user;

[0136] It can be understood that screening the alternative face regions can be to reduce redundant alternative face regions and incorrect face regions, etc. There can be multiple alternative face regions representing the same face, and the multiple repeated alternative face regions representing the same face can be screened to reduce redundant alternative face regions. In addition, some alternative face regions are incorrect, and the incorrect alternative face regions can be deleted to improve the accuracy of the face region. At least one screened alternative face region is determined as the target face region. The face detection model can screen the alternative face regions through a post-processing layer. Exemplarily, the post-processing layer can use an algorithm such as non-maximum suppression (NMS) or an algorithm based on the intersection over union (IoU) to achieve screening. The target face region can include the region of a sensitive face or the face region including privacy information. Optionally, the server can perform desensitization processing on the target face region to protect sensitive information and privacy information of the target face region, etc.

[0137] 306. Input the current face region image of the target user into a preset face pose detection model for regression calculation to obtain the target deflection angle of the target user's head in the target direction, where the face pose detection model is a lightweight neural network model;

[0138] 307. Compare the target deflection angle with a preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective.

[0139] Among them, the execution steps of steps 306-307 are similar to those of the above steps 103-104, and will not be elaborated here specifically.

[0140] In the embodiment of the present invention, the process of face region detection is described in detail. By flexibly selecting the corresponding number of prior boxes according to the feature size, the calculation amount required for the prior boxes is reduced, and the detection speed of the face region is improved.

[0141] Please refer to Figure 4 , the fourth embodiment of the head movement detection method in the embodiments of the present invention includes:

[0142] 401. When a motion detection request sent by the terminal is received, call the preset camera device in the terminal to collect the current person image of the target user;

[0143] 402. Input the current person image of the target user into the preset face detection model for detection to obtain the current face area image of the target user, where the face detection model is a lightweight neural network model;

[0144] Among them, the execution steps of steps 401-402 are similar to those of the above steps 101-102, and will not be elaborated here specifically.

[0145] 403. Construct a face pose image set for model training, and label the deflection angles of each image in the face pose image set in the target direction;

[0146] It can be understood that the server can obtain corresponding face pose images from the public face dataset to construct the face pose image set, and can perform preprocessing on the face pose images therein, including but not limited to geometric transformation, image enhancement, frequency transformation, etc., so as to improve the robustness of the image samples in the face pose image set. Finally, the server labels the true deflection angle of each image in the target direction based on the image annotation tool for model learning.

[0147] 404. Divide the face pose image set into a face pose training image set and a face pose verification image set according to the preset image division ratio;

[0148] It can be understood that the image division ratio can be adjusted according to actual needs. For example, 80% of the image samples in the image set are used for model training, and 20% of the image samples are used for model verification.

[0149] 405. Based on each image in the face pose training image set, perform regression training on the deflection angle of the preset Mobile-Net convolutional neural network model;

[0150] It can be understood that the Mobile-Net convolutional neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The server calls the pose feature extraction network to extract the target face pose features from the target training images in the face pose training image set; calls the deflection angle classification network to calculate the multi-class deflection interval probability distribution corresponding to the target face pose features according to the target face pose features and the preset deflection interval probability matrix; selects the deflection interval with the largest probability value from the multi-class deflection interval probability distribution as the target deflection interval corresponding to the target face pose features, and obtains the regression function corresponding to the target deflection interval; calls the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose features and the regression function, and obtains the deflection angle of the target training image in the target direction; based on the preset loss function and the annotation information of the deflection angle of the target training image in the target direction, calculates the loss value corresponding to the deflection angle of the target training image in the target direction. The loss value reflects the difference between the result of the model's regression calculation and the true result pre-annotated by the sample, that is, the smaller the loss value, the more accurate the model's regression calculation result.

[0151] Finally, adjust the network parameters of the Mobile-Net convolutional neural network model according to the loss value. For example, the gradient descent of the network parameters can be performed based on the stochastic gradient descent algorithm, and the loss value of the model after the gradient descent of the network parameters is calculated until the network of the model converges. For example, when the loss value reaches the global minimum or the loss value is less than the preset threshold, it is determined that the Mobile-Net convolutional neural network model converges, and the regression training ends.

[0152] 406. Based on each image in the face pose verification image set, verify the regression-trained Mobile-Net convolutional neural network model to obtain a face pose detection model;

[0153] It can be understood that when the model regression training ends, the server verifies the accuracy of the model based on the divided verification image set. If the verification result is accurate, the network parameters of the current Mobile-Net convolutional neural network model are saved to obtain a face pose detection model.

[0154] 407. Input the current face region image of the target user into the face pose detection model for regression calculation to obtain the target deflection angle of the target user's head in the target direction, where the face pose detection model is a lightweight neural network model;

[0155] 408. Compare the target deflection angle with the preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective;

[0156] Among them, the execution steps of steps 407-408 are similar to those of the above steps 103-104, and will not be elaborated here specifically.

[0157] In the embodiments of the present invention, the construction process of the face pose detection model is described in detail. By constructing a face pose image set and dividing it into a training set and a validation set, the training set is used to train the model, and the validation set is used to verify the trained model, so that the face pose detection model can detect more accurately.

[0158] The above describes the head movement detection method in the embodiments of the present invention. Next, the head movement detection device in the embodiments of the present invention will be described. Please refer to Figure 5 , an embodiment of the head movement detection device in the embodiments of the present invention includes:

[0159] An image acquisition module 501, configured to call a preset camera device in the terminal to acquire a current person image of a target user when receiving a movement detection request sent by the terminal;

[0160] A face detection module 502, configured to input the current person image of the target user into a preset face detection model for detection, and obtain a current face region image of the target user, where the face detection model is a lightweight neural network model;

[0161] A deflection calculation module 503, configured to input the current face region image of the target user into a preset face pose detection model for regression calculation, and obtain a target deflection angle of the head of the target user in a target direction, where the face pose detection model is a lightweight neural network model;

[0162] An effective detection module 504, configured to compare the target deflection angle with a preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective.

[0163] In the embodiments of the present invention, head movement detection is performed on a terminal device with limited computing power based on a lightweight face detection model and a face pose detection model, reducing the computing power requirements for detection, and thus can be deployed to more application scenarios.

[0164] Please refer to Figure 6 , another embodiment of the head movement detection device in the embodiments of the present invention includes:

[0165] An image acquisition module 501, configured to call a preset camera device in the terminal to acquire a current person image of a target user when receiving a movement detection request sent by the terminal;

[0166] A face detection module 502 is configured to input the current person image of the target user into a preset face detection model for detection, so as to obtain the current face region image of the target user, where the face detection model is a lightweight neural network model;

[0167] A deflection calculation module 503 is configured to input the current face region image of the target user into a preset face pose detection model for regression calculation, so as to obtain the target deflection angle of the head of the target user in the target direction, where the face pose detection model is a lightweight neural network model;

[0168] An effective detection module 504 is configured to compare the target deflection angle with a preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective;

[0169] A head position calculation module 505 is configured to calculate the position where the head of the target user is located after deflection;

[0170] A movement qualification detection module 506 is configured to count the number of effective movements of the target user within a preset time period; if the number of effective movements of the target user within the preset time period is greater than or equal to a preset threshold, it is determined that the current movement of the target user is qualified, and a preset positive feedback strategy for movement is executed; if the number of effective movements of the target user within the preset time period is less than the preset threshold, it is determined that the current movement of the target user is unqualified, and a preset negative feedback strategy for movement is executed.

[0171] Wherein, the face detection module 502 specifically includes:

[0172] An image input unit 5021 is configured to input the current person image of the target user into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network;

[0173] A face feature extraction unit 5022 is configured to call the face feature extraction network to extract face features of at least one feature size from the current person image of the target user;

[0174] An alternative face recognition unit 5023 is configured to call the face feature recognition network to identify alternative face regions in the face features of each feature size according to a preset number of prior boxes corresponding to each feature size;

[0175] A face screening unit 5024 is configured to call the face feature screening network to screen each of the alternative face regions to obtain the current face region image of the target user.

[0176] Among them, the head position calculation module 505 specifically includes:

[0177] A coordinate system construction unit 5051, configured to use the intersection point between the horizontal central axis and the vertical central axis of the face region image as the origin of the coordinate system, use the horizontal direction as the horizontal axis direction of the coordinate system, and use the vertical direction as the vertical axis direction of the coordinate system to construct a plane rectangular coordinate system;

[0178] A length calculation unit 5052, configured to calculate the length of the face region image in the target direction;

[0179] A coordinate calculation unit 5053, configured to perform trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the current target coordinate position of the target user's head in the plane rectangular coordinate system.

[0180] A stability correction unit 5054, configured to determine a stable coordinate interval of the target user's head in the plane rectangular coordinate system based on preset stability requirement information; if the target deflection angle is greater than a preset abnormal deflection threshold or the target coordinate position is not within the stable coordinate interval, it is determined that the current shooting environment is unstable, and stability correction is performed on the target coordinate position.

[0181] In the embodiment of the present invention, the modular design enables the hardware of each part of the clinical path construction device to focus on the realization of a certain function, maximizing the performance of the hardware. At the same time, the modular design also reduces the coupling between the modules of the device, making maintenance more convenient.

[0182] Above Figure 5 And Figure 6 The head motion detection device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the head motion detection device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0183] Figure 7FIG. 0 is a schematic structural diagram of a head movement detection device provided by an embodiment of the present invention. The head movement detection device 700 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 for storing application programs 733 or data 732 (for example, one or more mass storage devices). Among them, the memory 720 and the storage media 730 may be transient storage or persistent storage. The program stored in the storage media 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the head movement detection device 700. Further, the processor 710 may be configured to communicate with the storage media 730 and execute a series of instruction operations in the storage media 730 on the head movement detection device 700.

[0184] The head movement detection device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input / output interfaces 760, and / or one or more operating systems 731, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 7 the shown structural diagram of the head movement detection device does not limit the head movement detection device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0185] The present invention also provides a head movement detection device. The computer device includes a memory and a processor. When a computer-readable instruction stored in the memory is executed by the processor, the processor executes each step of the head movement detection method in the above embodiments.

[0186] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer executes each step of the head movement detection method.

[0187] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0188] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0189] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0190] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A head movement detection method, characterized in that, The described head movement detection method includes: Construct a face pose image set for model training, and label the deflection angles in the target direction for each image in the face pose image set; According to the preset image division ratio, divide the face pose image set into a face pose training image set and a face pose verification image set; Based on each image in the face pose training image set, perform regression training on the preset Mobile-Net convolutional neural network model for the deflection angle; Based on each image in the face pose verification image set, perform model verification on the Mobile-Net convolutional neural network model after regression training to obtain a face pose detection model; When receiving a motion detection request sent by the terminal, call the preset camera device in the terminal to collect the current person image of the target user; Input the current person image of the target user into the preset face detection model for detection to obtain the current face region image of the target user, where the face detection model is a lightweight neural network model; Input the current face region image of the target user into the face pose detection model for regression calculation to obtain the target deflection angle of the head of the target user in the target direction; Compare the target deflection angle with the preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, determine that the current head movement is effective; The Mobile-Net convolutional neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The regression training of the preset Mobile-Net convolutional neural network model based on each image in the face pose training image set includes: calling the pose feature extraction network to extract the target face pose feature from the target training image in the face pose training image set; calling the deflection angle classification network to calculate the multi-class deflection interval probability distribution corresponding to the target face pose feature according to the target face pose feature and the preset deflection interval probability matrix; selecting the deflection interval with the largest probability value from the multi-class deflection interval probability distribution as the target deflection interval corresponding to the target face pose feature, and obtaining the regression function corresponding to the target deflection interval; calling the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose feature and the regression function to obtain the deflection angle of the target training image in the target direction; calculating the loss value corresponding to the deflection angle of the target training image in the target direction based on the preset loss function and the annotation information of the deflection angle of the target training image in the target direction; performing stochastic gradient descent on the network parameters of the Mobile-Net convolutional neural network model according to the loss value, and calculating the loss value again until the loss value is less than the preset threshold, then determining that the Mobile-Net convolutional neural network model converges and ending the regression training.

2. The head movement detection method according to claim 1, wherein After comparing the target deflection angle with a preset standard deflection angle, if the target deflection angle reaches the standard deflection angle and it is determined that the current head movement is effective, the following steps are further included: Taking the intersection point between the horizontal axis and the vertical axis of the face region image as the origin of the coordinate system, with the horizontal direction as the horizontal axis direction of the coordinate system and the vertical direction as the vertical axis direction of the coordinate system, a plane rectangular coordinate system is constructed; Calculating the length of the face region image in the target direction; Performing trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the target coordinate position of the target user's head in the plane rectangular coordinate system currently.

3. The head movement detection method according to claim 2, wherein After performing trigonometric geometric operations based on the length of the face region image in the target direction, the target deflection angle, and the standard deflection angle to determine and output the target coordinate position of the target user's head in the plane rectangular coordinate system currently, the following steps are further included: Based on the preset stability requirement information, determining the stable coordinate interval of the target user's head in the plane rectangular coordinate system; If the target deflection angle is greater than the preset abnormal deflection threshold or the target coordinate position is not within the stable coordinate interval, it is determined that the current shooting environment is unstable, and stability correction is performed on the target coordinate position.

4. The head movement detection method according to claim 1, wherein The step of inputting the current person image of the target user into a preset face detection model for detection to obtain the current face region image of the target user includes: Inputting the current person image of the target user into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network; Invoking the face feature extraction network to extract face features of at least one feature size from the current person image of the target user; Invoking the face feature recognition network to identify candidate face regions in the face features of each feature size according to a preset number of prior boxes corresponding to each feature size; Invoking the face feature screening network to screen each of the candidate face regions to obtain the current face region image of the target user.

5. The head movement detection method according to any one of claims 1-4, characterized in that, After comparing the target deflection angle with a preset standard deflection angle, if the target deflection angle reaches the standard deflection angle and it is determined that the current head movement is effective, the following steps are further included: Counting the number of effective movements of the target user within a preset time period; If the number of effective movements of the target user within the preset time period is greater than or equal to a preset threshold, it is determined that the current movement of the target user is qualified, and a preset positive movement feedback strategy is executed; If the number of effective movements of the target user within the preset time period is less than the preset threshold, it is determined that the current movement of the target user is unqualified, and a preset negative movement feedback strategy is executed.

6. A head movement detection device, characterized in that, The head movement detection device includes: A face pose detection model construction module, which is used to construct a face pose image set for model training, and label the deflection angles of each image in the face pose image set in the target direction; divide the face pose image set into a face pose training image set and a face pose verification image set according to a preset image division ratio; perform regression training on the preset Mobile-Net convolutional neural network model based on each image in the face pose training image set; perform model verification on the Mobile-Net convolutional neural network model after regression training based on each image in the face pose verification image set to obtain a face pose detection model; An image acquisition module, which is used to call a preset camera device in the terminal to acquire the current person image of the target user when receiving a motion detection request sent by the terminal; A face detection module, which is used to input the current person image of the target user into a preset face detection model for detection to obtain the current face region image of the target user, wherein the face detection model is a lightweight neural network model; A deflection calculation module, which is used to input the current face region image of the target user into the face pose detection model for regression calculation to obtain the target deflection angle of the head of the target user in the target direction; An effective detection module, which is used to compare the target deflection angle with a preset standard deflection angle. If the target deflection angle reaches the standard deflection angle, it is determined that the current head movement is effective; The Mobile-Net convolutional neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The regression training of the preset Mobile-Net convolutional neural network model for the deflection angle based on each image in the face pose training image set includes: calling the pose feature extraction network to extract the target face pose feature from the target training image in the face pose training image set; calling the deflection angle classification network to calculate the multi-class deflection interval probability distribution corresponding to the target face pose feature according to the target face pose feature and the preset deflection interval probability matrix; selecting the deflection interval with the largest probability value from the multi-class deflection interval probability distribution as the target deflection interval corresponding to the target face pose feature, and obtaining the regression function corresponding to the target deflection interval; calling the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose feature and the regression function to obtain the deflection angle of the target training image in the target direction; calculating the loss value corresponding to the deflection angle of the target training image in the target direction based on the preset loss function and the annotation information of the deflection angle of the target training image in the target direction; performing stochastic gradient descent on the network parameters of the Mobile-Net convolutional neural network model according to the loss value, and calculating the loss value again until the loss value is less than the preset threshold, determining that the Mobile-Net convolutional neural network model converges, and ending the regression training.

7. A head movement detection device, characterized in that, The head movement detection device includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the head movement detection device executes each step of the head movement detection method according to any one of claims 1-5.

8. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, each step of the head movement detection method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Head angle prediction model training method, prediction method, device and medium

    CN108920999A

  • Head angle labeling method, prediction model training method, prediction method, device and medium

    CN108921000A