A wave-for-help recognition system and method
By combining the information of video stream and audio stream, using deep learning models to wave help recognition, the problem of unstable posture key point recognition in the prior art is solved, and higher recognition accuracy and stability are achieved.
Patent Information
- Application Number
- CN202211259423.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-10-14
AI Technical Summary
The existing wave-wave recognition system has unstable and inaccurate recognition points for human posture due to factors such as occlusion, viewing angle, lighting, scale and motion blur, and cannot meet the high accuracy requirements of real use scenarios.
Combining the information of the video stream and audio stream, the MFCC acoustic features and face images are obtained through the feature extraction unit, the deep learning model is used to detect hand-waving movements, amplitude and frequency, and the comprehensive weighting method is used to determine the helping action, and the posture, sound and expression information are fused to improve the recognition accuracy.
By combining video and audio information, the accuracy of wave help recognition is improved, and the stability and reliability of judgment of the help state is enhanced, especially in complex environments, the wave help action can be better recognized.
Smart Images

Figure CN115641610B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence recognition, and specifically relates to a wave-for-help recognition system and method. Background Art
[0002] With the development of deep learning and machine vision, applications based on artificial intelligence have gradually matured and been widely used. In existing wave-for-help recognition systems, most are based on visual images to obtain pose key points, and then a deep learning model or artificial logic rules are used to establish a help-seeking recognition model to determine whether it is a waving action, and at the same time, the waving frequency is detected to identify help-seeking behaviors.
[0003] This method relies on the recognition of key points. Due to factors such as occlusion, perspective, lighting, scale, and motion blur, the recognition of human body pose key points will be interfered, resulting in unstable and inaccurate recognition results, and unable to meet the requirements of high accuracy in real usage scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide a stable and highly reliable wave-for-help recognition method. In view of the above problems, several strategies are proposed to improve and enhance the effect of help-seeking recognition.
[0005] The technical solution adopted by the present invention to achieve the above purpose is: a wave-for-help recognition system, including: a feature extraction unit, a calculation module, and a wave-for-help detection unit;
[0006] The feature extraction unit is used to receive the audio stream and video stream sent by the camera, extract MFCC acoustic features and obtain the preprocessed face image, and send them to the wave-for-help detection unit respectively. At the same time, after extracting the human skeleton pose information in the video stream, obtain the human skeleton key point information and transmit it to the calculation module;
[0007] The calculation module is used to perform waving action detection, waving amplitude detection, and detection of the sitting, lying, or standing posture of the person according to the human skeleton key point information, and send the detection results including the stretch index, waving action state, and sitting, lying, or standing posture of the person to the wave-for-help detection unit; at the same time, according to the human key point coordinates, obtain the stretch index, and then obtain the waving frequency, and send the waving frequency to the wave-for-help detection unit;
[0008] The wave-for-help detection unit is used to process the MFCC acoustic features extracted by the feature extraction unit, the sent face image, and the detection results sent by the calculation module, and use the comprehensive weight method to determine whether the person is performing a wave-for-help action.
[0009] The feature extraction unit includes: a sound feature extraction module, a facial feature preprocessing module, and a human body pose detection module;
[0010] A voice feature extraction module, which is used to receive the audio stream sent by the camera, process the continuous speech in the audio stream, obtain N-dimensional MFCC features, and send them to the waving for help detection unit;
[0011] A facial feature preprocessing module, which is used to obtain images from the camera video stream, perform human face detection by machine learning method, remove the background and non-face areas to obtain the face point coordinates, perform face feature normalization processing on the face image according to the face point coordinates, and send the processed face image to the waving for help detection unit;
[0012] A human body posture detection module, which is used to receive the video stream sent by the camera, extract the human skeleton posture information through a deep neural network, obtain the human skeleton key point information, and transmit it to the calculation module.
[0013] The said calculation module includes: a waving action detection module and a stretching degree index calculation module;
[0014] The waving action detection module is used to perform normalization preprocessing of coordinates and scales according to the human body key point information to obtain key point features, judge whether the person in the standing, sitting or lying posture has a waving action state, and obtain the waving amplitude information according to the waving action state; at the same time, according to the posture detection model, output the waving posture confidence of the person corresponding to the standing, sitting or lying waving action state at this moment; and send the human body corresponding standing, sitting or lying posture information, waving amplitude information and waving posture confidence to the waving for help detection unit;
[0015] The stretching degree index calculation module is used to obtain the stretching degree index according to the human body key point coordinates to reflect the posture stretching degree of the human body, and calculate the waving frequency according to the stretching degree index; send the waving frequency to the waving for help detection unit.
[0016] The said waving for help detection unit includes: a voice event recognition module, a facial expression recognition module, and a waving for help detection module;
[0017] The voice event recognition module is used to perform event detection by using a deep learning model according to the MFCC features extracted by the feature extraction unit, and obtain a voice classification result including the classified speech and the confidence information of the current frame voice output;
[0018] The facial expression recognition module is used to extract facial features from the face image sent by the feature extraction unit through a deep learning model, and perform training and inference to obtain a facial expression classification result including the classified human facial expression and the confidence information of the current frame facial expression output;
[0019] A wave-for-help detection module, which is used to determine whether a person is performing a wave-for-help action by using the comprehensive weight method based on the voice classification result, the facial expression classification result, and the posture information of the human body corresponding to standing, sitting, and lying, the waving amplitude information, the waving posture confidence, and the waving frequency sent by the calculation module.
[0020] A recognition method for a wave-for-help recognition system, including the following steps:
[0021] 1) The audio stream sent by the camera is sent to the sound feature extraction module, and the video stream is sent to the facial feature preprocessing module and the human body posture detection module respectively;
[0022] 2-1) The sound feature extraction module receives the audio stream sent by the camera, processes the continuous speech, obtains the MFCC acoustic features, and sends them to the wave-for-help detection unit;
[0023] 2-2) The facial feature preprocessing module obtains images through the camera video stream, performs human face detection by machine learning method, removes the background and non-face regions to obtain the face point coordinates, normalizes the face features of the face image according to the face point coordinates, and sends the processed face image to the wave-for-help detection unit;
[0024] 2-3) The human body posture detection module receives the video stream sent by the camera, extracts the human body skeleton posture information through the deep neural network, obtains the human body skeleton key point information, and transmits it to the calculation module;
[0025] 3-1) The calculation module performs normalization preprocessing of coordinates and scales according to the human body key point information, obtains the key point features, judges whether the person with the posture information of standing, sitting, and lying has a waving action state, and obtains the waving amplitude information according to the waving action state; at the same time, according to the posture detection model, the waving posture confidence of the person corresponding to standing, sitting, and lying at this moment is output; and the posture information of the human body corresponding to standing, sitting, and lying, the waving amplitude information, and the waving posture confidence are sent to the wave-for-help detection unit;
[0026] 3-2) The calculation module obtains the stretch index δ according to the human body key point information HPS and calculates the waving frequency according to the stretch index δ HPS and sends the waving frequency to the wave-for-help detection unit;
[0027] 4) The wave-for-help detection unit uses a deep learning model to perform event detection based on the MFCC acoustic features extracted by the feature extraction unit, including the classified speech and the confidence information of the corresponding current-frame sound output in the speech classification result; at the same time, based on the face image sent by the feature extraction unit, facial features are extracted through a deep learning model, and training and inference are performed to obtain a facial expression classification result including the classified facial expressions of the human body and the confidence information of the corresponding current-frame facial expression output.
[0028] 5) The wave-for-help detection unit determines whether a person is performing a wave-for-help action using the comprehensive weight method based on the speech classification result, the facial expression classification result, and the posture information of the human body standing, sitting, or lying, the wave amplitude information, the wave posture confidence, and the wave frequency sent by the calculation module.
[0029] In step 2-1), after processing the continuous speech, MFCC acoustic features are obtained, specifically:
[0030] The continuous speech of the audio stream is sequentially subjected to pre-emphasis, framing, and windowing operations to obtain the preprocessed sound information.
[0031] For the preprocessed sound information, fast Fourier transform, Mei filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction are sequentially performed, and finally N-dimensional MFCC acoustic features are obtained.
[0032] The specific content of step 3-1) is as follows:
[0033] 3-1-1) The calculation module obtains the standing, sitting, or lying posture information of the person at this moment and the confidence of the current-frame wave posture corresponding to standing, sitting, or lying through a posture detection model based on the human body key point information; the posture detection model is a CNN model.
[0034] 3-1-2) At the same time, based on the human body key point information, the included angle data between each joint is obtained, and based on the included angle data between the forearm and the upper arm and the included angle data between the upper arm and the shoulder in the included angle data between each joint, the wave amplitude information of the human body during the wave action is obtained.
[0035] 3-1-3) The calculation module sends the standing, sitting, or lying posture information, wave amplitude information, and wave posture confidence of the human body to the wave-for-help detection unit.
[0036] 8. The recognition method of a wave-for-help recognition system according to claim 5, characterized in that in step 3-2), the calculation module obtains the stretch index δ HPS , specifically:
[0037] 3-2-1) For the human key point information of a given pose, the calculation module constructs a matrix structure as follows:
[0038] X = [x1,…,x n ∈ R D×n
[0039] where D is the key point dimension, n is the number of key points, x n is the coordinate of the nth key point, and R is the set of real numbers;
[0040] 3-2-2) Let be the mean of all key point coordinates in X, and define Then the variance sum δ HPS of the main axis of the human pose key points is defined as:
[0041]
[0042] where U is the projection matrix, x i is the coordinate of the ith key point, d is a positive number less than the key point dimension D, and λ j is the covariance matrix the jth eigenvalue, and Tr represents the trace of the matrix;
[0043] 3-2-3) According to the stretch index δ HPS calculate the waving frequency, specifically:
[0044] δ HPS is the variance sum of the main axis of the human pose key points, that is, it reflects the stretch of the human pose;
[0045] When the hand waves to both sides of the human body, the value of δ HPS is the largest; when the hand waves to the top of the head and is in a straight line with the human torso, the value of δ HPS is the smallest. By statistically analyzing the periodic transformation law of δ HPS on the time axis, the waving frequency information during the waving action is obtained; the obtained waving frequency information is sent to the waving for help detection unit.
[0046] The specific content of step 4) is as follows:
[0047] 4-1) For voice event classification, specifically:
[0048] During model training, it includes normalizing the input MFCC features, simultaneously reading the label text corresponding to the expression name for representation to generate a label vector, binding the MFCC features and the label vector, then performing preprocessing, and next sending them into a deep learning model to learn the parameters to obtain a voice classification model;
[0049] Among them, the deep learning model is any one of the CNN model, the RCNN model, and the LSTM model;
[0050] During model prediction, the input MFCC features are normalized, then preprocessed, and then sent to the speech classification model for prediction. The obtained data is post-processed to obtain a speech classification result including the classified speech and the confidence information of the corresponding current frame sound output;
[0051] 4-2) For facial expression classification, specifically:
[0052] The deep learning model is used to extract facial features from the face image;
[0053] Among them, the deep learning model is any one of the convolutional neural network, the deep belief network, the deep autoencoder, and the recurrent neural network;
[0054] After completing the facial feature extraction, the convolutional neural network is used to classify the facial expression to obtain a facial expression classification result including the classified facial expression of the human body and the confidence information of the corresponding current frame facial expression output.
[0055] The specific content of step 5) is as follows:
[0056] (1) The wave-for-help detection unit uses the speech classification result including the classified speech and the confidence information of the corresponding current frame sound output, the facial expression classification result including the classified facial expression of the human body and the confidence information of the corresponding current frame facial expression output, the body's standing, sitting, and lying posture information, wave amplitude information, wave gesture confidence, and wave frequency sent by the calculation module, and uses the comprehensive weight method to obtain the global strategy of waving for help:
[0057]
[0058] Among them, is the wave-for-help confidence at the current time t, is the wave-for-help gesture confidence at the current time t, is the sound event detection confidence at the current time t, is the facial expression confidence at the current time t; w pose , w sound , w face are the wave gesture weight, the sound event weight, and the facial expression weight respectively;
[0059] Among them:
[0060] w pose +w sound +w face =1
[0061] w pose , w sound , w face is a preset assignment, w pose > w sound , w pose > w face ;
[0062] (2) Obtain by means of exponential weighted moving average That is:
[0063]
[0064]
[0065]
[0066] where β represents the weighting coefficient, is the confidence level of the waving for help gesture in the current frame, is the confidence level information output by the voice detection in the current frame, is the confidence level information output by the facial expression in the current frame;
[0067] (3) For the confidence level of the waving for help gesture in the current frame is obtained by integrating the confidence level of the waving gesture in the current frame, the waving amplitude information, the waving frequency information and the gesture bonus information, that is:
[0068]
[0069] where, is the confidence level of the waving gesture in the current frame, is the waving amplitude information in the current frame, is the waving frequency information in the current frame, is the gesture bonus coefficient;
[0070] (4) The waving amplitude information in the current frame That is:
[0071]
[0072] where, is the detected waving amplitude, is the maximum preset waving amplitude, is a value between 0 and 1, and the larger the waving amplitude, the higher this value;
[0073] (5) The waving frequency information in the current frame Specifically:
[0074]
[0075] Among them, is the detected waving frequency, is the preset maximum waving frequency, is a value between 0 and 1. The greater the waving frequency, the higher this value;
[0076] (6) Finally, is judged with the preset threshold. If it is greater than the threshold, the current frame is in the waving for help state; if n consecutive frames are in the waving for help state, then a waving for help alarm is issued.
[0077] The present invention has the following beneficial effects and advantages:
[0078] 1. The present invention introduces sound information, combines the video stream with the audio stream, and uses the audio information to assist in improving the accuracy of distress recognition;
[0079] 2. The present invention introduces face facial expression recognition and assists in judging the distress state when the facial expression is panicked or frightened;
[0080] 3. For posture recognition in the present invention, when it is considered to be in the waving state, the standing / sitting / lying and other information of the person's posture is judged, and at the same time, the waving amplitude of the arm is obtained. When the waving amplitude of the arm is large, the standing / sitting / lying posture information of the person is combined at the same time to assist in judging the distress state;
[0081] 4. The waving action in the present invention is a periodic frequency action. Through the stretch evaluation index, the periodicity of the waving action is judged, so as to obtain the action frequency of the waving, and this is used to assist in judging whether the person is waving normally or making a distress action
[0082] 5. The waving for help fusion strategy of the present invention: According to the information data obtained from the waving posture model, the sound event detection model, and the expression recognition model, a multi-state fusion strategy is proposed, making the result of the final waving for help determination more robust. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is the system framework diagram of the present invention;
[0084] Figure 2 is the working flow chart of the feature extraction by the sound feature extraction module of the present invention;
[0085] Figure 3 The present invention is the schematic diagram of establishing a voice classification model;
[0086] Figure 4 The circuit diagram of the balance management module of the present invention;
[0087] Figure 5 The schematic diagram of the human body key points of the present invention;
[0088] Among them, Figure 5 The corresponding relationships of the key point numbers in Figure 5 are as follows: 0: nose, 1: neck, 2: right shoulder, 3: right elbow, 4: right wrist, 5: left shoulder, 6: left elbow, 7: left wrist, 8: middle hip, 9: right hip, 10: right knee, 11: right ankle, 12: left hip, 13: left knee, 14: left ankle, 15: right eye, 16: left eye, 17: right ear, 18: left ear, 19: left thumb, 20: left little finger, 21: left heel, 22: right thumb, 23: right little finger, 24: right heel. Specific implementation manner
[0089] As Figure 1 shown, it is the system framework diagram of the present invention. The overall structure of the system of the present invention is divided into two major parts, the front-end camera and the back-end application server.
[0090] The front-end Camera accesses the back-end AppServer through RTSP or API, and multiple modules in the AppServer perform artificial intelligence analysis and processing on the video stream and audio stream.
[0091] The AppServer includes: a feature extraction unit, a calculation module, and a waving for help detection unit. The functions of each unit / module are as follows:
[0092] The feature extraction unit is used to receive the audio stream and video stream sent by the camera, extract the MFCC acoustic features and obtain the preprocessed face image, and send them to the waving for help detection unit respectively. At the same time, after extracting the human skeleton pose information in the video stream, obtain the human skeleton key point information and transmit it to the calculation module;
[0093] The calculation module is used to detect the waving action, waving amplitude, and the sitting, lying or standing posture of the person according to the human skeleton key point information, and send the detection results including the stretch index, waving action state, and sitting, lying or standing posture of the person to the waving for help detection unit; at the same time, according to the human key point coordinates, obtain the stretch index, and then obtain the waving frequency, and send the waving frequency to the waving for help detection unit;
[0094] The waving for help detection unit is used to process the MFCC acoustic features extracted by the feature extraction unit, the sent face image, and the detection results sent by the calculation module, and use the comprehensive weight method to determine whether the person is performing a waving for help action.
[0095] Among them, the feature extraction unit includes: a sound feature extraction module, a facial feature preprocessing module, and a human body posture detection module;
[0096] A voice feature extraction module, which is used to receive the audio stream sent by the camera, process the continuous speech in the audio stream, obtain N-dimensional MFCC features, and send them to the waving for help detection unit;
[0097] A facial feature preprocessing module, which is used to obtain images from the camera video stream, perform human face detection by machine learning method, remove the background and non-face areas to obtain the face point coordinates, perform face feature normalization processing on the face image according to the face point coordinates, and send the processed face image to the waving for help detection unit;
[0098] A human body posture detection module, which is used to receive the video stream sent by the camera, extract the human skeleton posture information through a deep neural network, obtain the human skeleton key point information, and transmit it to the calculation module.
[0099] The calculation module includes: a waving action detection module and a stretching degree index calculation module;
[0100] The waving action detection module is used to perform normalization preprocessing of coordinates and scales according to the human key point information to obtain key point features, judge whether the person in the standing, sitting or lying posture has a waving action state, and obtain the waving amplitude information according to the waving action state; at the same time, according to the posture detection model, output the waving posture confidence of the person corresponding to the standing, sitting or lying waving action state at this moment; and send the corresponding standing, sitting or lying posture information, waving amplitude information and waving posture confidence of the human body to the waving for help detection unit;
[0101] The stretching degree index calculation module is used to obtain the stretching degree index according to the human key point coordinates to reflect the posture stretching degree of the human body, and calculate the waving frequency according to the stretching degree index; send the waving frequency to the waving for help detection unit.
[0102] The waving for help detection unit includes: a voice event recognition module, a facial expression recognition module, and a waving for help detection module;
[0103] The voice event recognition module is used to perform event detection using a deep learning model according to the MFCC features extracted by the feature extraction unit, and obtain a voice classification result including the classified speech and the confidence information of the current frame voice output;
[0104] The facial expression recognition module is used to extract facial features from the face image sent by the feature extraction unit through a deep learning model, and perform training and inference to obtain a facial expression classification result including the classified facial expressions of the human body and the confidence information of the current frame facial expression output;
[0105] The waving for help detection module is used to determine whether a person is making a waving for help action by using the comprehensive weight method based on the voice classification result, the facial expression classification result, and the information of the corresponding standing, sitting, and lying postures of the human body, the waving amplitude information, the waving posture confidence, and the waving frequency sent by the calculation module.
[0106] As Figure 1 shown, the working process method of the present invention is specifically implemented based on the data streams transmitted by each module. The method of the present invention includes the following steps:
[0107] 1) The audio stream sent by the camera goes to the sound feature extraction module, and the video stream is sent to the facial feature preprocessing module and the human body posture detection module respectively;
[0108] 2-1) The sound feature extraction module receives the audio stream sent by the camera, processes the continuous speech, obtains the MFCC acoustic features, and sends them to the waving for help detection unit;
[0109] 2-2) The facial feature preprocessing module obtains images through the camera video stream, performs human face detection by machine learning method, removes the background and non-face areas to obtain the face point coordinates, normalizes the face features of the face image according to the face point coordinates, and sends the processed face image to the waving for help detection unit;
[0110] 2-3) The human body posture detection module receives the video stream sent by the camera, extracts the human body skeleton posture information through the deep neural network, obtains the human body skeleton key point information, and transmits it to the calculation module;
[0111] 3-1) The calculation module performs normalization preprocessing of coordinates and scales according to the human body key point information, obtains the key point features, determines whether the person in the standing, sitting, and lying posture information has a waving action state, and obtains the waving amplitude information according to the waving action state; at the same time, according to the posture detection model, the waving posture confidence of the person corresponding to standing, sitting, and lying at this moment is output; and the posture information, waving amplitude information, and waving posture confidence of the human body corresponding to standing, sitting, and lying are sent to the waving for help detection unit;
[0112] 3-2) The calculation module obtains the stretch index δ HPS according to the human body key point information, and calculates the waving frequency according to the stretch index δ HPS and sends the waving frequency to the waving for help detection unit;
[0113] 4) The wave-for-help detection unit uses a deep learning model to perform event detection based on the MFCC acoustic features extracted by the feature extraction unit, including the classified speech and the confidence information of the corresponding current frame sound output in the speech classification result; at the same time, according to the face image sent by the feature extraction unit, facial features are extracted through a deep learning model, and training and inference are performed to obtain a facial expression classification result including the classified facial expression of the human body and the confidence information of the corresponding current frame facial expression output.
[0114] 5) The wave-for-help detection unit determines whether a person is performing a wave-for-help action by using the comprehensive weight method based on the speech classification result, the facial expression classification result, and the posture information (standing, sitting, lying) of the human body, the waving amplitude information, the waving posture confidence, and the waving frequency sent by the calculation module.
[0115] As Figure 2 shown, it is the working flowchart of the feature extraction by the sound feature extraction module of the present invention, which specifically includes the following steps:
[0116] Perform preprocessing on continuous speech, including steps such as pre-emphasis, framing, and windowing.
[0117] For the sound information after preprocessing, perform steps such as fast Fourier transform, Mei filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction. Finally, 39-dimensional MFCC features (Mel-scale Frequency Cepstral Coefficients, abbreviated as MFCC) are extracted in this embodiment.
[0118] As Figure 3 shown, it is the schematic diagram of establishing a speech classification model. Based on a deep learning model, including but not limited to convolutional neural networks or recurrent neural networks such as CNN / RCNN / LSTM, the input speech features are recognized.
[0119] During model training, it includes normalizing the input MFCC features, simultaneously reading the label text representing the true expression to generate a label vector, binding the features and the labels, then performing preprocessing, and next sending them into the CNN / RCNN / LSTM model to learn the parameters to obtain a reliable speech classification model.
[0120] As Figure 4 shown, during model prediction, the input 39-dimensional MFCC features are normalized, then preprocessed, and next sent into the speech classification model for prediction. The obtained data is post-processed to obtain a speech classification result including the classified speech and the confidence information of the corresponding current frame sound output.
[0121] As Figure 1As shown, the facial feature preprocessing module in this system obtains images through a video stream, performs face detection through machine learning, and then removes the background and non-face regions.
[0122] Normalization of facial features: Normalize the brightness of the face image content to make the distribution of facial brightness pixels as uniform as possible, and at the same time enhance the facial contrast.
[0123] The facial expression recognition module extracts features through deep learning models, including convolutional neural network (CNN), deep belief network (DBN), deep autoencoder (DAN), and recurrent neural network (RNN).
[0124] After feature extraction, classify the facial expressions. Facial expression recognition can be performed in a deep learning manner to form an end-to-end model, or traditional machine learning methods, such as classifiers like SVM used in this embodiment, can be used for classification to obtain whether a person's face has expressions of panic or fluster.
[0125] As Figure 1 and Figure 5 shown, the human body pose detection module trains and predicts the key points of the human body pose through the deep learning human body pose network model structure, specifically as shown in the schematic diagram of human key points in Figure 5 .
[0126] In the waving for help recognition method according to the present invention, step 3-1) is specifically as follows:
[0127] 3-1-1) The waving action detection module, based on the human key point information (as shown in Figure 5 ), through the pose detection model, obtains the standing, sitting or lying pose information of the person at this moment, as well as the confidence of the current frame waving pose corresponding to standing, sitting, and lying; the pose detection model is a CNN model; it learns whether the person with the pose information has a waving state and gives the confidence of the person being standing, sitting, or lying at this moment. When in the sitting and lying states, it plays a better auxiliary role in judging whether the output is a distress signal.
[0128] 3-1-2) At the same time, based on the human key point information, obtain the included angle data between each joint. According to the included angle data between the forearm and the upper arm in the included angle data between each joint (that is, as shown in Figure 5 , the included angle formed between key points 2, 3, and 4), and the included angle data between the upper arm and the shoulder (the included angles formed between key points 1-4 and between key points 1 and key points 5-7), obtain the waving amplitude information of the human body during the waving behavior;
[0129] 3-1-3) The stretch index calculation module uses the standing, sitting, and lying pose information of the human body (similarly, according to the appendix Figure 5The key point information, waving amplitude information, and waving gesture confidence identified therein are sent to the waving for help detection unit.
[0130] In step 3-2), the calculation module obtains the stretch index δ according to the human key point information HPS , specifically:
[0131] Introduce the δ HPS index to calculate the stretch of the human body posture, so as to assist in judging the waving frequency of the waving detection. δ HPS is the sum of the distribution variances of the human body posture key points on their main axes.
[0132] 3-2-1) As shown by Figure 4 , for the human key point information of a given posture, the stretch index calculation module constructs a matrix structure as:
[0133] X = [x1,…,x n ∈ R D×n
[0134] where D is the key point dimension, n is the number of key points, x n is the coordinate of the nth key point, and R is the set of real numbers;
[0135] 3-2-2) Let be the mean of all key point coordinates in X, and define then the variance sum δ of the main axis of the human body posture key points HPS is defined as Eq1:
[0136]
[0137] Eq1 is just one form of the δ HPS expression, and it can also be written in the following form Eq2:
[0138]
[0139] where U is the projection matrix, x i is the coordinate of the ith key point, d is a positive number less than the key point dimension D, in this embodiment, D is taken as 2 and d is taken as 1, λ j is the covariance matrix the jth eigenvalue, and Tr represents the trace of the matrix;
[0140] Therefore, δ HPS is the variance sum of the main axis of the human body posture key points, used to reflect the stretch of the human posture.
[0141] 3-2-3) The stretch index calculation module calculates the waving frequency according to the stretch index δ HPS , specifically:
[0142] Apply the δ HPS index to the waving motion. Since the waving motion is also a periodic activity, when the hand waves to both sides of the body during this motion, the δ HPS value is large, and when the hand waves to the top of the head and forms a straight line with the human torso, the δ HPS value is small. By statistically analyzing the periodic transformation law of δ HPS on the time axis, the waving frequency during the waving motion can be obtained, and the obtained waving frequency information is sent to the waving for help detection unit.
[0143] The advantages of δ HPS are as follows: This index is only related to human movements, independent of information such as camera perspective and scale, and is not affected by pose interference. It can effectively reflect the periodic data of a person's waving frequency, thereby determining whether a person is waving normally or waving violently for help.
[0144] In step 5) of the present invention, it specifically includes the following steps:
[0145] (1) The waving for help detection unit uses the comprehensive weight method to obtain the global strategy for waving for help based on the speech classification result including the classified speech and the confidence information of the current frame sound output, the facial expression classification result including the classified human facial expression and the confidence information of the current frame facial expression output, the posture information of the human body corresponding to standing, sitting, and lying sent by the calculation module, the waving amplitude information, the waving posture confidence, and the waving frequency:
[0146]
[0147] Among them, is the confidence of waving for help at the current moment t, is the confidence of the waving for help posture at the current moment t, is the confidence of the sound event detection at the current moment t, is the facial expression confidence at the current moment t; w pose , w sound , w face are the waving posture weight, the sound event weight, and the facial expression weight respectively;
[0148] Among them:
[0149] w pose + w sound + w face = 1
[0150] w pose , w sound , w face are preset values, w pose > w sound , wpose > w face ; In this embodiment, since it is a wave for help, the value of w pose can be higher. It can be set as follows:
[0151] w pose = 0.8, w sound = 0.1, w face = 0.1
[0152] The above is only an example, and other values can also be selected according to the situation in actual applications.
[0153] To avoid false detection problems in single-frame detection of the model, it is obtained by using the exponential weighted moving average method, that is, step (2) is as follows:
[0154] (2) Obtain through the exponential weighted moving average method Taking into account the current data and historical data, and assigning a higher weight to the current data, that is:
[0155]
[0156]
[0157]
[0158] Among them, β represents the weighting coefficient, is the confidence level of the wave-for-help gesture in the current frame, is the confidence information output by the sound detection in the current frame, is the confidence information output by the facial expression in the current frame;
[0159] There is a deviation at the initial moment, so deviation correction is introduced. The formula is as follows:
[0160]
[0161]
[0162]
[0163] (3) For the confidence level of the wave-for-help gesture in the current frame It is comprehensively obtained from the confidence level of the wave gesture in the current frame, the wave amplitude information, the wave frequency information, and the gesture bonus information, that is:
[0164]
[0165] Among them, is the confidence level of the wave gesture in the current frame, is the wave amplitude information in the current frame, is the waving frequency information of the current frame, is the posture bonus coefficient;
[0166] (4) Waving amplitude information of the current frame That is:
[0167]
[0168] Among them, is the detected waving amplitude, is the maximum preset waving amplitude, is a value between 0 and 1, and the larger the waving amplitude, the higher this value;
[0169] (5) Waving frequency information of the current frame Specifically:
[0170]
[0171] Among them, is the detected waving frequency, is the maximum preset waving frequency, is a value between 0 and 1, and the larger the waving frequency, the higher this value;
[0172] In practical applications, gamma transformation can be performed according to the situation to perform non-linear compensation on the waving amplitude information and the waving frequency information, and adjust their final weight values in the waving posture determination, as follows:
[0173]
[0174]
[0175] γ is an empirical value, generally taking a value slightly less than 1.
[0176] is the posture bonus coefficient, and its calculation is as follows:
[0177] If the posture detection model detects that the person is in a sitting / lying posture, we believe that the expected probability of waving for help in the sitting / lying posture is greater, and it can be set to the empirical value 0.1, or this bonus value can be increased or decreased:
[0178]
[0179] Otherwise, if the person's posture is standing:
[0180]
[0181] (6) Finally, V t (sos)Judge with a preset threshold. If it is greater than the threshold, the current frame is in the state of waving for help; if it is in the state of waving for help for consecutive n frames, then issue a warning for waving for help.
[0182] The above are only the embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, expansions, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A wave-for-help recognition system, characterized in that, Including: A feature extraction unit, a calculation module, and a waving for help detection unit; The feature extraction unit is used to receive the audio stream and video stream sent by the camera, extract the MFCC acoustic features and obtain the preprocessed face image, and send them to the waving for help detection unit respectively. At the same time, after extracting the human skeleton pose information in the video stream, obtain the human skeleton key point information and transmit it to the calculation module; The calculation module is used to detect the waving action, waving amplitude, and the sitting, lying or standing posture of the person according to the human skeleton key point information, and send the detection results including the stretch index, waving action state, and the sitting, lying or standing posture of the person to the waving for help detection unit; At the same time, according to the human key point coordinates, obtain the stretch index, and then obtain the waving frequency, and send the waving frequency to the waving for help detection unit; The calculation module obtains the stretch index δ based on the human key point information HPS , specifically as follows: 3-2-1) For the human key point information of a given posture, the calculation module constructs a matrix structure as: X = [x1, …, x n ∈ R D×n where D is the dimension of the key points, n is the number of key points, x n is the coordinate of the nth key point, and R is the set of real numbers; 3-2-2) Let be the mean of all key point coordinates in X, and define Then the variance sum δ of the main axis of the human body pose key points HPS is defined as: where U is the projection matrix, and x i is the coordinate of the i-th key point, d is a positive number less than the key point dimension D, and λ j is the covariance matrix is the j-th eigenvalue, and Tr represents the trace of the matrix; 3-2-3) According to the stretch index δ HPS Calculate the waving frequency, specifically as follows: δ HPS is the sum of variances of the main axes of human body pose key points, that is, it reflects the degree of human body pose extension; When the hand swings to both sides of the body, δ HPS reaches its maximum value; when the hand swings to the top of the head and forms a straight line with the body trunk, δ HPS reaches its minimum value. By statistically analyzing the periodic transformation law of δ HPS on the time axis, the waving frequency information during the waving motion is obtained; the obtained waving frequency information is sent to the waving for help detection unit; The waving for help detection unit is used to process the MFCC acoustic features extracted by the feature extraction unit, the sent face image, and the detection results sent by the calculation module, and use the comprehensive weight method to determine whether the person is making a waving for help action.
2. The wave-for-help recognition system according to claim 1, characterized in that, The feature extraction unit includes: a voice feature extraction module, a facial feature preprocessing module, and a human body posture detection module; The voice feature extraction module is used to receive the audio stream sent by the camera, process the continuous speech in the audio stream, obtain the N-dimensional MFCC features, and send them to the waving for help detection unit; The facial feature preprocessing module is used to obtain an image through the camera video stream, detect the human face through machine learning, remove the background and non-face areas, obtain the face point coordinates, normalize the face features of the face image according to the face point coordinates, and send the processed face image to the waving for help detection unit; The human body posture detection module is used to receive the video stream sent by the camera, and after extracting the human skeleton pose information through a deep neural network, obtain the human skeleton key point information and transmit it to the calculation module.
3. The wave-for-help recognition system according to claim 1, characterized in that, The calculation module includes: a waving action detection module and a stretch index calculation module; The waving action detection module is used to perform normalization preprocessing of coordinates and scales according to the human key point information to obtain key point features, judge whether the person in the standing, sitting or lying posture has a waving action state, and obtain the waving amplitude information according to the waving action state; At the same time, according to the posture detection model, output the waving posture confidence of the person corresponding to the standing, sitting or lying waving action state at this moment; And send the corresponding standing, sitting or lying posture information, waving amplitude information and waving posture confidence of the human body to the waving for help detection unit; The stretch index calculation module is used to obtain the stretch index according to the human key point coordinates to reflect the posture stretch of the human body, and calculate the waving frequency according to the stretch index; Send the waving frequency to the waving for help detection unit.
4. The wave-for-help recognition system according to claim 1, wherein The waving for help detection unit includes: a sound event recognition module, a facial expression recognition module, and a waving for help detection module; A voice event recognition module, which is used to perform event detection based on the MFCC features extracted by the feature extraction unit using a deep learning model, and obtain a voice classification result including the classified speech and the confidence information of the current frame voice output; A facial expression recognition module, which is used to extract facial features from the face image sent by the feature extraction unit through a deep learning model, and perform training and inference to obtain a facial expression classification result including the classified facial expression of the human body and the confidence information of the current frame facial expression output; A waving for help detection module, which is used to determine whether a person is performing a waving for help action using the comprehensive weight method based on the voice classification result, the facial expression classification result, and the posture information of the human body standing, sitting, or lying, the waving amplitude information, the waving posture confidence, and the waving frequency sent by the calculation module.
5. The recognition method of a wave-for-help recognition system according to claim 1, characterized in that It includes the following steps: 1) The audio stream sent by the camera goes to the voice feature extraction module, and the video stream is sent to the facial feature preprocessing module and the human body posture detection module respectively; 2-1) The voice feature extraction module receives the audio stream sent by the camera, processes the continuous speech, obtains the MFCC acoustic features, and sends them to the waving for help detection unit; 2-2) The facial feature preprocessing module obtains images from the camera video stream, performs human face detection through machine learning methods, removes the background and non-face regions to obtain the face point coordinates, normalizes the face image according to the face point coordinates, and sends the processed face image to the waving for help detection unit; 2-3) The human body posture detection module receives the video stream sent by the camera, extracts the human body skeleton posture information through a deep neural network, obtains the human body skeleton key point information, and transmits it to the calculation module; 3-1) The calculation module performs normalization preprocessing of coordinates and scales based on the human body key point information to obtain key point features, determines whether the person with the standing, sitting, or lying posture information has a waving action state, and obtains the waving amplitude information according to the waving action state; at the same time, obtains the waving posture confidence of the person corresponding to standing, sitting, or lying at this moment according to the output of the posture detection model; and sends the posture information of the human body corresponding to standing, sitting, or lying, the waving amplitude information, and the waving posture confidence to the waving for help detection unit; 3-2) The calculation module obtains the stretch index δ based on the human key point information HPS , and based on the stretch index δ HPS calculates the waving frequency and sends the waving frequency to the waving for help detection unit; 4) The waving for help detection unit performs event detection based on the MFCC acoustic features extracted by the feature extraction unit using a deep learning model to obtain a voice classification result including the classified speech and the confidence information of the current frame voice output; at the same time, extracts facial features from the face image sent by the feature extraction unit through a deep learning model, and performs training and inference to obtain a facial expression classification result including the classified facial expression of the human body and the confidence information of the current frame facial expression output; 5) The waving for help detection unit determines whether a person is performing a waving for help action using the comprehensive weight method based on the voice classification result, the facial expression classification result, and the posture information of the human body corresponding to standing, sitting, or lying, the waving amplitude information, the waving posture confidence, and the waving frequency sent by the calculation module.
6. The recognition method of a wave-for-help recognition system according to claim 5, characterized in that, In step 2-1), after processing the continuous speech, MFCC acoustic features are obtained, specifically: For the continuous speech of the audio stream, pre-emphasis, framing, and windowing operations are performed in sequence to obtain the preprocessed sound information; For the preprocessed sound information, fast Fourier transform, Mei filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction are performed in sequence, and finally N-dimensional MFCC acoustic features are obtained.
7. The recognition method of a waving-for-help recognition system according to claim 5, characterized in that, The said step 3-1) is specifically: 3-1-1) The calculation module obtains the standing, sitting, or lying posture information of the person at this moment, as well as the current frame waving posture confidence corresponding to standing, sitting, and lying, according to the human key point information through the posture detection model; the posture detection model is a CNN model; 3-1-2) At the same time, according to the human key point information, the included angle data between each joint is obtained, and according to the included angle data between the forearm and the upper arm, and the included angle data between the upper arm and the shoulder in the included angle data between each joint, the waving amplitude information of the human body during the waving behavior is obtained; 3-1-3) The calculation module sends the standing, sitting, or lying posture information, waving amplitude information, and waving posture confidence of the human body to the waving for help detection unit.
8. The recognition method of a waving for help recognition system according to claim 5, characterized in that, The said step 4) is specifically: 4-1) For the sound event classification, specifically: During model training, it includes normalizing the input MFCC features, and at the same time reading the label text corresponding to the expression name for representation to generate a label vector, and binding the MFCC features and the label vector, then performing preprocessing, and then sending it to the deep learning model to learn the parameters to obtain a speech classification model; Among them, the deep learning model is any one of the CNN model, RCNN model, and LSTM model; During model prediction, the input MFCC features are normalized, then preprocessed, and then sent to the speech classification model for prediction, and the obtained data is post-processed to obtain a speech classification result including the classified speech and the confidence information of the corresponding current frame sound output; 4-2) For the facial expression classification, specifically: The deep learning model extracts facial features from the face image; Among them, the deep learning model is any one of the convolutional neural network, deep belief network, deep autoencoder, and recurrent neural network; After completing the facial feature extraction, the convolutional neural network is used to classify the facial expressions to obtain a facial expression classification result including the classified facial expressions of the human body and the confidence information of the corresponding current frame facial expression output.
9. The recognition method of a wave-for-help recognition system according to claim 5, characterized in that, The said step 5) is specifically: (1) The waving for help detection unit uses the comprehensive weight method to obtain the global strategy of waving for help according to the speech classification result including the classified speech and the confidence information of the corresponding current frame sound output, the facial expression classification result including the classified facial expressions of the human body and the confidence information of the corresponding current frame facial expression output, the standing, sitting, or lying posture information, waving amplitude information, waving posture confidence, and waving frequency of the human body sent by the calculation module: Among them, is the confidence level of waving for help at the current moment t, is the confidence level of the waving-for-help gesture at the current moment t, is the confidence level of voice event detection at the current moment t, is the confidence level of facial expression at the current moment t; w pose , w sound , w face are the weights of the waving gesture, the voice event, and the facial expression respectively; Among them: w pose +w sound +w face = 1 w pose , w sound , w face For preset assignment, w pose > w sound , w pose > w face ; (2) Obtained by exponentially weighted moving average That is: where β represents a weighting coefficient, is the confidence level of the waving for help gesture in the current frame, is the confidence information output by the sound detection in the current frame, is the confidence information output by the facial expression in the current frame; (3) Confidence of the waving for help gesture in the current frame It is obtained by comprehensively considering the confidence of the waving gesture in the current frame, the waving amplitude information, the waving frequency information, and the gesture bonus information, that is: Among them, is the confidence level of the waving gesture in the current frame, is the waving amplitude information in the current frame, is the waving frequency information in the current frame, is the gesture bonus coefficient; (4) Current frame wave amplitude information That is: Among them, is the detected waving amplitude, is the maximum preset waving amplitude, is a value between 0 and 1. The larger the waving amplitude, the higher this value is; (5) Current frame wave frequency information Specifically: wherein, is the detected waving frequency, is the maximum preset waving frequency, is a value between 0 and 1, and the larger the waving frequency, the higher this value; (6)Finally, it is judged against a preset threshold. If it is greater than the threshold, the current frame is in the waving for help state; if it is in the waving for help state for n consecutive frames, a waving for help alarm is triggered.
Citation Information
Patent Citations
Gesture help-seeking recognition method and device and storage medium
CN112299172A
KR20200084451A
Cited By
Unmanned aerial vehicle-mounted real-time action detection method and system for search and rescue scene
CN118196900A