A Cross-modal Sound Source Localization and Gesture Recognition Method Combining Audio and Visual Information

Through the cross-modal sound source positioning and gesture recognition method of audio and visual information, the gestures of the initiator of the command are accurately positioned and identified in complex scenarios in a complex scenario using audio and visual information, the problem of difficulty in accurately positioning and identifying gesture recognition in complex scenarios in the prior art is solved, and the accuracy of gesture recognition in high dynamic and high noise environments is achieved.

CN116631059BActive Publication Date: 2025-06-10BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310600840.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2025-06-10
Estimated Expiration
2043-05-25

AI Technical Summary

Technical Problem

In complex large-scale outdoor scenarios, existing gesture recognition technologies are difficult to accurately locate and identify gesture commands from the initiator of the command, especially in the context of messy backgrounds.

Method used

The cross-modal sound source positioning and gesture recognition method of audio-visual joint use is adopted to accurately locate the position of the initiator of the command through audio information, and to identify the gesture of the initiator of the command using visual information. Specific steps include: establishing a spatial polar coordinate system, pre-processing of audio information, using convolutional neural network for sound source positioning, using YOLOv7 for human body detection, using AlphaPose to extract bone information of gestures, and using spatiotemporal graph convolutional network for gesture recognition.

Benefits of technology

In complex scenarios, the precise positioning and gesture recognition of the command initiator are realized, which improves the accuracy and robustness of gesture recognition, and enables the intelligent robot to effectively complete the corresponding instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631059B_ABST
    Figure CN116631059B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for cross-modal sound source localization and gesture recognition combining sound and image. The specific implementation process of this method is as follows: Step 1. Establish a spatial polar coordinate system; Step 2. Divide the three-dimensional space into different sub-spaces; Step 3. Preprocess the audio information; Step 4. Convolutional neural network model; Step 5. Use YOLOv7 for human detection; Step 6. Use AlphaPose to extract the skeletal information of gestures; Step 7. Use a spatio-temporal graph convolutional network for gesture recognition; accurately locate the position of the instruction initiator through the audio information, and then recognize the gesture of the instruction initiator through the visual information, so as to complete the corresponding instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an identification method, in particular to a technology for cross-modal sound source localization and gesture recognition based on audio-visual combination, belonging to the field of computer vision technology. Background Art

[0002] Gesture recognition is one of the representative tasks in computer vision. Precise perception and recognition of human gestures are important prerequisites for intelligent interaction and human-machine collaboration. In recent years, it has become a widely concerned research field. For example, in application fields such as behavior analysis, intelligent driving, and medical control, the research on body language interaction is of great significance. However, for large-scale outdoor complex environments with high dynamics and strong confrontation, such as battlefields, current methods are difficult to accurately locate the position of the instruction initiator and recognize their gestures, so as to execute corresponding instructions.

[0003] To address gesture recognition in complex environments, traditional methods only use visual information and construct long-term associations through recurrent neural networks. They can obtain more behavioral features by using a global context storage unit to focus on information nodes in each frame. There are also some methods aiming to aggregate the features of spatio-temporal image regions using attention mechanisms to effectively remove the influence of noise and improve the recognition accuracy. However, these methods still cannot quickly and effectively locate the key area (the gesture of the instruction initiator) in complex environments, which is a major challenge in gesture recognition tasks in large-scale outdoor complex environments. The cross-modal sound source localization and gesture recognition method based on audio-visual combination aims to adopt multi-modal data to accurately locate the instruction initiator through audio information, so as to solve the problem that intelligent robots cannot recognize effective gestures due to cluttered backgrounds in complex scenarios, and then effectively recognize the gestures of the instruction initiator through visual information, thereby enhancing the accuracy of gesture recognition in high-dynamics and strong-confrontation outdoor complex scenarios. In addition, traditional sound source localization methods mainly rely on signal processing and mathematical algorithms, such as beamforming, cross-correlation, least squares and other methods. These methods can achieve sound source localization to a certain extent, but their accuracy and robustness are affected by factors such as environmental noise.

[0004] To solve these problems, the present invention patent application discloses a cross-modal sound source localization and gesture recognition method based on audio-visual combination. This method combines sound source information and visual information. First, it uses a convolutional neural network to extract features from sound signals, which improves the accuracy and robustness of sound source localization to a certain extent and accurately locates the instruction initiator. Then, considering spatio-temporal correlation at the same time, it uses the joint information of gestures to recognize the gestures of the instruction initiator through a spatio-temporal graph convolutional neural network. Finally, it effectively realizes gesture recognition in large-scale outdoor complex scenarios, improves the accuracy of gesture recognition in uncertain and high-dynamics complex scenarios, and enables intelligent robots to effectively complete corresponding instructions. Summary of the Invention

[0005] In order to solve the problem that it is difficult for machines to determine the position of the commander in complex scenarios and unable to accurately recognize their gesture commands, the present invention aims to achieve accurate recognition and tracking of gesture commands by combining audio-visual information, locate the position of the command initiator from the sound, and recognize the posture of the command initiator from the visual information, so as to meet the requirements of different environments and scenarios, and has the advantages of high efficiency, reliability, and high robustness.

[0006] In order to achieve accurate recognition of gesture commands in complex environments, the technical solution adopted by the present invention is a cross-modal sound source localization and gesture recognition method combining sound and image. The position of the command initiator is accurately located through audio information, and then the gesture of the command initiator is recognized through visual information, so as to complete the corresponding command. The specific implementation process of this method is as follows:

[0007] Step 1. Establish a spatial polar coordinate system;

[0008] Taking the intelligent robot as the center, establish a spatial polar coordinate system This represents the position of the sound source relative to the intelligent robot. Among them, the initial position of the center of the intelligent robot is represented as (0, 0, 0), which represents the origin of the spatial polar coordinate system; r represents the distance from the sound source to the center of the intelligent robot, represents the azimuth angle between the sound source and the center of the intelligent robot, and θ represents the elevation angle between the sound source and the center of the intelligent robot.

[0009] Step 2. Divide the three-dimensional space into different sub-spaces;

[0010] Within the range of the polar radius R in the spatial polar coordinate system, the three-dimensional space is divided into Z sub-spaces of equal size and non-overlapping, that is, each sub-space is independent of each other, and each sub-space has a unique three-dimensional coordinate representation.

[0011] Step 3. Preprocess the audio information;

[0012] The intelligent robot is equipped with multiple microphones, and the microphones are placed on the same horizontal plane. After receiving the sound recorded by the multi-microphones, preprocess the audio information, convert the time-domain signal into a frequency-domain signal, and then extract the features of the frequency-domain signal, and change the extracted features into a form suitable for processing by the convolutional neural network.

[0013] Step 4. Convolutional neural network model;

[0014] On the preprocessed sound signal, use the convolutional neural network model for training. The convolutional neural network model consists of 4 convolutional layers and 4 pooling layers. Among them, the convolutional layer is used to perform convolutional operations on different local matrices and the convolutional kernel matrix of the input image to extract image features; the pooling layer is used to compress the features extracted by the convolutional layer and reduce the feature dimension.

[0015] The transformed three-dimensional data is used as the input of the convolutional neural network, and the output of the convolutional neural network is used as the training feature, denoted as f'. Then, classification is performed based on the training feature to determine the subspace where the sound source is located, realizing sound source localization. It is expressed as:

[0016] g z = classify(f')

[0017] where classify(·) represents the classifier function, and g z represents the subspace where the target sound source is predicted. According to the subspace position, the intelligent robot moves in front of the instruction initiator, and the direction angle of the intelligent robot's rotation is The moving distance is r·cosθ.

[0018] Step 5. Use YOLOv7 for human detection;

[0019] After obtaining the position of the instruction initiator based on the audio information, the intelligent robot moves to the side of the instruction initiator, uses YOLOv7 to process the visual information, and detects the human target;

[0020] Step 6. Use AlphaPose to extract the skeletal information of the gesture;

[0021] After detecting the human target through YOLOv7, use AlphaPose to extract the hand joint points of the human body to obtain the joint point information of the instruction initiator's hand. The joint point information includes the positions and postures of key points such as fingers, wrists, elbows, and shoulders, providing important features for subsequent gesture recognition.

[0022] Step 7. Use the spatio-temporal graph convolutional network for gesture recognition;

[0023] Use the spatio-temporal graph convolutional network to model and process the gesture joint point information extracted by AlphaPose, thereby identifying the category of the gesture. The spatio-temporal graph convolutional network uses a multi-layer graph convolutional neural network and performs convolutional operations in space and time to extract the spatial and temporal features of the gesture, having high accuracy and robustness in gesture recognition.

[0024] Furthermore, in the said Step 2, the three-dimensional space sound source localization is characterized through probability distribution, transforming the linear regression problem into a non-linear classification problem. By extracting position features from the sound source signals received by the array, different classifiers are used to determine which subspace the sound source belongs to.

[0025] Furthermore, the specific implementation manner of the said Step 3 is as follows:

[0026] (1) Pre-emphasis;

[0027] Perform high-pass filtering on the original audio signal to enhance the energy of high-frequency signals and reduce the energy of low-frequency signals; the process of pre-emphasis is expressed as follows:

[0028] s'(n) = s(n) - αs(n - 1)

[0029] Where n represents the sampling point index of the time-domain signal, s(n) and s(n - 1) are the input signals, representing the sampling values of the time-domain signal at two adjacent sampling points, s'(n) represents the pre-emphasized signal, and α is the pre-emphasis coefficient.

[0030] (2) Frame segmentation;

[0031] Divide the audio signal into several frames, and the length of each frame is a fixed time window. The process of frame segmentation is expressed as follows:

[0032] s m (n) = s'(n)·w(n - mR)

[0033] Where s m (n) represents the m-th frame signal, w(n) represents the window function, the Hamming window function is used, and R represents the frame offset, with a value of half of the frame length.

[0034] (3) STFT transformation;

[0035] Perform STFT transformation on the sound signal to convert the time-domain signal into a frequency-domain signal; specifically, segment the time-domain signal, the length of each segment of the time-domain signal is the window length, there is an overlap between two adjacent segments of the time-domain signal, and perform Fourier transform on each segment of the time-domain signal to obtain its frequency-domain representation, that is, decompose the time-domain signal into several frequency bands in the frequency domain. The process of STFT transformation is expressed by the formula as follows:

[0036]

[0037] Where S m (k) represents the frequency-domain representation of the m-th frame, N represents the number of sampling points per frame, and k = n - mR. In addition, j represents the imaginary unit and e represents the natural constant.

[0038] (4) Feature extraction;

[0039] Feature extraction extracts representative features from the frequency-domain signal to provide reliable input data for the subsequent sound source localization task. The MFCC method is used to extract the sound features at different frequencies.

[0040] First, pass the frequency-domain signal through the Mel filter bank to obtain the coefficients at different frequencies, and the formula is expressed as follows:

[0041]

[0042]

[0043] Among them, H m (k) represents the k-th filter of the Mel filter bank, h(m) represents the frequency corresponding to the m-th Mel frequency, and E m represents the Mel frequency coefficient of the m-th frame.

[0044] Perform discrete cosine transform DCT on the filtered Mel frequency coefficients: Perform DCT transformation on the Mel frequency coefficients after taking the logarithm to obtain a set of MFCC coefficients. The MFCC coefficients represent the speech features of the audio signal and are expressed as follows:

[0045]

[0046] Among them, f a represents the a-th MFCC coefficient, and M represents the number of MFCC coefficients. The MFCC coefficients are used as the feature representation for the sound source localization task.

[0047] (5) Format conversion;

[0048] Convert the data extracted by feature extraction into a form suitable for processing by a convolutional neural network, that is, convert the frequency-domain signal into an image form, so that the data is regarded as two-dimensional image data. By performing spectral analysis on the signal, extract the amplitude spectrum or phase spectrum of the signal in the frequency domain, and then use the two-dimensional fast Fourier transform FFT to convert it into an image form. The amplitude spectrum Q m (k) and the phase spectrum P m (k) are calculated as follows:

[0049] Q m (k) = |S m (k)|

[0050] P m (k) = arg{S m (k)}

[0051] Among them, |·| represents the modulus operation, and arg{·} represents the argument operation. Regard the amplitude spectrum or phase spectrum as two-dimensional image data, and obtain a three-dimensional data set by stacking the amplitude spectrum or phase spectrum at different time steps. The first dimension represents time, the second dimension represents frequency, and the third dimension represents amplitude or phase, which is used as the input data of the convolutional neural network.

[0052] Furthermore, the implementation process of step 5 is as follows:

[0053] (1) Data input;

[0054] Obtain video data as input, split the video data into T frames, input the image frames into the YOLOv7 network, perform multiple convolution and pooling operations on the image frames, and obtain a series of feature maps.

[0055] (2) Anchor box and feature map processing;

[0056] YOLOv7 uses anchor boxes to predict the position and size of the target; input the feature maps obtained in the data input stage into the convolutional layer and pooling layer for processing to obtain feature maps of different scales.

[0057] (3) Target prediction;

[0058] In the feature map, each pixel point corresponds to an anchor box. By classifying and regressing each anchor box, it is predicted whether there is a target object in each box and its position and size are estimated. The objective function is expressed as follows:

[0059] L = L cls + λ coord L coord + λ obj L obj + λ noobj L noobj

[0060] Among them, L cls represents the classification loss, L coord represents the position loss, L obj represents the loss of the existence of the target object, L noobj represents the loss of the non-existence of the target object. λ coord 、λ obj 、λ noobj are weight parameters. Specifically, the classification loss is calculated using the cross-entropy loss function, the position loss is calculated using the mean squared error loss function, and the loss of the existence of the target object and the loss of the non-existence of the target object are respectively expressed as:

[0061]

[0062]

[0063] Among them, G represents the size of the feature map, B represents the number of anchor boxes predicted by each pixel point, represents whether there is a target object in the j-th anchor box on the i-th pixel point, represents whether the j-th anchor box on the i-th pixel point does not contain the target object. x i 、y i represent the center point coordinates of the predicted box, represents the center point coordinates of the actual target object. Ci Indicates the confidence level of whether there is a target object within the predicted bounding box. Indicates the confidence level of the actual existence of the target object.

[0064] Redundant predicted bounding boxes are removed through non-maximum suppression (NMS), and the prediction result is denoted as γ = {b 1 , b 2 , …, b T}, where b t = (x t , y t , w t , h t ) represents the center coordinates (x t , y t ) and width and height (w t , h t ) of the instruction initiator in the t-th frame image, thereby determining the position of the instruction initiator in each frame image.

[0065] Furthermore, in step 6, the instruction initiator in each frame image is extracted, and hand keypoint detection is performed on it through AlphaPose. After AlphaPose, for the t-th frame image, the keypoint information corresponding to the hand of the instruction initiator is denoted as U t , where represents the coordinate of the d-th keypoint of the hand of the instruction initiator in the t-th frame image. The hand keypoint information of the instruction initiator in each frame image is concatenated to obtain the keypoint information of the entire gesture action, denoted as U = {U 1 , U 2 , …, U T}.

[0066] Furthermore, in step 7, after the spatio-temporal graph convolutional network obtains the keypoint information of the hand of the instruction initiator, the keypoint information U t of the hand of the instruction initiator in each frame is regarded as a graph, each keypoint is used as a graph node, and the edges between the graph nodes represent the connection relationships between the keypoints. The coordinates and time information of each node are concatenated to obtain a three-dimensional tensor where T represents the number of time steps, D represents the number of key points, and 3 represents the feature dimension of each node. The spatio-temporal graph convolutional neural network is used to perform a convolution operation on this three-dimensional tensor to obtain a new three-dimensional tensor where F represents the number of feature channels. This three-dimensional tensor is regarded as the feature representation of the gesture, and global pooling or a convolutional neural network is used for classification or regression tasks. By performing a weighted sum on the feature representations of the gestures of the instruction initiator in all frames, an overall gesture representation is obtained, and it is transformed into a probability distribution through the softmax function, that is, the probability that the gesture belongs to each category is obtained, thereby obtaining the final gesture recognition result.

[0067] In complex large-scale outdoor scenarios, due to the influence of factors such as noise, existing gesture recognition technologies are difficult to accurately locate and recognize the gesture commands of the instruction initiator. To solve this problem, the present invention uses audio information for sound source localization and visual information for gesture recognition, and through the combination of sound and image, accurately locates and recognizes the gestures of the instruction initiator in complex scenarios. Brief Description of the Drawings

[0068] The following is a further detailed description of the specific implementation manner of the present invention, such as the network structure involved, with reference to the accompanying drawings and through the description of the embodiments, so as to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0069] The present invention will be described in more detail below in conjunction with the accompanying drawings and embodiments.

[0070] Figure 1 It is a schematic flowchart of a cross-modal sound source localization and gesture recognition method combining sound and image according to the present invention;

[0071] Figure 2 It is a flowchart of the preprocessing of audio information according to the present invention;

[0072] Figure 3 It is a structure parameter diagram of a convolutional network model for sound source localization according to the present invention;

[0073] Figure 4 It is a structure diagram of the spatio-temporal graph convolutional layer according to the present invention. Detailed Description of the Preferred Embodiments

[0074] The present invention will be described in detail below in conjunction with the accompanying drawings and embodiments.

[0075] The technical solution adopted by the present invention is a cross-modal sound source localization and gesture recognition method combining sound and image. First, the position of the instruction initiator is accurately located through audio information, and then the gesture of the instruction initiator is recognized through visual information, so as to complete the corresponding instruction.

[0076] S1. Establish a spatial polar coordinate system;

[0077] Taking the intelligent robot as the center, establish a spatial polar coordinate system This is used to represent the position of the sound source relative to the intelligent robot. Among them, the initial position of the center of the intelligent robot is represented as (0, 0, 0), which represents the origin of the spatial polar coordinate system; r represents the distance from the sound source to the center of the intelligent robot, represents the azimuth angle between the sound source and the center of the intelligent robot, and θ represents the elevation angle between the sound source and the center of the intelligent robot.

[0078] S2. Divide the three-dimensional space into different subspaces;

[0079] In the range where the polar radius is R in the spatial polar coordinate system, the three-dimensional space is divided into Z subspaces of equal size and non-overlapping. That is, each subspace is independent, and each subspace has a unique three-dimensional coordinate representation. If the subspace is smaller, the number of them will be more, that is, the larger the value of Z, the higher the classification complexity, and at the same time, the higher the positioning accuracy. Considering that the subspace is small enough, the problem of sound source localization in three-dimensional space can be processed through probability distribution, thus transforming the linear regression problem into a non-linear classification problem to reduce the computational amount. By extracting the position features from the sound source signals received by the array, different classifiers can be used to determine which subspace the sound source belongs to.

[0080] S3. Audio information preprocessing;

[0081] In the present invention, the intelligent robot is equipped with multiple microphones. Before collecting audio information, it is necessary to determine the arrangement method of the microphones. Generally, the arrangement of the microphone array will affect the accuracy of sound source localization. Place the microphones on the same horizontal plane, and the distance between adjacent microphones is 10 cm to ensure the acquisition effect and accuracy of the sound signals.

[0082] After receiving the sounds recorded by multiple microphones, preprocess the audio information. First, convert the time-domain signal into a frequency-domain signal, then extract the features of the frequency-domain signal, and finally transform the features into a form suitable for processing by a convolutional neural network. The main process is as follows:

[0083] (1) Pre-emphasis;

[0084] In the present invention, high-pass filtering is performed on the original signal to enhance the energy of the high-frequency signal and reduce the energy of the low-frequency signal; the process of pre-emphasis is expressed as follows:

[0085] s'(n) = s(n) - αs(n - 1)

[0086] where s(n) represents the input signal, s'(n) represents the pre-emphasized signal, and α is the pre-emphasis coefficient.

[0087] (2) Frame segmentation;

[0088] Divide the signal into several frames, and the length of each frame is a fixed time window. The process of frame segmentation is expressed as follows:

[0089] s m (n) = s'(n)·w(n - mR)

[0090] where s m (n) represents the m-th frame signal, w(n) represents the window function, the Hamming window function is used, and R represents the frame offset, and its value is half of the frame length.

[0091] (4) STFT transformation;

[0092] After pre-emphasis and framing operations, the present invention performs STFT (Short-time Fourier Transform) transformation on the sound signal to convert the time-domain signal into a frequency-domain signal. STFT is an analysis method based on the Fourier transform. It divides a long-time signal into several short time segments. The signal within each short time segment is regarded as stationary and undergoes Fourier transform. STFT transformation can convert the time-domain signal into a frequency-domain signal and extract the sound features at different frequencies, which is an important basis for subsequent feature extraction and sound source localization.

[0093] Specifically, the present invention segments the time-domain signal. The length of each segment is the window length, and there is a certain overlap between adjacent segments. Fourier transform is performed on each segment of the signal to obtain its frequency-domain representation, that is, the signal is decomposed into several frequency bands in the frequency domain. The process of STFT transformation is expressed by the following formula:

[0094]

[0095] where S m (k) represents the frequency-domain representation of the m-th frame, N represents the number of sampling points per frame, and k = n - mR.

[0096] (4) Feature extraction;

[0097] In the process of sound source localization, feature extraction from the frequency-domain signal is a very important step. Feature extraction can extract representative features from the frequency-domain signal and provide reliable input data for subsequent sound source localization tasks. The present invention uses the MFCC (Mel Frequency Cepstral Coefficients) method to extract the sound features at different frequencies.

[0098] First, the frequency-domain signal is passed through the Mel filter bank to obtain the coefficients at different frequencies. The formula is expressed as follows:

[0099]

[0100]

[0101] where H m (k) represents the m-th filter of the Mel filter bank, h(m) represents the frequency corresponding to the m-th Mel frequency, and E m represents the Mel frequency coefficient of the m-th frame.

[0102] Next, perform a discrete cosine transform (DCT) on the filtered Mel frequency coefficients. DCT is a technique that converts a time-domain signal into a frequency-domain signal, which can convert the Mel frequency coefficients into MFCC coefficients. Specifically, perform a DCT transformation on the Mel frequency coefficients after taking the logarithm to obtain a set of MFCC coefficients. These coefficients can represent the speech features of the audio signal, and the formula is as follows:

[0103]

[0104] where f a represents the a-th MFCC coefficient, and M represents the number of MFCC coefficients. The present invention uses the MFCC coefficients as the feature representation for the sound source localization task.

[0105] (5) Format conversion;

[0106] Convert the extracted feature data into a form suitable for processing by a convolutional neural network, that is, convert the frequency-domain signal into an image form, so that the data can be regarded as two-dimensional image data. The present invention performs spectral analysis on the signal, extracts the amplitude spectrum or phase spectrum of the signal in the frequency domain, and then uses a two-dimensional fast Fourier transform (FFT) to convert it into an image form. The amplitude spectrum Q m (k) and the phase spectrum P m (k) are calculated as follows:

[0107] Q m (k) = |S m (k)|

[0108] P m (k) = arg{S m (k)}

[0109] where |·| represents the modulus operation, and arg{·} represents the argument operation. The present invention regards the amplitude spectrum or phase spectrum as two-dimensional image data, and by stacking the amplitude spectrum or phase spectrum at different time steps, a three-dimensional data set is obtained, where the first dimension represents time, the second dimension represents frequency, and the third dimension represents amplitude or phase, so that it can be used as the input data of the convolutional neural network.

[0110] S4. Convolutional neural network model;

[0111] Train using a convolutional neural network model on the preprocessed sound signal. The convolutional neural network model constructed by the present invention consists of 4 convolutional layers and 4 pooling layers. Among them, the convolutional layer is mainly used to perform a convolutional operation on different local matrices and the convolutional kernel matrix of the input image to extract image features; the pooling layer is mainly used to compress the features extracted by the convolutional layer, reduce the feature dimension, thereby reducing the amount of calculation, preventing overfitting, and improving the calculation speed.

[0112] The three-dimensional data after format transformation is used as the input of the convolutional neural network, and the output of the convolutional neural network is used as the training feature, denoted as f'. Then, classification is performed based on the training feature to determine the subspace where the sound source is located, and sound source localization is achieved. This process is expressed as:

[0113] g z = classify(f')

[0114] where classify(·) represents the classifier function, and g z represents the subspace where the target sound source is predicted. Finally, according to the position of the subspace, the intelligent robot moves in front of the instruction initiator. Among them, the direction angle of rotation of the intelligent robot is The moving distance is r·cosθ.

[0115] S5. Use YOLOv7 for human detection;

[0116] YOLOv7 is an efficient object detection algorithm that can quickly and accurately locate objects in images. In the present invention, we use YOLOv7 to detect humans in videos. YOLOv7 can process a large number of images in a short time and accurately locate humans. Through YOLOv7, we can obtain the position information of the instruction initiator, preparing for subsequent gesture recognition. Its main process is as follows:

[0117] (1) Data input;

[0118] Obtain video data as input, split the video data into T frames, input the image frames into the YOLOv7 network, and perform multiple convolution and pooling operations on the image frames to obtain a series of feature maps.

[0119] (2) Anchor box and feature map processing;

[0120] In order to detect target objects of different sizes and ratios, YOLOv7 uses Anchor boxes to predict the position and size of the target. Anchor boxes are a set of boxes with fixed sizes and ratios that cover different regions of the input image. In addition, the feature maps obtained in the data input stage are input into convolutional layers and pooling layers for processing to obtain feature maps of different scales.

[0121] (3) Target prediction;

[0122] In the feature map, each pixel point corresponds to an Anchor box. By classifying and regressing each Anchor box, it is predicted whether there is a target object in each box, and its position and size are estimated. The objective function is expressed as follows:

[0123] L = L cls + λcoord L coord + λ obj L obj + λ noobj L noobj

[0124] Among them, L cls represents the classification loss, L coord represents the position loss, L obj represents the loss of the presence of the target object, and L noobj represents the loss of the absence of the target object. λ coord , λ obj , λ boobj are weight parameters. Specifically, the classification loss is calculated using the cross-entropy loss function, the position loss is calculated using the mean squared error loss function, and the losses of the presence and absence of the target object are respectively expressed as:

[0125]

[0126]

[0127] Among them, G represents the size of the feature map, B represents the number of Anchor boxes predicted for each pixel point, indicates whether there is a target object in the j-th Anchor box at the i-th pixel point, indicates whether the j-th Anchor box at the i-th pixel point does not contain a target object. x i , y i represent the center coordinates of the predicted box, represents the center coordinates of the actual target object. C i represents the confidence level of whether there is a target object within the predicted box, represents the confidence level of the presence of the actual target object.

[0128] In addition, in the prediction results, there may be a situation where multiple predicted boxes cover the same target object. Therefore, we need to eliminate redundant predicted boxes through non-maximum suppression (NMS) and only retain the best prediction result. The prediction result is denoted as β = {b 1 , b 2 , …, b T}, where b t = (x t , y t , w t , h t ) represents the center coordinates (x t , y t ) and width and height (w t , h t), thereby determining the position of the instruction initiator in each frame of the image.

[0129] S6. Use AlphaPose to extract the skeletal information of the gesture;

[0130] AlphaPose is a human pose estimation algorithm that can quickly and accurately extract human joint point information. In the present invention, AlphaPose is used to extract the hand joint points of the human body detected by YOLOv7, and the joint point information of the hand of the instruction initiator is obtained. These joint point information include the positions and postures of key points such as fingers, wrists, elbows, and shoulders, providing important features for subsequent gesture recognition.

[0131] In the present invention, the instruction initiator in each frame of the image is extracted, and its hand joint points are detected by AlphaPose. After passing through AlphaPose, for the t-th frame of the image, the joint point information corresponding to the hand of the instruction initiator is denoted as U t , where represents the coordinate of the d-th joint point of the hand of the instruction initiator in the t-th frame of the image. We can splice the hand joint point information of the instruction initiator in each frame of the image to obtain the joint point information of the entire gesture action, denoted as U = {U 1 , U 2 , …, U T}.

[0132] S7. Use a spatio-temporal graph convolutional network for gesture recognition;

[0133] The spatio-temporal graph convolutional network is a video action recognition algorithm based on the graph convolutional neural network, which can effectively utilize time and space information to classify the actions of the video. In the present invention, the spatio-temporal graph convolutional network is used to model and process the gesture joint point information extracted by AlphaPose, so as to recognize the category of the gesture. The spatio-temporal graph convolutional network uses multiple layers of graph convolutional neural networks and performs convolutional operations in space and time, thereby extracting the spatial and temporal features of the gesture, and has high accuracy and robustness in gesture recognition.

[0134] Specifically, after obtaining the joint point information of the hand of the instruction initiator, the joint point information U t of the hand of the instruction initiator in each frame is regarded as a graph, each joint point is used as a graph node, and the edges between the graph nodes represent the connection relationships between the joint points. In the present invention, the coordinates and time information of each node are spliced together to obtain a three-dimensional tensor where T represents the number of time steps, D represents the number of key points, and 3 represents the feature dimension of each node (including xy coordinates and time information). In the present invention, a spatio-temporal graph convolutional neural network is used to perform convolutional operations on this three-dimensional tensor to obtain a new three-dimensional tensor Among them, F represents the number of feature channels. In the present invention, this three-dimensional tensor is regarded as the feature representation of the gesture, and global pooling or a convolutional neural network is used for classification or regression tasks.

[0135] Finally, an overall gesture representation is obtained by weighted summation of the feature representations of the gestures of the instruction initiator in all frames, and it is transformed into a probability distribution through the softmax function, that is, the probability that the gesture belongs to each category is obtained, and thus the final gesture recognition result is obtained.

[0136] The schematic flowchart of Embodiment 1 is as Figure 1 shown.

[0137] The flowchart of audio information preprocessing in Embodiment 2 is as Figure 2 shown. In the present invention, audio information preprocessing needs to be performed before inputting the audio information into the convolutional neural network to convert the format of the audio information into a format suitable for input to the convolutional neural network. The specific process is as follows:

[0138] Pre-emphasis: Pre-emphasis is a high-pass filter that can strengthen high-frequency signals and weaken low-frequency signals, making the audio signal more stable in subsequent processing.

[0139] Framing: Framing is to divide the audio signal into several segments, each segment is called a frame for subsequent processing. When framing, the frame length and frame shift parameters need to be set. Usually, the frame length is selected to be 20 - 40 ms, and the frame shift is 10 - 20 ms.

[0140] STFT conversion: Fourier transform is performed on each frame to obtain the frequency-domain signal of that frame. Usually, the short-time Fourier transform (STFT) is used for calculation.

[0141] Feature extraction: The frequency-domain signal is filtered through a Mel filter bank. The Mel filter bank is a group of filters for filtering the signal in the frequency domain, which can convert the frequency-domain signal into Mel frequency coefficients. In the present invention, 40 Mel filters are selected. The discrete cosine transform (DCT) is performed on the filtered Mel frequency coefficients. DCT can convert the Mel frequency coefficients into MFCC coefficients. In the present invention, we select 13 MFCC coefficients as the features of the audio information.

[0142] Format conversion: The frequency-domain signal features are converted into an image form so that the data can be regarded as two-dimensional image data. The general conversion method is to extract the amplitude spectrum or phase spectrum of the signal in the frequency domain, and then use the two-dimensional fast Fourier transform (FFT) to convert it into an image form.

[0143] The structural parameter diagram of the convolutional network model for sound source localization in Embodiment 3 is as Figure 3As shown. In the present invention, kernel_size represents the convolutional kernel size, stride represents the stride, pad represents the edge padding parameter, pooling represents the pooling method, dropout represents the neuron inefficiency rate, iterations represents the number of iterations, and batch_size represents the batch size.

[0144] The structure diagram of the spatio-temporal graph convolutional layer in Embodiment 4 is as Figure 4 shown. In the present invention, the entire model is trained from beginning to end in a backpropagation manner. Specifically, spatio-temporal graph convolution is divided into spatial graph convolution and temporal graph convolution, where spatial graph convolution is the core part, and temporal graph convolution includes two BN layers, one ReLU activation layer, one Dropout layer, and one convolutional layer. One spatial graph convolution plus one temporal convolution is one layer, for a total of 10 layers, but the first layer does not have a residual structure.

Claims

1. An audio-visual combined cross-modal sound source localization and gesture recognition method, characterized in that, the specific implementation process of this method is as follows: Step 1. Establish a spatial polar coordinate system; Establish a spatial polar coordinate system centered on the intelligent robot This is used to represent the position of the sound source relative to the location of the intelligent robot; among them, the initial position of the center of the intelligent robot is represented as (0, 0, 0), which represents the origin of the spatial polar coordinate system; r represents the distance from the sound source to the center of the intelligent robot represents the direction angle between the sound source and the center of the intelligent robot, and θ represents the elevation angle between the sound source and the center of the intelligent robot Step 2. Divide the three-dimensional space into different sub-spaces; Within the range of the polar radius R in the spatial polar coordinate system, the three-dimensional space is divided into Z sub-spaces of equal size and non-overlapping, that is, each sub-space is independent, and each sub-space has a unique three-dimensional coordinate representation; Step 3. Preprocess the audio information; The intelligent robot is equipped with multiple microphones, and the microphones are placed on the same horizontal plane; after receiving the sound recorded by the multi-microphones, preprocess the audio information, convert the time-domain signal into a frequency-domain signal, then extract the features of the frequency-domain signal, and change the extracted features into a form suitable for processing by a convolutional neural network; Step 4. Convolutional neural network model; On the preprocessed sound signal, use a convolutional neural network model for training; the convolutional neural network model consists of 4 convolutional layers and 4 pooling layers; use the three-dimensional data after format transformation as the input of the convolutional neural network, and the output of the convolutional neural network as the training feature, denoted as f'; classify according to the training feature to determine the sub-space where the sound source is located, and realize sound source localization; expressed as: g z = classify(f′) Among them, classify(·) represents the classifier function, and g z represents the subspace where the target sound source is predicted to be located; according to the subspace position, the intelligent robot moves in front of the instruction initiator, and the direction angle of rotation of the intelligent robot is The moving distance is r·cosθ; Step 5. Use YOLOv7 for human detection; After the intelligent robot obtains the position of the instruction initiator according to the audio information, it moves to the side of the instruction initiator, uses YOLOv7 to process the visual information, and detects the human target; Step 6. Use AlphaPose to extract the skeletal information of the gesture; After detecting the human target through YOLOv7, use AlphaPose to extract the hand joint points of the human body to obtain the joint point information of the hand of the instruction initiator; the joint point information includes the positions and postures of key points such as fingers, wrists, elbows, and shoulders, providing important features for subsequent gesture recognition; Step 7. Use a spatio-temporal graph convolutional network for gesture recognition; Use a spatio-temporal graph convolutional network to model and process the gesture joint point information extracted by AlphaPose, so as to recognize the category of the gesture; the spatio-temporal graph convolutional network uses a multi-layer graph convolutional neural network and performs convolutional operations in space and time to extract the spatial and temporal features of the gesture.

2. The audio-visual combined cross-modal sound source localization and gesture recognition method according to claim 1, characterized in that, in the step 2, the three-dimensional space sound source localization is characterized by a probability distribution, turning a linear regression problem into a non-linear classification problem; extracting position features from the sound source signals received by the array, and using different classifiers to determine which sub-space the sound source belongs to.

3. The audio-visual combined cross-modal sound source localization and gesture recognition method according to claim 1, characterized in that, the specific implementation manner of the step 3 is as follows: (1) Pre-emphasis; Perform high-pass filtering on the original audio signal, expressed as follows: s′(n) = s(n) - αs(n - 1) Among them, n represents the sampling point index of the time-domain signal, s(n) and s(n - 1) are the input signals, representing the sampling values of the time-domain signal at two adjacent sampling points, s'(n) represents the pre-emphasized signal, and α is the pre-emphasis coefficient; (2) Frame segmentation; The audio signal is segmented into several frames, and the length of each frame is a fixed time window; the frame segmentation process is shown as follows: s m (n) = s′(n)·w(n - mR) where s m (n) represents the m-th frame signal, w(n) represents the window function, and the Hamming window function is used. R represents the frame offset and takes a value of half of the frame length; (3) STFT transformation; Perform STFT transformation on the sound signal to convert the time-domain signal into a frequency-domain signal; specifically, segment the time-domain signal, the length of each segment of the time-domain signal is the window length, there is an overlap between two adjacent segments of the time-domain signal, and perform Fourier transform on each segment of the time-domain signal to obtain its frequency-domain representation, that is, decompose the time-domain signal into several frequency bands in the frequency domain; the process of STFT transformation is expressed by the formula as follows: Among them, S m (k) represents the frequency-domain representation of the m-th frame, N represents the number of sampling points per frame, and k = n - mR; in addition, j represents the imaginary unit and e represents the natural constant; (4) Feature extraction; Feature extraction extracts representative features from the frequency-domain signal to provide reliable input data for the subsequent sound source localization task; use the MFCC method to extract sound features at different frequencies; First, pass the frequency-domain signal through the Mel filter bank to obtain the coefficients at different frequencies, and the formula is expressed as follows: Among them, H m (k) represents the m-th filter of the Mel filter bank, h(m) represents the frequency corresponding to the m-th Mel frequency, and E m represents the Mel frequency coefficient of the m-th frame; Perform discrete cosine transform DCT on the filtered Mel frequency coefficients: perform DCT transformation on the Mel frequency coefficients after taking the logarithm to obtain a set of MFCC coefficients, and the MFCC coefficients represent the speech features of the audio signal, which are expressed as follows: Among them, f a represents the a-th MFCC coefficient, and M represents the number of MFCC coefficients; the MFCC coefficients are used as the feature representation for the sound source localization task; (5) Format conversion; Convert the data extracted by feature extraction into a form suitable for processing by a convolutional neural network, that is, convert the frequency-domain signal into an image form, so that the data is regarded as two-dimensional image data; by performing spectral analysis on the signal, extract the amplitude spectrum or phase spectrum of the signal in the frequency domain, and then use the two-dimensional fast Fourier transform FFT to convert it into an image form; amplitude spectrum Q m (k) and phase spectrum P m (k) are calculated as follows: Q m (k) = |S m (k)| P m (k) = arg{S m (k)} Where |·| represents the modulo operation, and arg{·} represents the argument operation; regard the amplitude spectrum or phase spectrum as two-dimensional image data, and by stacking the amplitude spectrum or phase spectrum at different time steps, obtain a three-dimensional data set, where the first dimension represents time, the second dimension represents frequency, and the third dimension represents amplitude or phase, as the input data of the convolutional neural network.

4. The cross-modal sound source localization and gesture recognition method combining sound and image according to claim 1, characterized in that, The implementation process of step 5 is as follows: (1) Data input; Obtain video data as input, segment the video data into T frames, input the image frames into the YOLOv7 network, and perform multiple convolution and pooling operations on the image frames to obtain a series of feature maps; (2) Anchor box and feature map processing; YOLOv7 uses Anchor boxes to predict the position and size of the target; input the feature maps obtained in the data input stage into the convolutional layer and pooling layer for processing to obtain feature maps of different scales; (3) Target prediction; In the feature map, each pixel point corresponds to an Anchor box. By classifying and regressing each Anchor box, it is predicted whether there is a target object in each box, and its position and size are estimated; the objective function is expressed as follows: L = L cls + λ coord L coord + λ obj L obj + λ noobj L noobj Among them, L cls represents the classification loss, L coord represents the position loss, L obj represents the loss of the presence of the target object, L noobj represents the loss of the absence of the target object; λ coord and λ obj and λ noobj are weight parameters; the classification loss is calculated using the cross-entropy loss function, the position loss is calculated using the mean squared error loss function, and the loss of the presence of the target object and the loss of the absence of the target object are respectively expressed as: Among them, G represents the size of the feature map, and B represents the number of Anchor boxes predicted for each pixel point. Indicates whether there is a target object in the j-th Anchor box on the i-th pixel point. Indicates whether the j-th Anchor box on the i-th pixel point does not contain a target object; x i and y i represent the center point coordinates of the predicted box. represents the center point coordinates of the actual target object; C i represents the confidence level of whether there is a target object in the predicted box. represents the confidence level of the existence of the actual target object. Redundant prediction boxes are removed by non-maximum suppression (NMS), and the prediction result is denoted as β = {b 1 , b 2 , …, b T}, where b t = (x t , y t , w t , h t ) represents the center coordinates (x t , y t ) and the width and height (w t , h t ) of the instruction initiator in the t-th frame image, thereby determining the position of the instruction initiator in each frame of the image.

5. The cross-modal sound source localization and gesture recognition method combining sound and image according to claim 1, characterized in that, In step 6, the instruction initiator in each frame of image is extracted, and its hand joint points are detected by AlphaPose; after AlphaPose, for the t-th frame of image, the joint point information corresponding to the hand of the instruction initiator is denoted as U t , where represents the coordinate of the d-th joint point of the hand of the instruction initiator in the t-th frame of image; the hand joint point information of the instruction initiator in each frame of image is concatenated to obtain the joint point information of the entire gesture, denoted as U = {U 1 , U 2 , …, U T}.

6. The cross-modal sound source localization and gesture recognition method combining sound and image according to claim 1, characterized in that, In step 7, after the spatio-temporal graph convolutional network obtains the joint point information of the instruction initiator's hand, the joint point information U of the instruction initiator's hand in each frame is t regarded as a graph, each joint point is used as a graph node, and the edges between the graph nodes represent the connection relationships between the joint points; the coordinates and time information of each node are concatenated to obtain a three-dimensional tensor where T represents the number of time steps, D represents the number of key points, and 3 represents the feature dimension of each node; the spatio-temporal graph convolutional neural network is used to perform a convolutional operation on this three-dimensional tensor to obtain a new three-dimensional tensor where F represents the number of feature channels; this three-dimensional tensor is regarded as the feature representation of the gesture, and global pooling or a convolutional neural network is used for classification or regression tasks; the overall gesture representation is obtained by weighted summation of the feature representations of the instruction initiator's gestures in all frames, and it is transformed into a probability distribution through the softmax function, that is, the probability that the gesture belongs to each category is obtained, and thus the final gesture recognition result is obtained.

Citation Information

Patent Citations

  • Intelligent robot

    CN108000529A

  • Traffic police command gesture recognition method based on skeleton joint point sequence

    CN110837778A