Voice interaction method and system based on lip movement and voice fusion recognition and garbage truck
Through the lip movement and speech fusion recognition method, combined with high-definition cameras and deep learning technology, the problem of low recognition accuracy in high-noise environments in traditional voice interaction systems is solved, and higher recognition accuracy and system robustness are achieved, adapting to complex environments and generating reliable operating instructions.
Patent Information
- Application Number
- CN202510635118.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional voice interaction systems are difficult to effectively filter background noise in high noise environments, resulting in low speech recognition accuracy and cannot meet the application needs of autonomous driving vehicles in complex environments.
The method based on lip movement and speech fusion recognition is adopted, and the lip movement images are collected through high-definition cameras and pre-processed. The lip region is located in combination with a deep convolutional neural network and a cascade classifier, and the speech signal noise reduction is reduced by combining the deep convolutional neural network and short-time Fourier transform. The feature recognition is used for CNN+LSTM and LSTM+CTC architectures, and finally the operation instructions are generated through the weighted average method.
It significantly improves the recognition accuracy in high-noise environments, enhances the system's adaptability and robustness to complex environments, generates more reliable operating instructions, avoids confusion of similar pronunciation instructions, and improves the safety and reliability of autonomous vehicles.
Smart Images

Figure CN120356470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent environmental sanitation, and particularly relates to a voice interaction method, system and garbage truck based on lip movement and voice fusion recognition. Background Art
[0002] Currently, common voice interaction systems in autonomous vehicles generally use multiple microphone arrays for voice collection, and adopt noise reduction algorithms (such as spectral subtraction, Wiener filtering, etc.) to preprocess the collected voice signals to reduce the interference of environmental noise on voice recognition. Then, feature extraction and recognition are performed on the processed voice signals, and voice commands are converted into text or operation commands.
[0003] However, in practical applications, autonomous vehicles may face various complex environmental conditions, and high-noise environments pose a severe challenge to the performance of voice interaction systems. Taking a garbage transfer station as an example, such places usually have mechanical noises generated by the operation of various large mechanical equipment, impact noises during garbage loading and unloading, and engine noises of vehicles driving. The multiple noises are superimposed on each other, forming an extremely complex and high-intensity noise environment. In such a high-noise environment, even with the above noise reduction algorithms, traditional voice interaction systems are still difficult to effectively filter background noise. Relevant experimental data show that in high-noise environments such as garbage transfer stations where the noise intensity reaches above 80 decibels, the voice recognition accuracy of traditional voice interaction systems may be lower than 50%, which far fails to meet the requirements of autonomous vehicles for the accuracy and reliability of voice interaction in practical applications.
[0004] Therefore, it is necessary to develop a voice interaction method, system and garbage truck based on lip movement and voice fusion recognition. Summary of the Invention
[0005] The purpose of the present invention is to provide a voice interaction method, system and garbage truck based on lip movement and voice fusion recognition, which can significantly improve the recognition accuracy in high-noise environments.
[0006] In the first aspect, a voice interaction method based on lip movement and voice fusion recognition includes the following steps:
[0007] Collect high-definition lip movement images, and perform preprocessing and lip region positioning on the lip movement images to obtain a sequence of lip images;
[0008] Collect voice signals, and perform noise reduction processing on the voice signals to obtain noise-reduced voice signals;
[0009] Extract lip movement features of the sequence of lip images, and extract voice features of the noise-reduced voice signals;
[0010] Identify the lip movement feature and the speech feature respectively to obtain a lip movement recognition result and a speech recognition result;
[0011] Fuse the lip movement recognition result and the speech recognition result to generate a final operation instruction;
[0012] Execute corresponding operations based on the operation instruction.
[0013] Optionally, the steps of collecting high-definition lip movement images, preprocessing the lip movement images, and positioning the lip region to obtain a lip image sequence include:
[0014] Install a high-definition camera in the cab and in front of the driver directly to ensure that the lip image is clear and occupies the central area of the image;
[0015] Collect high-definition lip movement images in real time, convert the images into grayscale images, and perform denoising and contrast enhancement processing;
[0016] Use a cascade classifier based on Haar features to detect the face region, locate the lip region, crop out the lip image, and adjust the image size to a preset size. Installing the high-definition camera in the cab and in front of the driver directly can ensure that a clear, complete lip image in the central area of the image is collected, providing high-quality image data for subsequent lip region positioning and feature extraction, and reducing recognition errors caused by image quality problems. Converting the image into a grayscale image can simplify the image processing process and reduce the computational complexity; denoising processing can effectively remove noise interference in the image and improve the image quality; contrast enhancement processing makes the lip features more obvious, facilitating subsequent lip region detection and positioning. The cascade classifier based on Haar features is an efficient face detection method that can quickly and accurately detect the face region and then locate the lip position. Cropping out the lip image and adjusting the size helps to unify the format of the input data, facilitating subsequent feature extraction and model processing, and improving the recognition efficiency and accuracy.
[0017] Optionally, the steps of collecting a speech signal and performing noise reduction processing on the speech signal to obtain a noise-reduced speech signal include:
[0018] Collect the speech signal through an audio interface;
[0019] Use the short-time Fourier transform to convert the speech signal into a spectral representation;
[0020] Use a speech enhancement model based on a deep convolutional neural network to perform noise reduction processing on the spectrum to obtain an enhanced speech signal spectrum;
[0021] The enhanced speech signal spectrum is converted back to the time-domain signal by using the inverse short-time Fourier transform. The speech signal is collected through the audio interface, which can stably and accurately collect the speech signal and provide the original data for subsequent noise reduction and recognition processing. The short-time Fourier transform converts the speech signal in the time domain into the spectral representation in the frequency domain, which is convenient for analyzing the frequency characteristics of the speech signal and provides a suitable signal representation form for subsequent noise reduction processing. The deep convolutional neural network has powerful feature extraction and pattern recognition capabilities, can effectively learn the noise characteristics in the speech signal, remove them from the speech signal, and obtain the enhanced speech signal spectrum, improve the quality of the speech signal, and provide a more accurate input for subsequent speech recognition. The inverse short-time Fourier transform converts the enhanced speech signal spectrum in the frequency domain back to the time-domain signal, enabling it to be processed by subsequent speech recognition models, and realizing the process of noise reduction in the frequency domain and restoration of the time-domain signal.
[0022] Optionally, extracting the lip movement features of the lip image sequence includes:
[0023] Performing normalization processing on the cropped lip image;
[0024] Using the Canny edge detection algorithm to extract the lip contour, and using the optical flow method to calculate the optical flow features of the lip movement;
[0025] Using an improved convolutional neural network model to extract the features of the lip image. The input is a continuous lip image sequence, and the output is a lip movement feature vector. Performing normalization processing on the cropped lip image can make the lip image have better consistency and comparability in the feature extraction process, and improve the stability and accuracy of feature extraction. The Canny edge detection algorithm can accurately extract the lip contour and provide a basis for the analysis of lip movements; the optical flow method can calculate the optical flow features of the lip movement, reflecting the movement direction and speed of the lip between different frames. These features can effectively describe the dynamic changes of the lip and provide rich information for lip movement recognition. The improved convolutional neural network model can automatically learn the deep features in the lip image, capture the spatio-temporal information of the lip movement by processing the continuous lip image sequence, and output the lip movement feature vector. These feature vectors can more comprehensively and accurately represent the lip movement and improve the performance of lip movement recognition.
[0026] Optionally, extracting the speech features of the noise-reduced speech signal includes:
[0027] Performing frame segmentation on the noise-reduced speech signal;
[0028] Calculate the MFCC features of each frame of the speech signal and extract the spectral features of the speech signal. The speech signal has time-varying characteristics. Frame processing can divide the continuous speech signal into multiple short-time frames. Within each frame, the speech signal can be approximately regarded as stationary, which is convenient for subsequent feature extraction and analysis. MFCC (Mel Frequency Cepstral Coefficients) features are a commonly used speech feature. It can simulate the auditory characteristics of the human ear, better reflect the spectral features of the speech signal, and has good robustness to the changes of the speech signal. It is an important feature representation method in speech recognition.
[0029] Optionally, the separately identifying the lip movement feature and the speech feature to obtain a lip movement recognition result and a speech recognition result includes:
[0030] Use a lip movement recognition model with a CNN+LSTM hybrid architecture to identify the lip movement feature and output the probability distribution of the lip movement recognition result;
[0031] Use a speech recognition model with an LSTM+CTC architecture to identify the speech feature and output the probability distribution of the speech recognition result. CNN (Convolutional Neural Network) can effectively extract the spatial information in the lip movement feature, and LSTM (Long Short-Term Memory Network) is good at processing sequence data and capturing the temporal relationship of lip movements. The CNN+LSTM hybrid architecture combines the advantages of both, can more accurately identify the lip movement feature, output the probability distribution of the lip movement recognition result, and provide a basis for subsequent fusion decision-making. LSTM can process the temporal information of the speech feature, and the CTC (Connectionist Temporal Classification) algorithm can solve the problem of inconsistent lengths of the input sequence and the output sequence in speech recognition, directly decode the speech feature sequence, and output the probability distribution of the speech recognition result, improving the accuracy and efficiency of speech recognition.
[0032] Optionally, the fusing the lip movement recognition result and the speech recognition result to generate a final operation instruction includes:
[0033] Adopt the weighted average method to fuse the lip movement recognition result and the speech recognition result to obtain the probability distribution of the fused instruction;
[0034] When the probability of a certain instruction in the probability distribution of the fused instruction exceeds a preset threshold, generate this instruction as the final operation instruction. The weighted average method can comprehensively consider the importance of the lip movement recognition result and the speech recognition result, fuse the recognition results of the two modalities by reasonably allocating weights, obtain the probability distribution of the fused instruction, and make the fused result more accurately reflect the actual user intention. And by setting a preset threshold, it can effectively filter out the misrecognition results with low probabilities. Only when the probability of a certain instruction exceeds the threshold is the final operation instruction generated, improving the accuracy and reliability of the final instruction and reducing the occurrence of misoperations.
[0035] In a second aspect, a voice interaction system based on lip movement and voice fusion recognition according to the present invention includes:
[0036] A lip movement image acquisition module, configured to acquire high-definition lip movement images, preprocess the lip movement images, and locate the lip region, so as to obtain a sequence of lip images;
[0037] A voice signal acquisition and noise reduction module, configured to acquire voice signals and perform noise reduction processing on the voice signals to obtain noise-reduced voice signals;
[0038] A feature extraction module, configured to extract lip movement features of the sequence of lip images and extract voice features of the noise-reduced voice signals;
[0039] An identification module, configured to respectively identify the lip movement features and the voice features to obtain a lip movement identification result and a voice identification result;
[0040] A fusion and decision-making module, configured to fuse the lip movement identification result and the voice identification result to generate a final operation instruction;
[0041] An execution module, configured to perform corresponding operations based on the operation instruction.
[0042] Optionally, the lip movement image acquisition module includes:
[0043] An image acquisition unit, configured to acquire high-definition lip movement images;
[0044] An image preprocessing unit, configured to convert image information into a grayscale image and perform denoising and contrast enhancement processing;
[0045] A lip region location unit, configured to detect a face region using a cascade classifier based on Haar features and locate the lip region to crop out a lip image;
[0046] The voice signal acquisition and noise reduction module includes:
[0047] A voice signal acquisition unit, configured to acquire voice signals through an audio interface;
[0048] A spectrum conversion unit, configured to convert voice signals into a spectrum representation using short-time Fourier transform;
[0049] A noise reduction processing unit, configured to perform noise reduction processing on the spectrum using a voice enhancement model based on a deep convolutional neural network to obtain an enhanced voice signal spectrum;
[0050] A time domain conversion unit, configured to convert the enhanced voice signal spectrum back into a time domain signal using inverse short-time Fourier transform.
[0051] In a third aspect, a garbage truck according to the present invention adopts the voice interaction system based on lip movement and voice fusion recognition as described in the present invention.
[0052] Advantages of the present invention:
[0053] (1) Significantly improve the recognition accuracy in high-noise environments: In high-noise environments such as garbage transfer stations (where the noise intensity reaches above 80 decibels), the voice recognition accuracy of traditional voice interaction systems may be lower than 50%. The present invention performs recognition by fusing information from two modalities, lip movement and voice, and comprehensively utilizes the respective advantages of lip movements and voice signals. When the quality of the voice signal severely deteriorates due to a high-noise environment, lip movement information can serve as a crucial supplement to effectively make up for the deficiencies in voice recognition, thereby significantly improving the recognition accuracy, greatly enhancing the recognition accuracy, and better meeting the actual application requirements of autonomous driving vehicles in complex environments.
[0054] (2) Enhance the adaptability of the system to complex environments: Autonomous driving vehicles will encounter various complex environmental conditions during actual driving. Traditional voice interaction systems often struggle to effectively handle these complex environments. The recognition method of the present invention that fuses lip movement and voice enables the system to no longer solely rely on voice signals for recognition. Lip movement information is not directly affected by environmental noise and can provide additional recognition basis for the system. Whether in a garbage transfer station where mechanical noise, impact noise, and engine noise are superimposed on each other, or in other complex environments with various interference factors, the system can accurately recognize based on the comprehensive information of lip movement and voice, thus greatly enhancing the adaptability to complex environments.
[0055] (3) Improve the robustness of the system: Traditional voice interaction systems are easily interfered with in high-noise environments, resulting in unstable recognition results. By fusing lip movement and voice features, when the voice signal is severely interfered with or partially missing, the lip movement recognition result can be used as an auxiliary judgment basis to ensure that the system can still generate relatively accurate operation instructions based on the comprehensive information. This multi-modal fusion recognition method enables the system to maintain stable recognition performance in the face of various uncertainties and interferences, improves the robustness of the system, and provides a strong guarantee for the safe and reliable operation of autonomous driving vehicles.
[0056] (4) Generate more reliable operation instructions: The operation instructions of autonomous vehicles are directly related to driving safety and user experience. Traditional voice interaction systems have low recognition accuracy in high-noise environments, which may lead to the generation of incorrect or inaccurate operation instructions, increasing driving risks. By fusing lip movement and speech recognition results and processing them through a scientific fusion algorithm, the present invention can more accurately understand the user's intention and generate more reliable operation instructions. This helps autonomous vehicles make correct decisions in complex environments, improves driving safety and stability, and provides users with a better-quality and more reliable voice interaction experience.
[0057] (5) Avoid confusion of instructions with similar pronunciations: Through lip movement and speech fusion recognition, the problem of confusion of instructions with similar pronunciations (such as "start" and "stop") is effectively solved. Description of the Drawings
[0058] Figure 1 is a flowchart of the voice interaction method based on lip movement and speech fusion recognition described in the embodiments of the present application;
[0059] Figure 2 is a schematic block diagram of the voice interaction system based on lip movement and speech fusion recognition described in the embodiments of the present application;
[0060] Figure 3 is a schematic block diagram of the lip movement image acquisition module described in the embodiments of the present application;
[0061] Figure 4 is the voice signal acquisition and noise reduction module described in the embodiments of the present application. Detailed Embodiments
[0062] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention and not for limiting the protection scope of the present invention.
[0063] As Figure 1 shown, in the embodiments of the present application, a voice interaction method based on lip movement and speech fusion recognition includes the following steps:
[0064] Collect high-definition lip movement images, preprocess the lip movement images and locate the lip region to obtain a sequence of lip images.
[0065] Collect voice signals and perform noise reduction processing on the voice signals to obtain noise-reduced voice signals.
[0066] Extract the lip movement features of the lip image sequence and extract the speech features of the denoised speech signal.
[0067] Identify the lip movement features and speech features respectively to obtain the lip movement recognition result and the speech recognition result.
[0068] Fuse the lip movement recognition result and the speech recognition result to generate the final operation instruction;
[0069] Execute the corresponding operation based on the operation instruction.
[0070] This method performs recognition by fusing the information of two modalities, lip movement and speech, comprehensively utilizes the advantages of lip movements and speech signals, can effectively improve the recognition accuracy. Especially in a noisy environment, when the quality of the speech signal deteriorates, the lip movement information can be used as an important supplement to enhance the adaptability and robustness of the system to complex environments, and finally generate more reliable operation instructions.
[0071] In a possible embodiment, high-definition lip movement images are collected, and the lip movement images are preprocessed and the lip region is located to obtain a lip image sequence, including:
[0072] Install a high-definition camera in the cab and in front of the driver's line of sight, ensuring that the lip image is clear and occupies the central area of the image. Collect high-definition lip movement images in real time, convert the images to grayscale images, and perform denoising and contrast enhancement processing. Use a cascade classifier based on Haar features to detect the face region and locate the lip region, crop out the lip image, and adjust the image size to a preset size.
[0073] In this method, installing the high-definition camera in the cab and in front of the driver's line of sight can ensure that clear, complete lip images in the central area of the image are collected, providing high-quality image data for subsequent lip region location and feature extraction, and reducing recognition errors caused by image quality problems. Converting the image to a grayscale image can simplify the image processing process and reduce the computational complexity; denoising processing can effectively remove noise interference in the image and improve the image quality; contrast enhancement processing makes the lip features more obvious, facilitating subsequent lip region detection and location. The cascade classifier based on Haar features is an efficient face detection method that can quickly and accurately detect the face region and then locate the lip position. Cropping out the lip image and adjusting the size helps to unify the format of the input data, facilitating subsequent feature extraction and model processing, and improving the recognition efficiency and accuracy.
[0074] In one example, the camera is installed above the dashboard in the cab and in front of the driver's line of sight, about 50 cm away from the driver's face, and the angle is adjusted to face the driver's face directly. After installation, calibrate the camera to ensure that the lip image is clear and occupies the central area of the image.
[0075] A high-definition camera captures high-definition lip movement images (RGB images) in real time and transmits the image data to the in-vehicle computing unit through a USB 3.0 interface. The captured images are converted into grayscale images; a Gaussian filter is used to denoise the grayscale images, and the filter parameter is σ = 1.5; the denoised grayscale images are subjected to histogram equalization and the contrast of the images is enhanced.
[0076] A cascade classifier based on Haar features is used to detect the face region in the image, and the lip region is further located. The lip image is cropped to ensure that the size of the lip image is 224×224 pixels for subsequent processing.
[0077] In a possible embodiment, a voice signal is collected and the voice signal is subjected to noise reduction processing to obtain a noise-reduced voice signal, including:
[0078] The voice signal is collected through an audio interface; the voice signal is converted into a spectral representation by using short-time Fourier transform; a voice enhancement model based on a deep convolutional neural network is used to perform noise reduction processing on the spectrum to obtain an enhanced voice signal spectrum; the enhanced voice signal spectrum is converted back into a time-domain signal by using inverse short-time Fourier transform.
[0079] This method collects the voice signal through the audio interface, can stably and accurately collect the voice signal, and provides the original data for subsequent noise reduction and recognition processing. The short-time Fourier transform converts the voice signal in the time domain into a spectral representation in the frequency domain, which is convenient for analyzing the frequency characteristics of the voice signal and provides a suitable signal representation form for subsequent noise reduction processing. The deep convolutional neural network has strong feature extraction and pattern recognition capabilities, can effectively learn the noise characteristics in the voice signal, remove them from the voice signal, obtain an enhanced voice signal spectrum, improve the quality of the voice signal, and provide a more accurate input for subsequent voice recognition. The inverse short-time Fourier transform converts the enhanced voice signal spectrum in the frequency domain back into a time-domain signal, enabling it to be processed by subsequent voice recognition models, realizing the process of noise reduction in the frequency domain and restoration of the time-domain signal.
[0080] In an example, the driver issues a "start" command in a noisy environment. The collected voice signal is transmitted to the in-vehicle computing unit through the audio interface.
[0081] The voice signal is converted into a spectral representation S(t,f) by using short-time Fourier transform, and the formula is:
[0082]
[0083] Among them, \(x(n)\) is the speech signal, \(w(n)\) is the window function, \(N\) is the frame length, and \(j\) is the imaginary unit. \(S(t,f)\) contains information in both the time \(t\) and frequency \(f\) dimensions. By observing the change in \(t\), the frequency characteristics of the speech signal at different time periods can be understood; by observing the change in \(f\), the variation of the energy of the speech signal at different frequencies over time can be analyzed. During the noise reduction process, based on the value of \(S(t,f)\), it can be determined which time-frequency points are mainly noise and which are mainly speech.
[0084] Use a speech enhancement model based on a deep convolutional neural network (DCNN). The input is the spectrum of the speech signal, and the output is the spectrum of the enhanced speech signal.
[0085] Model training data: It includes speech spectrum samples in a high-noise environment and corresponding clean speech spectrum samples, with the number of samples being 10,000.
[0086] Model training objective: Minimize the mean squared error loss function, and the formula is:
[0087]
[0088] Among them, MSE is the mean squared error, \(K\) is the number of training samples, \(S\) clean (i) is the \(i\)-th clean speech spectrum, \(S\) enHanced (i) is the \(i\)-th enhanced speech spectrum. \(\|\cdot\|\) 2 represents the square of the two-norm of the vector, which is used to calculate the square of the Euclidean distance between two spectra and measure the magnitude of the difference between them.
[0089] Use the inverse short-time Fourier transform to convert the spectrum of the enhanced speech signal back to the time-domain signal \(x(n)'\), and the formula is:
[0090]
[0091] Among them, \(S(m,f)\) is the spectrum of the enhanced speech signal; \(M\) is the frame shift.
[0092] In a possible embodiment, extract the lip movement features of the lip image sequence (the lip movement features include optical flow features and lip movement feature vectors), including:
[0093] Normalize the cropped lip images; use the Canny edge detection algorithm to extract the lip contour and use the optical flow method to calculate the optical flow features of the lip movement; use an improved convolutional neural network model to extract the features of the lip images. The input is a continuous sequence of lip images, and the output is the lip movement feature vector.
[0094] This method normalizes the cropped lip images, enabling better consistency and comparability during feature extraction, and improving the stability and accuracy of feature extraction. The Canny edge detection algorithm can accurately extract the lip contour, providing a basis for the analysis of lip movements; the optical flow method can calculate the optical flow features of lip movements, reflecting the movement direction and speed of the lips between different frames. These features can effectively describe the dynamic changes of the lips and provide rich information for lip movement recognition. The improved convolutional neural network model can automatically learn the deep features in lip images, capture the spatio-temporal information of lip movements by processing continuous sequences of lip images, and output lip movement feature vectors. These feature vectors can more comprehensively and accurately represent lip movements and improve the performance of lip movement recognition.
[0095] In one example, when the driver issues a "start" command, the high-definition camera captures lip movement images.
[0096] The cropped lip images are normalized, and the image size is adjusted to 224×224 pixels.
[0097] The Canny edge detection algorithm is used to extract the lip contour. The formula is:
[0098]
[0099] Among them, G(x,y) is the gradient magnitude. T is the threshold, a pre-set value used to distinguish between edge and non-edge pixels. If G(x,y)>T, it is considered that there is an edge at the pixel point (x,y). If G(x,y)≤T, it is considered that there is no edge at the pixel point (x,y). Edge(x,y) represents the output result of edge detection, indicating whether there is an edge at the pixel point (x,y). Edge(x,y) = 1: indicates that an edge is detected at the pixel point (x,y). Edge(x,y) = 0: indicates that no edge is detected at the pixel point (x,y).
[0100] The optical flow method is used to calculate the optical flow features of lip movements and extract the direction and speed information of lip movements. The formula is:
[0101]
[0102] Among them, V(x,y) is the optical flow feature; I is the luminance function; is the partial derivative of luminance in the horizontal direction (x-axis) or the gradient of luminance in the x-axis direction. is the partial derivative of luminance or the gradient of luminance in the horizontal direction (x-axis).
[0103] Feature extraction of lip images is performed using an improved convolutional neural network (CNN) model. The improved CNN model is based on the ResNet-50 architecture and adds a SENet attention mechanism module. The input is a continuous sequence of lip images (30 frames per second), and the output is a lip movement feature vector. The extraction formula for the lip movement feature vector is:
[0104] F lip = CNN(I lip ) ⊙ SE(I lip )
[0105] where I lip is the sequence of lip images, F lip is the extracted lip movement feature vector (i.e., a high-dimensional feature vector), and ⊙ represents element-wise multiplication.
[0106] In a possible embodiment, extracting the speech features of the noise-reduced speech signal includes:
[0107] Performing frame segmentation on the noise-reduced speech signal; calculating the MFCC features of each frame of the speech signal and extracting the spectral features of the speech signal.
[0108] In this method, the speech signal has time-varying characteristics. Frame segmentation can divide the continuous speech signal into multiple short-time frames. Within each frame, the speech signal can be approximately regarded as stationary, which is convenient for subsequent feature extraction and analysis. MFCC (Mel Frequency Cepstral Coefficients) features are a commonly used speech feature. It can simulate the auditory characteristics of the human ear, better reflect the spectral features of the speech signal, and has good robustness to the changes of the speech signal. It is an important feature representation method in speech recognition.
[0109] In an example, the noise-reduced speech signal is framed, with each frame having a length of 25 ms and a frame shift of 10 ms. Calculate the MFCC features (Mel Frequency Cepstral Coefficients) of each frame of the speech signal and extract the spectral features of the speech signal. The formula is:
[0110] F audio = MFCC(S)
[0111] where S is the spectrum of the speech signal, and F audio is the extracted MFCC feature vector.
[0112] In a possible embodiment, the lip movement features and speech features are respectively recognized to obtain lip movement recognition results and speech recognition results, including:
[0113] Use a lip movement recognition model with a CNN+LSTM hybrid architecture to recognize lip movement features and output the probability distribution of lip movement recognition results; use a speech recognition model with an LSTM+CTC architecture to recognize speech features and output the probability distribution of speech recognition results.
[0114] CNN (Convolutional Neural Network) can effectively extract the spatial information in lip movement features, while LSTM (Long Short-Term Memory Network) is good at processing sequential data and capturing the temporal relationship of lip movements. The CNN+LSTM hybrid architecture combines the advantages of both, enabling more accurate recognition of lip movement features and outputting the probability distribution of lip movement recognition results, providing a basis for subsequent fusion decisions. LSTM can process the temporal information of speech features, and the CTC (Connectionist Temporal Classification) algorithm can solve the problem of inconsistent lengths between the input sequence and the output sequence in speech recognition, directly decoding the speech feature sequence and outputting the probability distribution of speech recognition results, improving the accuracy and efficiency of speech recognition.
[0115] In one example, the following details lip movement recognition and speech recognition:
[0116] 1. Lip movement recognition model training:
[0117] Adopt a lip movement recognition model with a CNN+LSTM hybrid architecture. The CNN part extracts the spatio-temporal features of the lip image, and the LSTM part performs temporal modeling on the feature sequence.
[0118] The training data includes 10,000 lip movement samples of drivers in different environments, labeled with corresponding speech command categories.
[0119] The training objective is to minimize the cross-entropy loss function:
[0120]
[0121] where \(L\) is the minimized cross-entropy loss function, \(y\) i is the true label, is the class probability distribution predicted by the model, and \(K\) is the number of training samples.
[0122] 2. Output of lip movement recognition results:
[0123] Input the extracted lip movement feature vector into the trained lip movement recognition model with a CNN+LSTM hybrid architecture, and the lip movement recognition model with a CNN+LSTM hybrid architecture outputs the probability distribution of lip movement recognition results.
[0124] Suppose the probability distribution output by the lip movement recognition model with a CNN+LSTM hybrid architecture is:
[0125] F lip= [0.85, 0.10, 0.05]
[0126] Among them, F lip represents the probability distribution of the lip movement recognition result.
[0127] Assume the output result is as follows: 0.85 represents the confidence of the "start" instruction, 0.10 represents the confidence of the "stop" instruction, and 0.05 represents the confidence of other instructions.
[0128] 3. Speech recognition model training:
[0129] Adopt a speech recognition model with an LSTM + CTC architecture. The LSTM part performs temporal modeling on the speech feature sequence, and the CTC part maps the feature sequence output by the model to text instructions. The training data includes 10,000 speech samples in a high-noise environment, labeled with corresponding text instructions.
[0130] The training objective is to minimize the CTC (Connectionist Temporal Classification) loss function:
[0131]
[0132] Among them, L CTC Minimize the CTC loss function, Y i is the true text instruction, X i is the speech feature sequence, and K is the number of training samples.
[0133] 4. Speech recognition result output:
[0134] Input the extracted MFCC feature vectors into the trained speech recognition model with an LSTM + CTC architecture. The speech recognition model with an LSTM + CTC architecture outputs the probability distribution of the speech recognition result.
[0135] Assume the probability distribution output by the speech recognition model with an LSTM + CTC architecture is:
[0136] F audio = [0.75, 0.15, 0.10]
[0137] Among them, F audio represents the probability distribution of the speech recognition result.
[0138] Assume the output result is as follows: 0.75 represents the confidence of the "start" instruction, 0.15 represents the confidence of the "stop" instruction, and 0.10 represents the confidence of other instructions.
[0139] In a possible embodiment, fuse the lip movement recognition result and the speech recognition result to generate the final operation instruction, including:
[0140] The weighted average method is used to fuse the lip movement recognition result and the speech recognition result to obtain the fused instruction probability distribution; when the probability of a certain instruction in the fused instruction probability distribution exceeds the preset threshold, that instruction is generated as the final operation instruction.
[0141] In this method, the weighted average method can comprehensively consider the importance of the lip movement recognition result and the speech recognition result. By reasonably allocating weights, the recognition results of the two modalities are fused to obtain the fused instruction probability distribution, making the fused result more accurately reflect the actual user intention. And by setting the preset threshold, the misrecognition results with relatively low probabilities can be effectively filtered out. Only when the probability of a certain instruction exceeds the threshold is the final operation instruction generated, improving the accuracy and reliability of the final instruction and reducing the occurrence of misoperations.
[0142] In an example, the fusion of the lip movement recognition result and the speech recognition result is described in detail as follows:
[0143] The weighted average method is used to fuse the lip movement recognition result and the speech recognition result, and the weight allocation is α = 0.6.
[0144] Fusion formula:
[0145] F final = α·F lip +(1 - α)·F audio
[0146] where F final is the fused instruction probability distribution,
[0147] Substitute specific data:
[0148] F final = 0.6×[0.85, 0.10, 0.05] + 0.4×[0.75, 0.15, 0.10] = [0.81, 0.12, 0.07]
[0149] Set the confidence threshold to 0.8 (for reference). The probability of the "start" instruction in the fused instruction probability distribution is 0.81, which is higher than the set confidence threshold. The probability of the confidence of the "stop" instruction is 0.12, and the probabilities of other instructions are 0.07.
[0150] Therefore, the final operation instruction "start" is generated and sent to the control system of the garbage truck. The control system executes the corresponding operation based on the operation instruction.
[0151] In a possible embodiment, it is also necessary to install a touch display screen and a high-fidelity speaker in the cab. If the system successfully recognizes and generates an operation instruction, the display screen shows "Instruction received: Start", and the speaker plays a voice prompt "Instruction received, start operation". If the system fails to recognize or the confidence level is insufficient, the display screen shows "Instruction not recognized, please re-enter", and the speaker plays a voice prompt "Instruction not recognized, please re-enter". The touch display screen adopts a graphical interface design, which displays real-time images of lip movements and voice waveform diagrams to help the driver intuitively understand the input state of the system. The high-fidelity speaker automatically adjusts the volume according to the noise level in the cab to ensure that the driver can clearly hear the prompt information.
[0152] In the embodiment of the present application, through the fusion recognition of lip movement and voice, the problem of confusion of similar pronunciation instructions (such as "start" and "stop") is effectively solved as follows:
[0153] First, using lip movement features provides an additional dimension for differentiation. The lip movements corresponding to different instructions are different. Even if the pronunciations are similar, the lip movement patterns are different. For example, when saying "start", the lips are slightly open and protruded, and when saying "stop", the lips are closed and tightened. The present application extracts lip movement features to capture these subtle differences. When speech recognition is confused by similar pronunciation instructions, lip movement recognition can assist in judgment based on lip movement features to avoid misjudgment.
[0154] Second, multi-modal fusion enhances recognition accuracy. In this method, lip movement and voice features are recognized separately and then fused. For similar pronunciation instructions, the weight of lip movement information is appropriately increased during fusion. Because in a complex environment, lip movement is less affected by interference and can provide a more direct basis for differentiation. For example, in a noisy environment, the voice signal is easily interfered, but lip movement recognition is normal. By fusing the information of the two modalities, comprehensive consideration can be made to improve the recognition accuracy of similar pronunciation instructions.
[0155] Third, the deep learning model captures complex feature relationships. In the extraction and recognition of lip movement and voice features, this method uses an improved convolutional neural network, a CNN+LSTM hybrid architecture lip movement recognition model, and an LSTM+CTC architecture voice recognition model. These models can automatically learn the complex feature relationships between lip movements and voice signals, and through a large amount of training data, master the subtle differences in lip movement and voice features of similar pronunciation instructions, so as to accurately distinguish them in actual recognition.
[0156] Fourth, the weighted average fusion algorithm optimizes the decision-making. When fusing the lip movement and speech recognition results, the weighted average method is adopted to reasonably adjust the weights of the two, so that the lip movement recognition result plays a greater role in decision-making, improving the accuracy of the final recognition result and effectively avoiding the confusion of similar pronunciation instructions.
[0157] Such as Figure 2As shown in the figure, in the embodiment of the present application, a voice interaction system based on lip movement and voice fusion recognition includes a lip movement image acquisition module, a voice signal acquisition and noise reduction module, a feature extraction module, a recognition module, a fusion and decision-making module, and an execution module. The lip movement image acquisition module and the voice signal acquisition and noise reduction module are respectively connected to the feature extraction module, the feature extraction module is connected to the recognition module, the recognition module is connected to the fusion and decision-making module, and the fusion and decision-making module is connected to the execution module. The lip movement image acquisition module is used to acquire high-definition lip movement images, preprocess the lip movement images, and locate the lip area to obtain a sequence of lip images. The voice signal acquisition and noise reduction module is used to acquire voice signals and perform noise reduction processing on the voice signals to obtain noise-reduced voice signals. The feature extraction module is used to extract lip movement features from the sequence of lip images and extract voice features from the noise-reduced voice signals. The recognition module is used to respectively recognize the lip movement features and the voice features to obtain lip movement recognition results and voice recognition results. The fusion and decision-making module is used to fuse the lip movement recognition results and the voice recognition results to generate a final operation instruction. The execution module performs corresponding operations based on the operation instruction,
[0158] As Figure 3 shown, in a possible embodiment, the lip movement image acquisition module includes an image acquisition unit, an image preprocessing unit, and a lip area positioning unit connected in sequence. The image acquisition unit is used to acquire high-definition lip movement images. The image preprocessing unit is used to convert the image information into a grayscale image and perform denoising and contrast enhancement processing. The lip area positioning unit is used to detect the face area using a cascade classifier based on Haar features and locate the lip area to crop out the lip image.
[0159] As Figure 4 shown, the voice signal acquisition and noise reduction module includes a voice signal acquisition unit, a spectrum conversion unit, a noise reduction processing unit, and a time domain conversion unit connected in sequence. Among them, the voice signal acquisition unit is used to acquire voice signals through an audio interface. The spectrum conversion unit is used to convert the voice signals into a spectrum representation using the short-time Fourier transform. The noise reduction processing unit is used to perform noise reduction processing on the spectrum using a voice enhancement model based on a deep convolutional neural network to obtain an enhanced voice signal spectrum. The time domain conversion unit is used to convert the enhanced voice signal spectrum back into a time domain signal using the inverse short-time Fourier transform.
[0160] In the embodiment of the present application, a garbage truck adopts the voice interaction system based on lip movement and voice fusion recognition as in the embodiment of the present application.
[0161] The above embodiments are preferred embodiments of the present invention. However, the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A voice interaction method based on lip movement and speech fusion recognition, characterized in that, Including the following steps: Collect high-definition lip movement images, preprocess the lip movement images and locate the lip region to obtain a lip image sequence; Collect voice signals and perform noise reduction processing on the voice signals to obtain noise-reduced voice signals; Extract the lip movement features of the lip image sequence and extract the voice features of the noise-reduced voice signals; Identify the lip movement features and the voice features respectively to obtain a lip movement recognition result and a voice recognition result; Fuse the lip movement recognition result and the voice recognition result to generate a final operation instruction; Execute corresponding operations based on the operation instruction.
2. The voice interaction method based on lip movement and voice fusion recognition according to claim 1, characterized in that The collecting of high-definition lip movement images, preprocessing the lip movement images and locating the lip region to obtain a lip image sequence includes: Install a high-definition camera inside the cab and in front of the driver to ensure that the lip image is clear and occupies the central area of the image; Collect high-definition lip movement images in real time, convert the images into grayscale images, and perform denoising and contrast enhancement processing; Use a cascade classifier based on Haar features to detect the face region, locate the lip region, crop the lip image, and adjust the image size to a preset size.
3. The voice interaction method based on lip movement and voice fusion recognition according to claim 1, characterized in that The collecting of voice signals and performing noise reduction processing on the voice signals to obtain noise-reduced voice signals includes: Collect voice signals through an audio interface; Use the short-time Fourier transform to convert the voice signals into a spectral representation; Use a voice enhancement model based on a deep convolutional neural network to perform noise reduction processing on the spectrum to obtain an enhanced voice signal spectrum; Use the inverse short-time Fourier transform to convert the enhanced voice signal spectrum back into a time-domain signal.
4. The voice interaction method based on lip movement and voice fusion recognition according to claim 1, characterized in that The extracting of the lip movement features of the lip image sequence includes: Perform normalization processing on the cropped lip images; Use the Canny edge detection algorithm to extract the lip contour and use the optical flow method to calculate the optical flow features of lip movement; Use an improved convolutional neural network model to extract features of lip images. The input is a continuous lip image sequence, and the output is a lip movement feature vector.
5. The voice interaction method based on lip movement and voice fusion recognition according to claim 1, characterized in that, The extracting of the voice features of the noise-reduced voice signals includes: Perform frame segmentation on the noise-reduced voice signals; Calculate the MFCC features of each frame of voice signals and extract the spectral features of the voice signals.
6. The method according to claim 1, wherein The respectively identifying of the lip movement features and the voice features to obtain a lip movement recognition result and a voice recognition result includes: Use a lip movement recognition model with a CNN+LSTM hybrid architecture to identify the lip movement features and output the probability distribution of the lip movement recognition result; Use a voice recognition model with an LSTM+CTC architecture to identify the voice features and output the probability distribution of the voice recognition result.
7. The method according to claim 1, wherein The fusing of the lip movement recognition result and the voice recognition result to generate a final operation instruction includes: Use the weighted average method to fuse the lip movement recognition result and the voice recognition result to obtain the probability distribution of the fused instruction; When the probability of a certain instruction in the probability distribution of the fused instruction exceeds a preset threshold, generate this instruction as the final operation instruction.
8. A voice interaction system based on the fusion recognition of lip movement and speech, characterized in that, Including: The lip movement image acquisition module is used to acquire high-definition lip movement images, preprocess the lip movement images and locate the lip region, so as to obtain a sequence of lip images; The voice signal acquisition and noise reduction module is used to acquire voice signals and perform noise reduction processing on the voice signals to obtain noise-reduced voice signals; The feature extraction module is used to extract the lip movement features of the sequence of lip images and extract the voice features of the noise-reduced voice signals; The recognition module is used to recognize the lip movement features and the voice features respectively to obtain lip movement recognition results and voice recognition results; The fusion and decision-making module is used to fuse the lip movement recognition results and the voice recognition results to generate a final operation instruction; The execution module performs corresponding operations based on the operation instruction.
9. The voice interaction system based on lip movement and voice fusion recognition according to claim 8, wherein, The lip movement image acquisition module includes: The image acquisition unit is used to acquire high-definition lip movement images; The image preprocessing unit is used to convert the image information into a grayscale image and perform denoising and contrast enhancement processing; The lip region location unit is used to detect the face region using a cascade classifier based on Haar features and locate the lip region to crop out the lip image; The voice signal acquisition and noise reduction module includes: The voice signal acquisition unit is used to acquire voice signals through an audio interface; The spectrum conversion unit is used to convert the voice signal into a spectrum representation using short-time Fourier transform; The noise reduction processing unit is used to perform noise reduction processing on the spectrum using a voice enhancement model based on a deep convolutional neural network to obtain an enhanced voice signal spectrum; The time domain conversion unit is used to convert the enhanced voice signal spectrum back into a time domain signal using inverse short-time Fourier transform.
10. A garbage truck, characterized in that, A voice interaction system based on lip movement and voice fusion recognition as claimed in claim 8 or 9 is adopted.
Citation Information
Cited By
AI multi-mode voice interaction method based on vehicle-mounted intelligent terminal and electronic equipment
CN121171226A