Speech recognition method and system based on artificial intelligence
The speech signal is processed through wavelet transformation and Mel frequency cepspectral coefficients, combined with the LBP value of the lip image to generate feature vectors, and the weight is adjusted using the cross-modal attention mechanism to build a deep learning network, which solves the problem of the reduction in accuracy of traditional speech recognition systems in noise environments and realizes accurate speech recognition in high noise environments.
Patent Information
- Application Number
- CN202510484618.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional voice recognition systems rely on audio signals and are susceptible to environmental noise, especially in complex noise or multi-person conversation scenarios, and the recognition accuracy is reduced, and the computing resources and time are consumed, which cannot meet the needs of instant voice recognition.
Wavelet transformation and threshold processing technology are used to denoise speech signals, voice features are extracted in combination with the Mel frequency cepspectral coefficient, and image feature vectors are generated in combination with the LBP value of the lip-moving image. The weights of speech and visual features are adjusted through the cross-modal attention mechanism to construct a deep learning network for fusion recognition.
Improve speech recognition accuracy and robustness in high noise environments, adaptively adjust feature fusion strategies, and provide more accurate and stable recognition results.
Smart Images

Figure CN120388575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and specifically provides a speech recognition method and system based on artificial intelligence. Background Art
[0002] At present, speech recognition technology has been widely applied in multiple fields, such as intelligent assistants, automatic translation, and barrier-free technology. However, traditional speech recognition systems mainly rely on audio signals and are easily affected by environmental noise, resulting in a decline in recognition accuracy and robustness. In addition, single audio input performs particularly poorly in complex noise or multi-person conversation scenarios. Multimodal speech recognition systems can effectively solve this problem, that is, simultaneously utilize audio and visual information for speech recognition to improve the recognition effect in noisy environments. The primary challenge is how to effectively integrate these two types of information and adjust their contribution degrees during the recognition process.
[0003] In the prior art, the publication number CN116580706B discloses a speech recognition method based on artificial intelligence. By collecting the speech audio information input by the user, converting the speech audio information into an audio spectrogram and obtaining multiple audio frames in the audio spectrogram, extracting the feature information in each audio frame, associating the feature information of multiple audio frames to obtain the data to be recognized, inputting the data to be recognized into a trained speech recognition model, determining the speech content corresponding to the speech audio information, and verifying the speech content to obtain and output the speech recognition result.
[0004] The main problems of the above method are as follows: It overall depends on speech audio information, so the recognition accuracy will be affected in complex noise environments. Different recording devices and recording conditions will lead to differences in audio quality, resulting in certain errors when generating spectrograms. Moreover, the process of feature extraction through audio spectrograms is complex, requiring a large amount of computing resources and time, and may not be able to meet the usage scenarios of real-time speech recognition.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present invention is to provide a speech recognition method and system based on artificial intelligence to solve the problems raised in the above background art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A speech recognition method based on artificial intelligence, the specific steps include:
[0009] Step 1: Collect the original speech signal, perform noise reduction processing on the original speech signal, decompose the speech signal into multiple sub-bands with different frequencies through wavelet transform. Each sub-band corresponds to the information of the signal at different time scales. Perform threshold processing on the high-frequency sub-bands to weaken the noise components, and recombine the processed sub-bands to obtain the denoised speech signal;
[0010] Step 2: Pre-emphasize the denoised speech signal through a high-pass filter, segment the pre-emphasized speech signal into short-time frames of 20 - 40 ms, with 50% of the frame length overlapping between frames. Window each frame signal, and convert the windowed frame signal from the time domain signal to the frequency domain signal through the fast Fourier transform. Generate Mel-frequency cepstral coefficients from the frequency domain signal, and combine the Mel-frequency cepstral coefficients with their first-order and second-order differences into a speech feature vector;
[0011] Step 3: Real-time collect the lip movement information of the speaker as the original visual image, perform grayscale processing on the original image to generate the first recognition image;
[0012] Step 4: Normalize the pixels of the first recognition image. For each pixel point in the first recognition image, select 8 neighboring pixel points around the pixel point, compare the gray value of the central pixel point with its neighboring pixel points to generate a binary number, and convert the binary number into a decimal number to generate the LBP value of the central pixel point. Calculate the LBP values of all pixel points in the first recognition image to form a new image, which is called the second recognition image;
[0013] Step 5: Divide the second recognition image into non-overlapping units of 8×8 pixels. Directly crop and discard the units that do not meet the size. Statistically analyze the LBP value distribution of all pixel points inside each unit to generate a histogram of a unit. Concatenate the histograms of all units end to end in the unit order to generate an image feature vector as the final feature representation of the visual signal;
[0014] Step 6: Convert the speech feature vector and the image feature vector into matrix representations, and dynamically adjust the weights of the speech feature and the image feature through a cross-modal attention mechanism to generate a fusion weight matrix;
[0015] Step 7: Construct a deep learning network model. Use the original speech signal and the original visual image as the training set, and use the speech command corresponding to the fusion weight matrix as the label, and input them into the deep learning network model for training;
[0016] Step 8: Input the real-time collected speech signal and visual image into the trained deep learning network model, obtain the current fusion weight matrix, and determine the current speech command according to the fusion weight matrix.
[0017] Furthermore, the principle for judging the high-frequency subband range is as follows:
[0018]
[0019] Among them, represents the frequency of the high-frequency subband, represents the sampling rate of the original signal, represents the corresponding layer number in wavelet transform.
[0020] Furthermore, the formula for decomposing the speech signal into multiple subbands with different frequencies through wavelet transform is:
[0021]
[0022]
[0023]
[0024] Among them, represents the Haar wavelet basis function, represents the time variable, represents the wavelet coefficient of the high-frequency subband, represents the original signal, represents the result of scaling and translation of the wavelet basis function is the scaling parameter, is the translation parameter;
[0025] Set a noise threshold, and generate subband coefficients according to the noise threshold and wavelet coefficients. The formula is:
[0026]
[0027] Among them, represents the subband coefficient after processing, represents the noise threshold;
[0028] Perform inverse wavelet transform on the subband coefficient after processing to generate the denoised signal. The formula is:
[0029]
[0030] Among them, represents the denoised signal, represents the subband coefficient, represents the result of scaling and translation of the wavelet basis function is the scaling parameter, is the translation parameter.
[0031] Furthermore, the principle for generating Mel Frequency Cepstral Coefficients is as follows:
[0032] First, pre-emphasize the denoised signal, and the formula is:
[0033]
[0034] where represents the value of the pre-emphasized signal at time corresponding to it, represents the value of the denoised signal at time corresponding to it, represents the value of the denoised signal at time corresponding to it, represents the pre-emphasis coefficient, represents the time variable;
[0035] After splitting the pre-emphasized signal into short-time frames, window each frame of the signal, and the formula is:
[0036]
[0037]
[0038] where represents the result after windowing the split frame signal, represents the split frame signal, represents the time variable, represents the Hamming window function, represents the time index, represents the frame length;
[0039] The formula for converting the windowed frame signal from the time domain signal to the frequency domain signal is:
[0040]
[0041] where Z represents the frequency domain signal, represents the time domain signal, represents the frame length, represents the imaginary unit, represents the frequency index, with a value range of [0, N - 1], represents the time index, with a value range of [0, N - 1];
[0042] The formula for the energy output of each filter is:
[0043]
[0044] Among them, represents the value of the th Mel filter, represents the frequency index, represents the th frequency of the Mel filter;
[0045] Take the logarithm of the energy output of each filter to generate log energy, and perform discrete cosine transform on the log energy to generate Mel cepstral coefficients. The formula is:
[0046]
[0047]
[0048] Among them, among them, represents the log energy of the th filter, represents the frequency-domain signal, represents the th value of the Mel filter, represents the frequency index, represents the frame length, represents the th Mel frequency cepstral coefficient, represents the number of Mel filters, represents the th log energy of the Mel filter, represents the index of the Mel filter, represents the index of the Mel frequency cepstral coefficient;
[0049] The formulas for calculating the first-order difference and second-order difference of the Mel frequency cepstral coefficients are:
[0050]
[0051]
[0052] Among them, represents the first-order difference of the th Mel frequency cepstral coefficient, represents the th second-order difference of the Mel frequency cepstral coefficient, represents the frame offset, determined by the size of the difference window, and the size of the difference window is , and the specific value is determined by the actual application scenario;
[0053] The principle of combining the Mel frequency cepstral coefficients, first-order differences, and second-order differences into a feature vector is:
[0054]
[0055] wherein, represents the final speech feature vector, represents the th Mel-frequency cepstral coefficient, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient,
[0056] Furthermore, the formula for generating the first recognition image by grayscale processing the original image is:
[0057]
[0058] wherein, represents the grayscale value of the pixel point, represents the red channel value of the pixel point of the original image, represents the green channel value of the pixel point of the original image, represents the blue channel value of the pixel point of the original image.
[0059] Furthermore, the principle for generating a binary number by comparing the grayscale value of the central pixel point with that of its neighboring pixel points is:
[0060] For a central pixel point with 8 neighboring pixel points, it can be expressed as:
[0061]
[0062] The formula for judging the value of the binary number is:
[0063] [[ID=5L]]
[0064] wherein, represents the grayscale value of the central pixel point, represents the grayscale value of the neighboring pixel point, is a comparison formula for judging based on the size relationship between the central pixel point and the neighboring pixel points;
[0065] The formula for converting the binary number into the LBP value of the pixel point is:
[0066]
[0067] wherein, represents the LBP value of the central pixel point, represents the number of neighboring pixel points, It represents a comparison formula for making a judgment based on the size relationship between the central pixel point and the neighborhood pixel points.
[0068] Furthermore, the principle for generating the histogram of one cell is as follows:
[0069] Statistically analyze the distribution of LBP values of all pixel points within one cell. Divide the LBP values into 256 intervals according to the range of [0, 255]. Each interval corresponds to an LBP value, and generate a histogram vector with a length of 256, where each element represents the number of pixel points with LBP values within the corresponding interval;
[0070] The formula for generating a feature vector by connecting the histograms of all cells in sequence is:
[0071]
[0072] Among them, represents the final image feature vector, represents the th number of LBP values in the th interval of the histogram of the
[0073] Furthermore, the principle for generating the fusion weight matrix is as follows:
[0074] Regard the speech feature vector as a 1×3L matrix, which can be expressed as:
[0075]
[0076] Among them, represents the speech feature matrix, represents the th Mel-frequency cepstral coefficient, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient;
[0077] Regard the image feature vector as a 1×256E matrix, which can be expressed as:
[0078]
[0079] Among them, represents the image feature matrix, represents the th number of LBP values in the Indicates the number of units;
[0080] The voice feature matrix and the image feature matrix have the same number of elements, and the formula for generating their similarity is:
[0081]
[0082] Wherein, Indicates the similarity of the corresponding voice feature matrix and image feature matrix elements, Indicates the column element of the voice feature matrix, Indicates the column element of the image feature matrix;
[0083] Arrange all the calculated similarity values in sequence to form a matrix , and the matrix is the correlation matrix between the voice feature matrix and the image feature matrix;
[0084] The formula for calculating the attention weight based on the correlation matrix is:
[0085]
[0086]
[0087] Wherein, Indicates the attention weight of the voice feature, Indicates the attention weight of the image feature, Indicates the correlation matrix, Indicates the transposed matrix of the correlation matrix;
[0088] The formula for using the calculated weights for feature fusion to generate a fusion weight matrix is:
[0089]
[0090] Wherein, Indicates the fusion weight matrix, Indicates the attention weight of the voice feature, Indicates the voice feature matrix, Indicates the attention weight of the image feature, Indicates the image feature matrix.
[0091] The present invention also provides an artificial intelligence-based speech recognition system, and the system is used to implement the above artificial intelligence-based speech recognition method, including:
[0092] The voice acquisition module is used to acquire the original voice signal, perform noise reduction processing on the original voice signal, decompose the voice signal into multiple sub-bands with different frequencies through wavelet transform, where each sub-band corresponds to the information of the signal at different time scales, perform threshold processing on the high-frequency sub-bands to weaken the noise components, and recombine the processed sub-bands to obtain the denoised voice signal;
[0093] The voice feature extraction module is used to pre-emphasize the denoised voice signal through a high-pass filter, segment the pre-emphasized voice signal into short-time frames of 20 - 40 ms, with 50% of the frame length overlapping between frames, window each frame signal, transform the windowed frame signal from the time domain signal to the frequency domain signal through the fast Fourier transform, generate Mel cepstral coefficients from the frequency domain signal, and combine the Mel cepstral coefficients and their first-order and second-order differences into a voice feature vector;
[0094] The image acquisition module is used to collect the lip movement information of the speaker in real time as the visual original image, perform grayscale processing on the original image to generate the first recognition image;
[0095] The image processing module is used to perform normalization processing on the pixels of the first recognition image. For each pixel point in the first recognition image, select 8 neighboring pixel points around the pixel point, compare the gray value of the central pixel point with its neighboring pixel points to generate a binary number, convert the binary number into a decimal number to generate the LBP value of the central pixel point, and calculate the LBP values of all pixel points in the first recognition image to form a new image, which is called the second recognition image;
[0096] The image feature extraction module is used to divide the second recognition image into non-overlapping units of 8×8 pixels, directly crop and discard the units that do not meet the size, count the distribution of the LBP values of all pixel points inside each unit to generate a histogram of a unit, and splice the histograms of all units end to end in the unit order to generate an image feature vector as the final feature representation of the visual signal;
[0097] The feature fusion module is used to convert the voice feature vector and the image feature vector into matrix representations, and dynamically adjust the weights of the voice feature and the image feature through a cross-modal attention mechanism to generate a fusion weight matrix;
[0098] The model construction module is used to construct a deep learning network model, use the original voice signal and the visual original image as the training set, and use the voice command corresponding to the fusion weight matrix as the label, and input them into the deep learning network model for training;
[0099] The voice recognition module inputs the real-time collected voice signals and visual images into the trained deep learning network model, obtains the current fusion weight matrix, and determines the current voice command according to the fusion weight matrix.
[0100] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0101] According to the multi-modal voice recognition method of the present invention, the voice signal features and visual signal features are comprehensively considered, and different processing methods are adopted for different signal features: wavelet transform and threshold processing technology are used to denoise the voice signal to improve the voice recognition accuracy in a high-noise environment. For the voice features after noise reduction processing, they are transformed into Mel frequency and then Mel frequency cepstral coefficients are generated to extract key voice features, which are closer to human senses; after graying the lip movement information image, the LBP value of each pixel point is calculated to generate the feature vector of the image. Combining the voice feature vector and the image feature vector, the attention weights of the two are adjusted through the cross-modal attention mechanism, and then a fusion weight matrix that can comprehensively reflect the voice features and image features is generated. Different fusion weight matrices correspond to different voice features and image features, and thus different voice commands. Using the original voice signal and visual original image as the training set and the voice command corresponding to the fusion weight matrix as the label, the deep learning network model is trained. Inputting the real-time collected voice signals and visual images into the trained model can generate different voice commands corresponding to different fusion weight matrices. The present invention can adaptively optimize the weight ratio of voice features and image features in the fusion weight matrix in different environments to adjust the feature fusion strategy, thereby providing more accurate and stable recognition results. Description of the Drawings
[0102] Figure 1 It is a schematic flowchart of the method of the embodiment of the present invention;
[0103] Figure 2 It is a schematic diagram of the system module of the embodiment of the present invention. Detailed Embodiments
[0104] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.
[0105] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0106] Embodiment:
[0107] Please refer to Figure 1 , the present invention provides a technical solution:
[0108] A speech recognition method based on artificial intelligence, the specific steps include:
[0109] Step 1: Collect the original speech signal, perform noise reduction processing on the original speech signal, decompose the speech signal into multiple sub-bands with different frequencies through wavelet transform, each sub-band corresponding to the information of the signal at different time scales, perform threshold processing on the high-frequency sub-bands to weaken the noise components, and recombine the processed sub-bands to obtain the denoised speech signal;
[0110] In this embodiment, the principle for judging the range of the high-frequency sub-band is:
[0111]
[0112] Among them, represents the frequency of the high-frequency sub-band, represents the sampling rate of the original signal, represents the corresponding layer number in the wavelet transform.
[0113] The formula based on which the wavelet transform is used is:
[0114]
[0115]
[0116]
[0117] Among them, represents the Haar wavelet basis function, represents the time variable, represents the wavelet coefficients of the high-frequency subband, represents the original signal, represents the result of stretching and translation of the wavelet basis function is the result of stretching and translation, is the stretching scale parameter, is the translation scale parameter;
[0118] The wavelet transform decomposes the speech signal into a series of basis functions called wavelets, which are generated by translating and scaling a single prototype waveform with different scale and position parameters, and can provide information about the speech signal in both time and scale simultaneously. The wavelet coefficients represent the characteristics of the signal at different scales and time positions. The numerical value reflects the amplitude of the signal at the corresponding wavelet scale and time position, and the sign reflects the positive and negative polarity of the energy. The scale parameter The larger it is, the lower the corresponding frequency, and the rougher the signal represented by the wavelet coefficients. The scale parameter The smaller it is, the higher the corresponding frequency, and the finer the signal represented by the wavelet coefficients; the translation parameter The larger it is, the later the corresponding time position, and the wavelet coefficients represent the characteristics of the signal in the later stage. The translation parameter The smaller it is, the earlier the corresponding time position, and the wavelet coefficients represent the characteristics of the signal in the earlier stage.
[0119] Set a noise threshold, and generate subband coefficients based on the noise threshold and wavelet coefficients. The formula is:
[0120]
[0121] where, represents the processed subband coefficients, represents the noise threshold;
[0122] The subband coefficients are the result of threshold processing on the wavelet coefficients, reflecting the morphology of the signal after denoising; hard threshold processing is used in the threshold processing. Wavelet coefficients not greater than the threshold are set to zero, and wavelet coefficients greater than the threshold remain unchanged, which can better retain the original characteristics of the signal.
[0123] Perform inverse wavelet transform on the processed subband coefficients to generate the denoised signal. The formula is:
[0124]
[0125] where, represents the denoised signal, represents the subband coefficients, represents the result of stretching and translation of the wavelet basis function is the result of stretching and translation, is the stretching scale parameter, is the translation scale parameter;
[0126] The purpose of the inverse wavelet transform is to reconstruct the original signal from the denoised wavelet coefficients. By combining wavelet coefficients at different scales, it reflects the process of reconstructing the original signal with multi-level information from details to overview. The accuracy of the result of the inverse wavelet transform is proportional to the number of retained wavelet coefficients and the accuracy of the retained wavelet coefficients.
[0127] Step 2: Pre-emphasize the denoised speech signal through a high-pass filter, segment the pre-emphasized speech signal into short-time frames of 20 - 40 ms, overlap the frames by 50% of the frame length, window each frame signal, transform the windowed frame signal from the time domain signal to the frequency domain signal through the fast Fourier transform, generate mel-frequency cepstral coefficients from the frequency domain signal, and combine the mel-frequency cepstral coefficients and their first-order and second-order differences into a speech feature vector;
[0128] In this embodiment, the principle for generating mel-frequency cepstral coefficients is as follows:
[0129] First, pre-emphasize the denoised signal, and the formula is:
[0130]
[0131] where, represents the value of the pre-emphasized signal at time corresponding, represents the value of the denoised signal at time corresponding, represents the value of the denoised signal at time corresponding, represents the pre-emphasis coefficient, represents the time variable;
[0132] The pre-emphasized signal reflects the enhancement effect on the original signal. The common value range of the pre-emphasis coefficient is 0.90 - 0.97. The larger the pre-emphasis coefficient, the more obvious the high-frequency enhancement effect; the smaller the pre-emphasis coefficient, the weaker the high-frequency enhancement effect. The pre-emphasized signal is proportional to the pre-emphasis coefficient and the intensity of the original signal.
[0133] After segmenting the pre-emphasized signal into short-time frames, window each frame signal, and the formula is:
[0134]
[0135]
[0136] where, Represents the result after windowing the segmented frame signal. Represents the segmented frame signal. Represents the time variable. Represents the Hamming window function. Represents the time index. Represents the frame length;
[0137] Windowing the signal reduces the discontinuity at both ends of the signal, making the spectral representation of the signal clearer and more concentrated on the actual frequency components of the signal. The purpose of choosing the Hamming window function is to effectively reduce spectral leakage and provide a good balance between time-domain and frequency-domain performance, reducing the sidelobe amplitude while maintaining a reasonable main lobe width.
[0138] The formula for converting the windowed frame signal from the time-domain signal to the frequency-domain signal is:
[0139]
[0140] Among them, Z Represents the frequency-domain signal. Represents the time-domain signal. Represents the frame length. Represents the imaginary unit. Represents the frequency index, and the value range is [0, N - 1]. Represents the time index, and the value range is [0, N - 1];
[0141] Using the Fourier transform to convert the time-domain signal to the frequency-domain signal reflects the distribution of the time-domain signal at each frequency index. As the frequency index increases, it means that the time-domain signal is multiplied by a faster oscillating waveform, directly affecting the amplitude and phase of the corresponding frequency components in the frequency-domain representation. The frequency-domain component is inversely proportional to the frequency index and directly proportional to the input signal samples.
[0142] The formula for the energy output of each filter is:
[0143]
[0144] Among them, Represents the value of the frequency index of the th Mel filter;
[0145] Taking the logarithm of the energy output of each filter to generate the log energy is based on the formula:
[0146]
[0147] Among them, represents the logarithmic energy of the th filter, represents the frequency-domain signal, represents the value of the th Mel filter, represents the frequency index, represents the frame length;
[0148] The logarithmic energy reflects the human ear's perception of sound intensity. The magnitude of the logarithmic energy is proportional to the frequency-domain components and proportional to the response values of the Mel filters;
[0149] Performing a discrete cosine transform on the logarithmic energy to generate Mel-frequency cepstral coefficients, the formula based on is:
[0150]
[0151] Among them, represents the th Mel-frequency cepstral coefficient, represents the number of Mel filters, represents the logarithmic energy of the th Mel filter, represents the index of the Mel filter, represents the index of the Mel-frequency cepstral coefficient;
[0152] The Mel-frequency cepstral coefficients reflect the spectral characteristics of the speech signal that match the human ear's auditory perception. The filter bank on the Mel scale is used to reflect the non-linear perception of frequency by the human ear. In the low-frequency region, the frequency resolution on the Mel scale is higher, and in the high-frequency region, the resolution is lower. The Mel-frequency cepstral coefficients are proportional to the magnitude of the logarithmic energy of the Mel filters and proportional to the number of Mel filters;
[0153] The formula for calculating the first-order difference of the Mel-frequency cepstral coefficients is:
[0154]
[0155] Among them, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the frame offset, determined by the size of the difference window, and the size of the difference window is , and the value of the difference window size is determined by the specific application scenario;
[0156] The formula for calculating the second-order difference of the Mel-frequency cepstral coefficients is:
[0157]
[0158] Among them, represents the second-order difference of the th Mel-frequency cepstral coefficient, represents the frame offset, which is determined by the size of the difference window. The size of the difference window is , and the specific value is determined through experiments;
[0159] The principle for combining the Mel-frequency cepstral coefficient, the first-order difference, and the second-order difference into a feature vector is as follows:
[0160]
[0161] Among them, represents the final speech feature vector, represents the th Mel-frequency cepstral coefficient, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient,
[0162] The first-order difference of the Mel-frequency represents the speed feature, that is, the change rate of the Mel-frequency cepstral coefficient over time, reflecting the dynamic transformation of the speech signal. The second-order difference represents the acceleration feature, that is, the change rate of the first-order difference of the Mel-frequency cepstral coefficient over time, which can capture more subtle transformations of the speech signal. Combining the Mel-frequency cepstral coefficient, the first-order difference of the Mel-frequency cepstral coefficient, and the second-order difference of the Mel-frequency cepstral coefficient to generate a feature vector can timely capture the fast-changing dynamic features in the speech signal, thereby improving the accuracy and robustness of speech recognition.
[0163] Step 3: Real-time collect the lip movement information of the speaker as the visual original image, and perform grayscale processing on the original image to generate the first recognition image;
[0164] In this embodiment, the formula for performing grayscale processing on the original image to generate the first recognition image is:
[0165]
[0166] Among them, represents the grayscale value of the pixel point, represents the red channel value of the original image pixel point, represents the green channel value of the original image pixel point, represents the blue channel value of the original image pixel point;
[0167] The grayscale processing formula reflects that green has the greatest weight on the grayscale, followed by red, and blue has the least. The selection of this weighting method is based on the optical perception principle of color images, ensuring that the conversion result is as visually consistent as possible with the brightness level of the original image's color. The purpose of grayscale conversion is to remove color information and only retain brightness information, simplifying the complexity of image processing.
[0168] Step 4: Normalize the pixels of the first recognition image. For each pixel point in the first recognition image, select 8 neighboring pixel points around the pixel point, compare the grayscale value of the central pixel point with that of its neighboring pixel points, generate a binary number, convert the binary number into a decimal number to generate the LBP value of the central pixel point, and calculate the LBP values of all pixel points in the first recognition image to form a new image, called the second recognition image.
[0169] In this embodiment, the principle for comparing the grayscale value of the central pixel point with that of its neighboring pixel points to generate a binary number is as follows:
[0170] For the central pixel point with 8 neighboring pixel points in the neighborhood, it can be expressed as:
[0171]
[0172] The formula for judging the value of the binary number is:
[0173]
[0174] Among them, represents the grayscale value of the central pixel point, represents the grayscale value of the neighboring pixel point, is a comparison formula for judging based on the size relationship between the central pixel point and the neighboring pixel points;
[0175] The formula for converting the binary number into the LBP value of the pixel point is:
[0176]
[0177] Among them, represents the LBP value of the central pixel point, represents the number of neighboring pixel points, represents the comparison formula for judging based on the size relationship between the central pixel point and the neighboring pixel points.
[0178] In the calculation of LBP, when the grayscale value of the neighboring pixel point is not less than the grayscale value of the central pixel point , the value is 1. When the grayscale value of the neighboring pixel point is less than the grayscale value of the central pixel point When the value is 0, it is judged successively from the upper left neighborhood pixel point among the 8 neighborhood pixel points around the central pixel point, and a binary sequence is generated. Finally, the binary sequence is converted into a decimal number, which corresponds to the LBP value of the central pixel point. The LBP value of the pixel point reflects the change of the gray value in the neighborhood. If the gray values of the pixel points in the neighborhood are generally greater than or equal to the central pixel point, the LBP value is higher, indicating that the texture feature of the image is stronger. If the gray values of the pixel points in the neighborhood are generally less than the central pixel point, the LBP value is lower, and the texture feature of the image is weaker. The more the number of edge pixel points with gray values greater than the central pixel point, the higher the LBP value.
[0179] Step 5: Divide the second recognition image into non-overlapping units with 8×8 pixels. The units that do not meet the size are directly cropped and discarded. Count the distribution of the LBP values of all pixel points inside each unit, generate a histogram of one unit, and splice the histograms of all units end to end according to the unit order to generate an image feature vector as the final feature representation of the visual signal.
[0180] In this embodiment, the principle for generating a histogram of one unit is as follows:
[0181] Count the distribution of the LBP values of all pixel points inside one unit. Divide the LBP values into 256 intervals according to the range of [0, 255]. Each interval corresponds to an LBP value, and generate a histogram vector with a length of 256, where each element represents the number of pixel points with LBP values in the corresponding interval.
[0182] The formula for generating a feature vector by connecting the histograms of all units in sequence is as follows:
[0183]
[0184] Among them, represents the final image feature vector, represents the th number of LBP values in the th interval in the histogram of the
[0185] th unit, and represents the number of units.
[0185] Divide the second recognition image into several non-overlapping units, each unit with a size of 8×8. The parts that do not meet the size are discarded. For each unit, count the respective situations of the LBP values of all pixel points inside, the number of each value of the LBP value in the range of [0, 255]. After processing all units, connect all units to generate a feature vector.
[0186] Step 6: Convert the speech feature vector and the image feature vector into matrix representations, and dynamically adjust the weights of the speech features and the image features through a cross-modal attention mechanism to generate a fusion weight matrix;
[0187] The principle for generating the fusion weight matrix is as follows:
[0188] Regard the speech feature vector as a 1×3L matrix, which can be expressed as:
[0189]
[0190] Where, represents the speech feature matrix, represents the th Mel-frequency cepstral coefficient, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient;
[0191] Regard the image feature vector as a 1×256E matrix, which can be expressed as:
[0192]
[0193] Where, represents the image feature matrix, represents the number of LBP values in the th unit's histogram in the th interval, represents the number of units;
[0194] Take the speech feature matrix and the image feature matrix with the same number of elements to generate their similarity. The formula is as follows:
[0195]
[0196] Where, represents the similarity between the corresponding speech feature matrix and image feature matrix elements, represents the element in the th column of the speech feature matrix, represents the element in the th column of the image feature matrix;
[0197] Arrange all the calculated similarity values in sequence to form a matrix The matrix is the correlation matrix between the speech feature matrix and the image feature matrix;
[0198] Calculate the attention weights according to the correlation matrix, and the formula is as follows:
[0199]
[0200]
[0201] Where, represents the attention weight of the speech feature, represents the attention weight of the image feature, represents the correlation matrix, represents the transpose matrix of the correlation matrix;
[0202] The purpose of calculating the attention weights through the softmax function is to convert the original correlation scores into probability values ranging from 0 to 1, and ensure that the sum of all attention weights is 1, ensuring that the model makes a reasonable weight distribution among different modalities, and avoiding the situation where the features of a certain modality are too prominent or ignored.
[0203] Use the calculated weights for feature fusion to generate a fusion weight matrix, and the formula is as follows:
[0204]
[0205] Where, represents the fusion weight matrix, represents the attention weight of the speech feature, represents the speech feature matrix, represents the attention weight of the image feature, represents the image feature matrix;
[0206] The fusion weight matrix is a matrix that fuses speech features and image features, and generates a comprehensive and more informative feature representation, reflecting the correlation of features in each modality. Each element in the fusion weight matrix represents the correlation between a specific speech feature and a specific image feature. The higher the weight value, the higher the similarity between the two.
[0207] Step 7: Construct a deep learning network model, use the original speech signal and the original visual image as the training set, and use the speech command corresponding to the fusion weight matrix as the label, and input them into the deep learning network model for training;
[0208] In this embodiment, the different features extracted from the original speech signal and the original visual image, as well as the differences in their attention weights, will affect the generation result of the fusion weight matrix. The generated different fusion weight matrices correspond to different speech commands according to their unique speech feature vectors and image feature vectors.
[0209] Step 8: Input the real-time collected voice signal and visual image into the trained deep learning network model to obtain the current fusion weight matrix, and determine the current voice command according to the fusion weight matrix;
[0210] Please refer to Figure 2 , the present invention also provides an artificial intelligence-based voice recognition system, which is used to implement the above artificial intelligence-based voice recognition method, including:
[0211] A voice acquisition module, which is used to collect the original voice signal, perform noise reduction processing on the original voice signal, decompose the voice signal into multiple sub-bands with different frequencies through wavelet transform, each sub-band corresponds to the information of the signal at different time scales, perform threshold processing on the high-frequency sub-band to weaken the noise component, and recombine the processed sub-bands to obtain the denoised voice signal;
[0212] A voice feature extraction module, which is used to pre-emphasize the denoised voice signal through a high-pass filter, divide the pre-emphasized voice signal into short-time frames of 20-40 ms, with 50% of the frame length overlapping between frames, window each frame signal, convert the windowed frame signal from the time domain signal to the frequency domain signal through the fast Fourier transform, generate Mel cepstral coefficients from the frequency domain signal, and combine the Mel cepstral coefficients and their first-order and second-order differences into a voice feature vector;
[0213] An image acquisition module, which is used to collect the lip movement information of the speaker in real time as the visual original image, perform grayscale processing on the original image to generate the first recognition image;
[0214] An image processing module, which is used to perform normalization processing on the pixels of the first recognition image. For each pixel point in the first recognition image, select 8 neighboring pixel points around the pixel point, compare the grayscale value of the central pixel point with its neighboring pixel points to generate a binary number, convert the binary number into a decimal number to generate the LBP value of the central pixel point, and calculate the LBP values of all pixel points in the first recognition image to form a new image, which is called the second recognition image;
[0215] An image feature extraction module, which is used to divide the second recognition image into non-overlapping units according to 8×8 pixels, directly cut off and discard the units that do not meet the size, count the distribution of the LBP values of all pixel points inside each unit to generate a histogram of a unit, and splice the histograms of all units end to end according to the unit order to generate an image feature vector as the final feature representation of the visual signal;
[0216] A feature fusion module for converting speech feature vectors and image feature vectors into matrix representations, dynamically adjusting the weights of speech features and image features through a cross-modal attention mechanism, and generating a fusion weight matrix;
[0217] A model construction module for constructing a deep learning network model, using the original speech signal and the original visual image as the training set, and using the speech command corresponding to the fusion weight matrix as the label, and inputting them into the deep learning network model for training;
[0218] A speech recognition module that inputs the real-time collected speech signal and visual image into the trained deep learning network model, obtains the current fusion weight matrix, and determines the current speech command according to the fusion weight matrix.
[0219] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0220] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0221] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0222] As described above, only the specific implementation manners of this application are provided, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application.
Claims
1. A speech recognition method based on artificial intelligence, characterized in that, The specific steps include: Step 1: Collect the original speech signal, perform noise reduction processing on the original speech signal, decompose the speech signal into multiple sub-bands with different frequencies through wavelet transform. Each sub-band corresponds to the information of the signal at different time scales. Perform threshold processing on the high-frequency sub-bands to weaken the noise components, and recombine the processed sub-bands to obtain the denoised speech signal; Step 2: Pre-emphasize the denoised speech signal through a high-pass filter, segment the pre-emphasized speech signal into short-time frames of 20 - 40 ms, with 50% overlap between frames. Window each frame signal, convert the windowed frame signal from the time domain signal to the frequency domain signal through the fast Fourier transform. Generate Mel cepstral coefficients from the frequency domain signal, and combine the Mel cepstral coefficients and their first-order and second-order differences into a speech feature vector; Step 3: Real-time collect the lip movement information of the speaker as the original visual image, perform grayscale processing on the original image to generate the first recognition image; Step 4: Normalize the pixels of the first recognition image. For each pixel point in the first recognition image, select 8 neighboring pixel points around the pixel point, compare the gray value of the central pixel point with its neighboring pixel points to generate a binary number, convert the binary number to a decimal number to generate the LBP value of the central pixel point, and calculate the LBP values of all pixel points in the first recognition image to form a new image, called the second recognition image; Step 5: Divide the second recognition image into non-overlapping units of 8×8 pixels, directly crop and discard the units that do not meet the size. Statistically analyze the LBP value distribution of all pixel points inside each unit to generate a histogram of a unit, and splice the histograms of all units end to end in the unit order to generate an image feature vector as the final feature representation of the visual signal; Step 6: Convert the speech feature vector and the image feature vector into matrix representations, and dynamically adjust the weights of the speech feature and the image feature through a cross-modal attention mechanism to generate a fusion weight matrix; Step 7: Construct a deep learning network, use the original speech signal and the original visual image as the training set, and use the speech command corresponding to the fusion weight matrix as the label, and input them into the deep learning network for training; Step 8: Input the real-time collected speech signal and visual image into the trained deep learning network, obtain the current fusion weight matrix, and determine the current speech command according to the fusion weight matrix.
2. The method for speech recognition based on artificial intelligence according to claim 1, wherein: The principle for judging the high-frequency sub-band range in Step 1 is: Among them, represents the frequency of the high-frequency sub-band, represents the sampling rate of the original signal, represents the corresponding layer number in the wavelet transform.
3. The method for speech recognition based on artificial intelligence according to claim 1, characterized in that: The formula for decomposing the speech signal into multiple sub-bands with different frequencies through wavelet transform in Step 1 is: Among them, represents the Haar wavelet basis function, represents the time variable, represents the wavelet coefficients of the high-frequency subband, represents the original signal, represents the result of scaling and translation of the wavelet basis function and is the scaling parameter, is the translation parameter; Set a noise threshold, and generate sub-band coefficients according to the noise threshold and wavelet coefficients. The formula is: Among them, represents the sub-band coefficient after processing, represents the noise threshold; Perform inverse wavelet transform on the processed sub-band coefficients to generate the denoised signal. The formula is: Among them, represents the signal after denoising, represents the subband coefficients, represents the result of stretching and translating the wavelet basis function and is the stretching scale parameter, is the translation scale parameter.
4. The voice recognition method based on artificial intelligence according to claim 1, characterized in that: The principle for generating Mel frequency cepstral coefficients in Step 2 is: First, pre-emphasize the denoised signal. The formula is: Among them, represents the value corresponding to the signal after pre-emphasis at time ; represents the value corresponding to the signal after denoising at time ; represents the value corresponding to the signal after denoising at time ; represents the pre-emphasis coefficient, represents the time variable; After the pre-emphasized signal is segmented into short-time frames, each frame of the signal is windowed, and the formula is as follows: Among them, represents the result after windowing the segmented frame signal, represents the segmented frame signal, represents the time variable, represents the Hamming window function, represents the time index, represents the frame length; The formula for converting the windowed frame signal from the time-domain signal to the frequency-domain signal is as follows: Among them, Z represents the frequency-domain signal, represents the time-domain signal, represents the frame length, represents the imaginary unit, represents the frequency index, and its value range is [0, N - 1], represents the time index, and its value range is [0, N - 1]; The formula for the energy output of each filter is as follows: Among them, represents the value of the th Mel filter, represents the frequency index, represents the th frequency of the Mel filter; Taking the logarithm of the energy output of each filter to generate logarithmic energy, and performing discrete cosine transform on the logarithmic energy to generate Mel cepstral coefficients, the formula is as follows: Among them, among them, represents the logarithmic energy of the th filter, represents the frequency-domain signal, represents the numerical value of the th Mel filter, represents the frequency index, represents the frame length, represents the th Mel-frequency cepstral coefficient, represents the number of Mel filters, represents the logarithmic energy of the th Mel filter, represents the index of the Mel-frequency cepstral coefficient; The formula for calculating the first-order difference and second-order difference of Mel frequency cepstral coefficients is as follows: Among them, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient, represents the frame offset, determined by the size of the difference window, and the size of the difference window is , and the specific value is determined by the actual application scenario; The principle for combining Mel frequency cepstral coefficients, first-order differences, and second-order differences into a feature vector is as follows: Among them, represents the final speech feature vector, represents the th Mel-frequency cepstral coefficient, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient, and represents the number of Mel-frequency cepstral coefficients.
5. A speech recognition method based on artificial intelligence according to claim 1, characterized in that: In step 3, the original image is grayscale processed to generate the first recognition image, and the formula is as follows: Among them, represents the gray value of the pixel point, represents the red channel value of the pixel point of the original image, represents the green channel value of the original image pixel point, represents the blue channel value of the pixel point of the original image.
6. The method for speech recognition based on artificial intelligence according to claim 1, characterized in that: In step 4, comparing the gray value of the central pixel with the gray values of its neighboring pixels to generate a binary number, the principle is as follows: For the central pixel with 8 neighboring pixels in the neighborhood, it can be expressed as: The formula for judging the value of the binary number is as follows: Among them, represents the gray value of the central pixel point, represents the gray value of the neighboring pixel points, is a comparison formula for judging based on the size relationship between the central pixel point and the neighboring pixel points; The formula for converting the binary number into the LBP value of the pixel is as follows: Among them, represents the LBP value of the central pixel point, represents the number of neighborhood pixel points, represents a comparison formula for judgment based on the size relationship between the central pixel point and the neighborhood pixel points.
7. A speech recognition method based on artificial intelligence according to claim 1, characterized in that: The principle for generating a histogram of a unit in step 5 is as follows: Statistical distribution of LBP values of all pixels within a unit, dividing the LBP values into 256 intervals according to the range of [0, 255], each interval corresponding to an LBP value, generating a histogram vector of length 256, where each element represents the number of pixel LBP values within the corresponding interval; The formula for connecting the histograms of all units in sequence to generate a feature vector is as follows: Among them, represents the final image feature vector, represents the th unit's number of LBP values in the th interval of the histogram, represents the number of units.
8. A speech recognition method based on artificial intelligence according to claim 1, characterized in that: The principle for generating the fusion weight matrix in step 6 is as follows: Regarding the speech feature vector as a 1×3L matrix, it can be expressed as: Among them, represents the speech feature matrix, represents the th Mel-frequency cepstral coefficient, represents the first-order difference of the th Mel-frequency cepstral coefficient, represents the second-order difference of the th Mel-frequency cepstral coefficient; represents the number of Mel-frequency cepstral coefficients; Regarding the image feature vector as a 1×256E matrix, it can be expressed as: Among them, represents the image feature matrix, represents the th unit's number of LBP values in the th interval of the histogram, represents the number of units; The speech feature matrix and the image feature matrix take the same number of elements to generate their similarity, and the formula is as follows: Among them, represents the similarity between the corresponding elements of the speech feature matrix and the image feature matrix, represents the column elements of the speech feature matrix, represents the column elements of the image feature matrix; Arrange all the calculated similarity values in sequence to form a matrix , and the matrix is the correlation matrix between the speech feature matrix and the image feature matrix; Calculating the attention weight according to the correlation matrix, and the formula is as follows: Among them, represents the attention weight of the speech feature, represents the attention weight of the image feature, represents the correlation matrix, represents the transpose matrix of the correlation matrix; Using the calculated weight for feature fusion to generate the fusion weight matrix, and the formula is as follows: Among them, represents the fusion weight matrix, represents the attention weight of the speech feature, represents the speech feature matrix, represents the attention weight of the image feature, represents the image feature matrix.
9. A voice recognition system based on artificial intelligence, characterized in that: The system is used to implement the artificial intelligence-based speech recognition method described in any one of the above claims 1-, including: A speech acquisition module for acquiring the original speech signal, performing noise reduction processing on the original speech signal, decomposing the speech signal into multiple sub-bands of different frequencies through wavelet transform, each sub-band corresponding to the information of the signal at different time scales, performing threshold processing on the high-frequency sub-bands to weaken the noise components, and recombining the processed sub-bands to obtain the denoised speech signal; A speech feature extraction module for pre-emphasizing the denoised speech signal through a high-pass filter, segmenting the pre-emphasized speech signal into short-time frames of 20~40ms, overlapping the frames by 50% of the frame length, windowing each frame of the signal, converting the windowed frame signal from the time-domain signal to the frequency-domain signal through fast Fourier transform, generating Mel cepstral coefficients from the frequency-domain signal through Mel filters, and combining the Mel cepstral coefficients and their first-order differences and second-order differences into a speech feature vector; An image acquisition module, which is used to collect the lip movement information of the speaker in real time as the visual original image, perform grayscale processing on the original image, and generate the first recognition image; An image processing module, which is used to perform normalization processing on the pixels of the first recognition image. For each pixel point in the first recognition image, select 8 neighboring pixel points around the pixel point, compare the gray value of the central pixel point with its neighboring pixel points, generate a binary number, convert the binary number into a decimal number to generate the LBP value of the central pixel point, calculate the LBP values of all pixel points in the first recognition image to form a new image, which is called the second recognition image; An image feature extraction module, which is used to divide the second recognition image into non-overlapping units according to 8×8 pixels, directly cut and discard the units that do not meet the size, count the distribution of the LBP values of all pixel points inside each unit, generate a histogram of a unit, splice the histograms of all units in sequence from beginning to end, generate an image feature vector, which is used as the final feature representation of the visual signal; A feature fusion module, which is used to convert the speech feature vector and the image feature vector into matrix representations, dynamically adjust the weights of the speech feature and the image feature through a cross-modal attention mechanism, and generate a fusion weight matrix; A model construction module, which is used to construct a deep learning network model, use the speech original signal and the visual original image as the training set, use the speech command corresponding to the fusion weight matrix as the label, and input them into the deep learning network model for training; A speech recognition module, which inputs the real-time collected speech signal and visual image into the trained deep learning network model, obtains the current fusion weight matrix, and determines the current speech command according to the fusion weight matrix.
Citation Information
Patent Citations
An AI-based speech recognition method
CN116580706B
PCNN spectrogram feature integration based emotion voice recognition system
CN107845390A
Noise-robust audio and video bimodal speech recognition method and system
CN111754992A
Cross-modal multi-feature fusion audio and video speech recognition method and system
CN112053690A
Voiceprint and face recognition verification system and method based on trusted execution environment
CN112491844A
Cited By
Induction cooker control method based on voice recognition
CN120954405A
Method for controlling an induction cooker based on voice recognition
CN120954405B
Multi-agent-based intention recognition method and device, equipment and medium
CN121034296A
An intention recognition method, device and equipment based on multi-agent, and a medium
CN121034296B
Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement
CN121281514A