Artificial intelligence-based speech recognition method and device, computer device and medium
By combining audio and image features in the audiovisual speech model and optimizing video encoder parameters using feature clustering and fully connected layers, the problem of poor speech recognition performance in unlabeled cases is solved, and more efficient speech recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2026-04-07
AI Technical Summary
Existing audiovisual speech models perform poorly in unlabeled speech recognition, resulting in high training costs and limiting their development and application.
Audiovisual features are obtained by adding audio features and image features. Then, the parameters of the video encoder are optimized by combining the predicted clustering results and the reference clustering results using a feature clustering model and a fully connected layer, thus achieving label-free training.
The speech recognition performance of the audiovisual speech model was improved in the case of no labels, and the training cost was reduced.
Smart Images

Figure CN116403569B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a speech recognition method, device, computer equipment, and medium based on artificial intelligence. Background Technology
[0002] Existing audiovisual speech models, which combine audio and visual modalities for pre-training, exhibit good speech recognition performance and high robustness. However, to improve speech recognition accuracy, these models require extensive labeled data for parameter optimization during training, leading to high training costs and limiting their development and application. Furthermore, lightweight or unlabeled methods significantly restrict the parameter optimization process of existing audiovisual speech models, thus reducing their speech recognition performance.
[0003] Therefore, how to improve the speech recognition performance of audiovisual speech models in the absence of labels has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a speech recognition method, device, computer equipment and medium based on artificial intelligence to solve the problem of poor speech recognition performance of audiovisual speech models in the absence of labels.
[0005] In a first aspect, embodiments of the present invention provide an artificial intelligence-based speech recognition method, the speech recognition method comprising:
[0006] Obtain preset audio and video, input the audio into a trained audio encoder to obtain M audio features, and input the video into a video encoder to obtain M image features, where M is an integer greater than 1;
[0007] Each of the audio features and each of the image features are added one by one to obtain M audiovisual features. The audio features, the image features and the M fused features of the audiovisual features are input into the trained feature clustering model to obtain the predicted clustering result. The predicted clustering result includes M1 feature sets, where M1 is a positive integer less than M.
[0008] The M audio features are input into the trained feature clustering model to obtain a reference clustering result. The reference clustering result is used as the training label of the video encoder. The parameters of the video encoder are updated using the training label and the predicted clustering result to obtain a trained video encoder. The reference clustering result includes M2 feature sets, where M2 is a positive integer less than M.
[0009] Acquire target audio and target video, input the target audio into the trained audio encoder to obtain M target audio features, and input the target video into the trained video encoder to obtain M target image features;
[0010] Each of the target audio features and each of the target image features are added one by one to obtain M target audiovisual features. The M target fusion features of the target audio features, the target image features and the target audiovisual features are then input into the trained fully connected layer to obtain the speech recognition result.
[0011] Secondly, embodiments of the present invention provide an artificial intelligence-based speech recognition device, the speech recognition device comprising:
[0012] The first feature extraction module is used to acquire preset audio and video, input the audio into a trained audio encoder to obtain M audio features, and input the video into a video encoder to obtain M image features, where M is an integer greater than 1;
[0013] The feature clustering module is used to add each of the audio features to each of the image features one by one to obtain M audiovisual features. The audio features, the image features and the M fused features of the audiovisual features are input into the trained feature clustering model to obtain the predicted clustering result. The predicted clustering result includes M1 feature sets, where M1 is a positive integer less than M.
[0014] The parameter update module is used to input the M audio features into the trained feature clustering model to obtain a reference clustering result, use the reference clustering result as the training label of the video encoder, and update the parameters of the video encoder using the training label and the predicted clustering result to obtain a trained video encoder. The reference clustering result includes M2 feature sets, where M2 is a positive integer less than M.
[0015] The second feature extraction module is used to acquire target audio and target video, input the target audio into the trained audio encoder to obtain M target audio features, and input the target video into the trained video encoder to obtain M target image features;
[0016] The speech recognition module is used to add each of the target audio features and each of the target image features one by one to obtain M target audiovisual features, and input the target audio features, the target image features and the M target audiovisual features into the trained fully connected layer to obtain the speech recognition result.
[0017] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech recognition method as described in the first aspect.
[0018] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech recognition method as described in the first aspect.
[0019] The beneficial effects of this invention compared to existing technologies are as follows: By inputting audio into a trained audio encoder to obtain M audio features, inputting video into a video encoder to obtain M image features, and then obtaining M audiovisual features and M fusion features, the M fusion features are input into a trained feature clustering model to obtain a predicted clustering result, the M audio features are input into a trained feature clustering model to obtain a reference clustering result, and the reference clustering result is used as the training label for the video encoder to obtain a trained video encoder. Then, the target audio is input into the trained audio encoder to obtain M target audio features, the target video is input into the trained video encoder to obtain M target image features, and then obtaining M target audiovisual features and M target fusion features, the M target fusion features are input into a trained fully connected layer to obtain a speech recognition result. In the absence of labels, a preset clustering algorithm is used to obtain a reference clustering result, which is then used as the training label for the video encoder. This makes the speech information represented by the predicted clustering result close to the speech information represented by the reference clustering result, resulting in a trained video encoder, which effectively improves the speech recognition effect. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of an application environment for an artificial intelligence-based speech recognition method provided in Embodiment 1 of the present invention;
[0022] Figure 2 This is a flowchart illustrating an artificial intelligence-based speech recognition method provided in Embodiment 1 of the present invention.
[0023] Figure 3 This is a schematic diagram of the structure of an artificial intelligence-based speech recognition device provided in Embodiment 2 of the present invention;
[0024] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Detailed Implementation
[0025] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0026] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0027] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0028] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0029] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0030] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0031] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0032] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0033] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0034] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0035] The first embodiment of this invention provides an artificial intelligence-based speech recognition method, which can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0036] See Figure 2This is a flowchart illustrating an artificial intelligence-based speech recognition method provided in Embodiment 1 of the present invention. The speech recognition method described above can be applied to... Figure 1 In a client application, the speech recognition method may include the following steps:
[0037] Step S201: Obtain preset audio and video, input the audio into the trained audio encoder to obtain M audio features, and input the video into the video encoder to obtain M image features.
[0038] The preset video consists of M consecutive frames, and the preset audio corresponds to the preset video. The preset audio is segmented according to the duration of each frame to obtain M sub-audio files that correspond one-to-one with the M frames, where M is an integer greater than 1.
[0039] In this embodiment, the audiovisual speech model may include an audio encoder, a video encoder, and a fully connected layer. When performing feature extraction on video and audio to complete speech recognition, the audio encoder first extracts features from the audio to obtain M audio features corresponding to M sub-audio. Then, the video encoder extracts features from the video to obtain M image features corresponding to M frames.
[0040] Since audio features are obtained by extracting features from audio using a trained audio encoder, they can effectively represent the speech information contained in the audio. However, image features are obtained by extracting features from video using an untrained video encoder, and their ability and performance in representing speech information contained in the video are relatively poor. Therefore, in this embodiment, in order to improve the ability of the video encoder to represent speech information contained in the video, the extracted audio features and image features are used as the basis for subsequent parameter optimization of the video encoder. By optimizing, a trained video encoder is obtained, thereby improving the speech recognition performance of the audiovisual speech model in the absence of labels.
[0041] The steps described above—obtaining preset audio and video, inputting the audio into a trained audio encoder to obtain M audio features, and inputting the video into a video encoder to obtain M image features—result in audio features that can effectively represent the speech information contained in the audio, and image features that have a weaker ability to represent the speech information contained in the corresponding video. These serve as the basis for optimizing the video encoder and improve the speech recognition performance of the audiovisual speech model.
[0042] Step S202: Add each audio feature to each image feature one by one to obtain M audiovisual features. Input the M fused features of audio features, image features and audiovisual features into the trained feature clustering model to obtain the predicted clustering result, which includes M1 feature sets.
[0043] Among them, for the corresponding audio and video, the audio and video features can be combined to obtain the audio-visual features. These audio-visual features can combine sound information and image information to more effectively represent the corresponding speech information. As the optimization basis of the video encoder and the basis of speech recognition, they can effectively improve the speech recognition effect of the audio-visual speech model.
[0044] Therefore, when performing speech recognition, the audio and video features are first combined to obtain the audio features, and then the audio features, video features and audio features are further fused to optimize the video encoder based on the fused features.
[0045] Specifically, M audio features are added one by one to M image features to obtain M audiovisual features. Based on the audio, image, and audiovisual features, M fused features are obtained. These M fused features are then input into a trained feature clustering model to obtain a predicted clustering result. This model clusters similar fused features together to obtain a predicted clustering result comprising M1 feature sets. This predicted clustering result can effectively predict the distribution information representing the M fused features based on the number of feature sets, the composition of each feature set, and their location information, thereby further predicting the speech information contained in the corresponding audio and video. Here, M1 is a positive integer less than M.
[0046] In one embodiment, M fusion features are obtained based on audio features, image features, and audiovisual features. Each audio feature, as well as the corresponding image features and audiovisual features, can be spliced and fused in the same order to obtain the corresponding fusion feature.
[0047] Optionally, the training process for the audio encoder and the feature clustering model includes:
[0048] The sample audio is input into the audio encoder to obtain M sample audio features. The M sample audio features are then input into the feature clustering model to obtain the sample prediction clustering results.
[0049] The sample audio is clustered according to a preset clustering algorithm to obtain the sample reference clustering results;
[0050] The loss function is calculated based on the sample prediction clustering results and the sample reference clustering results. The parameters of the audio encoder and the feature clustering model are then corrected in reverse using the gradient descent method until the loss function converges, resulting in a trained audio encoder and a trained feature clustering model.
[0051] The training samples for the audio encoder and the feature clustering model are a large number of sample audios. By inputting the sample audios into the audio encoder, M sample audio features are obtained. These sample audio features are used to represent the semantic information contained in the corresponding sample audios. Then, the M sample audio features are input into the feature clustering model. By clustering similar sample audio features into a feature set, a sample prediction clustering result containing several feature sets is obtained. This sample prediction clustering result can effectively predict the distribution information representing the M sample audio features based on the number of feature sets, the composition of each feature set, and the location information, thereby further predicting the speech information contained in the corresponding sample audios.
[0052] Meanwhile, in order to optimize the parameters of the audio encoder and the feature clustering model, in this embodiment, the sample audio is clustered by a preset clustering algorithm to obtain the sample reference clustering result. The sample reference clustering result can effectively represent the speech information contained in the sample audio. Therefore, the sample reference clustering result is used as the training label to optimize the parameters of the audio encoder and the feature clustering model, so as to improve the feature extraction accuracy of the audio encoder and the feature clustering accuracy of the feature clustering model.
[0053] Specifically, a loss function is calculated based on the sample prediction clustering results and the sample reference clustering results. The magnitude of this loss function can characterize the difference between the sample prediction clustering results obtained by the audio encoder and the feature clustering model and the sample reference clustering results obtained by the preset clustering algorithm. When the difference is large, it indicates that the feature extraction accuracy of the audio encoder and the clustering accuracy of the feature clustering model are low. It is necessary to correct the parameters of the audio encoder and the feature clustering model in reverse according to the gradient descent method until the loss function converges, so as to obtain a trained audio encoder and a trained feature clustering model, thereby improving the speech recognition effect of the audiovisual speech model.
[0054] Optionally, the sample audio is clustered according to a preset clustering algorithm to obtain sample reference clustering results, including:
[0055] The sample audio is segmented into frames to obtain M sub-audio files. Mel-cepstrum analysis is then performed on the sub-audio files to obtain the Mel-spectral cepstrum coefficients of the M sub-audio files.
[0056] Obtain N preset values, and use the mean clustering algorithm to cluster M Mel-spectrum cepstral coefficients under each preset value to obtain N sample clustering results. The sample clustering results include M3 feature sets, where N is an integer greater than 1 and M3 is a positive integer less than M.
[0057] Calculate the dispersion of each sample clustering result, and determine the sample clustering result corresponding to the dispersion that meets the preset conditions as the sample reference clustering result.
[0058] Mel-Cepstral Analysis is a method for analyzing the spectrum of speech based on the mechanism of human hearing. Since the pitch of a sound that the human ear can hear is not linearly proportional to its frequency, the frequency scale of Mel-Cepstral Analysis is more in line with the auditory characteristics of the human ear, and can obtain Mel-Cepstral coefficients that effectively represent speech information.
[0059] Specifically, the sample audio is first segmented into frames to obtain M sub-audio units. Then, Mel-Cepstral Analysis is performed on each sub-audio unit to obtain Mel-Cepstral Coefficients, which represent the speech information of the corresponding sub-audio units. Next, to determine the distribution information of the sample audio and further characterize the speech information contained within it, in this embodiment, the mean-clustering algorithm is used to cluster the M Mel-Cepstral Coefficients, resulting in N sample clustering results.
[0060] Specifically, to improve the accuracy of sample clustering results, in this embodiment, N preset values are first obtained. Under any preset value, the mean clustering algorithm is used to cluster M Mel-spectrum cepstral coefficients. By clustering similar Mel-spectrum cepstral coefficients into a feature set, a sample clustering result containing M3 feature sets is obtained, where M3 is consistent with the corresponding preset value. Therefore, N clustering operations are performed on the M Mel-spectrum cepstral coefficients based on the N preset values to obtain N sample clustering results.
[0061] Since different preset values can affect the clustering results when clustering Mel-spectrum cepstral coefficients, in this embodiment, in order to improve the parameter optimization effect of the video encoder and feature clustering model, the best clustering result is further selected from the N sample clustering results and determined as the sample reference clustering result, so as to use the sample reference clustering result as the sample label for parameter optimization.
[0062] Specifically, in this embodiment, for the obtained N sample clustering results, the dispersion of each sample clustering result is calculated. The dispersion can reflect the representativeness of the cluster center of the sample clustering result to each Mel-spectrum cepstral coefficient. Therefore, by determining the sample clustering result corresponding to the dispersion that meets the preset conditions as the sample reference clustering result, the best sample clustering result can be selected from the N sample clustering results, thereby improving the accuracy of the sample reference clustering result.
[0063] Optionally, the dispersion of each sample clustering result is calculated, and the sample clustering result corresponding to the dispersion that meets the preset conditions is determined as the sample reference clustering result, including:
[0064] For any sample clustering result, determine the cluster center of the sample clustering result, calculate the distance between each Mel-spectrum cepstral coefficient in the sample clustering result and the cluster center, and calculate the dispersion of the sample clustering result based on the distance;
[0065] Compare the dispersion of the clustering results of N samples, and determine the sample clustering result corresponding to the minimum dispersion as the sample reference clustering result.
[0066] The dispersion reflects the representativeness of the cluster centers of the sample clustering results to each Mel-spectral cepstral coefficient. In this embodiment, for any sample clustering result, the cluster centers of the sample clustering result are first determined. The distance between each Mel-spectral cepstral coefficient in the sample clustering result and the cluster center is calculated, and the dispersion of the sample clustering result is calculated based on the distance. Therefore, the smaller the dispersion, the higher the representativeness of the cluster centers of the sample clustering result to each Mel-spectral cepstral coefficient.
[0067] Therefore, by comparing the dispersion of the clustering results of N samples, the clustering result corresponding to the minimum dispersion is determined as the sample reference clustering result, so as to effectively characterize the distribution information of the Mel-spectral cepstral coefficients of M samples, and further characterize the speech information contained in the corresponding sample audio.
[0068] In one embodiment, the distance between each Mel-spectral cepstral coefficient and each cluster center in the sample clustering result is calculated, and the mean distance is used as the dispersion of the sample clustering result.
[0069] The above steps involve adding each audio feature to each image feature to obtain M audiovisual features, then inputting the M fused features of audio, image, and audiovisual features into a trained feature clustering model to obtain a predicted clustering result. The predicted clustering result includes M1 feature sets. By fusing audio, image, and audiovisual features and obtaining the predicted clustering result based on the trained feature clustering model, the distribution information representing the fused features is effectively predicted, thereby further predicting the speech information contained in the corresponding audio and video, improving the speech information representation ability, and thus improving the speech recognition effect.
[0070] Step S203: Input M audio features into the trained feature clustering model to obtain reference clustering results. Use the reference clustering results as training labels for the video encoder. Use the training labels and predicted clustering results to update the parameters of the video encoder to obtain a trained video encoder. The reference clustering results include M2 feature sets.
[0071] The M audio features are obtained by feature extraction from the audio using a trained audio encoder, effectively representing the speech information contained in the audio. In this embodiment, the M audio features are further input into a trained feature clustering model to obtain a reference clustering result, which includes M² feature sets, where M² is a positive integer less than M. Based on the number of corresponding feature sets, the composition of each feature set, and their location information, the distribution information of the M audio features can be effectively represented, thereby further representing the speech information contained in the corresponding audio.
[0072] Since the reference clustering result is obtained by processing audio using a trained audio encoder and a trained feature clustering model, it can effectively represent the speech information contained in the audio. The predicted clustering result is obtained by processing audio and corresponding video using a trained audio encoder, a video encoder to be trained, and a trained feature clustering model, and can predict and represent the speech information contained in the audio and video. However, since the video encoder is not trained, the accuracy of predicting and representing the speech information contained in the audio and video is low. Therefore, it is necessary to use the reference clustering result, which can effectively represent the speech information contained in the audio, as the training label to continuously train the parameters of the video encoder so that after the video encoder is trained, the reference clustering result and the predicted clustering result can represent the same speech information.
[0073] Therefore, the reference clustering results are used as training labels for the video encoder. The parameters of the video encoder are updated using the training labels and the predicted clustering results, so that the predicted clustering results are closer to the reference clustering results. This allows the predicted clustering results and the reference clustering results to represent the same speech information, thus obtaining a well-trained video encoder and improving the speech recognition performance.
[0074] Optionally, the reference clustering results are used as training labels for the video encoder. The parameters of the video encoder are updated using the training labels and the predicted clustering results to obtain a trained video encoder, including:
[0075] Substitute the predicted clustering results and the reference clustering results into the preset loss function to calculate the clustering loss;
[0076] Based on the clustering loss, adjust the parameters of the video encoder until the clustering loss converges, and obtain the trained video encoder.
[0077] The predicted clustering result includes M fused features and M1 feature sets obtained from clustering, while the reference clustering result includes M audio features and M2 feature sets obtained from clustering. The greater the difference in the position of features and the difference in the number of feature sets between the predicted clustering result and the reference clustering result, the greater the clustering loss of the video encoder.
[0078] Therefore, in this embodiment, the M feature distances between the M fusion features in the predicted clustering result and the M audio features in the reference clustering result are first calculated, and the M feature distances and the number of the two feature sets are substituted into the preset loss function to calculate the clustering loss. Then, the parameters of the video encoder are adjusted based on the clustering loss until the clustering loss converges, and the trained video encoder is obtained.
[0079] The preset loss function formula is as follows:
[0080]
[0081] In the formula, L is the clustering loss function, M1 is the number of feature sets in the predicted clustering result, M2 is the number of feature sets in the reference clustering result, M is the number of fused features and audio features, and d i Let be the feature distance between the i-th fused feature and the i-th audio feature, where i = 1, 2, ..., M.
[0082] The above steps involve inputting M audio features into a trained feature clustering model to obtain a reference clustering result, using the reference clustering result as the training label for the video encoder, and updating the parameters of the video encoder using the training label and the predicted clustering result to obtain a trained video encoder. The reference clustering result includes M2 feature sets. By making the predicted clustering result closer to the reference clustering result, the predicted clustering result and the reference clustering result can represent the same speech information, thus obtaining a trained video encoder and improving the speech recognition effect.
[0083] Step S204: Obtain the target audio and target video. Input the target audio into the trained audio encoder to obtain M target audio features. Input the target video into the trained video encoder to obtain M target image features.
[0084] Among them, the target audio and target video are the audio and video to be performed on speech recognition, and the target video and target audio correspond to each other.
[0085] The target audio is input into a trained audio encoder to obtain M target audio features, which can effectively represent the speech information contained in the audio. The target video is input into a trained video encoder to obtain M target image features, which can effectively represent the speech information contained in the video. The obtained M target audio features and M target image features are used as the basis for speech recognition, which improves the speech recognition effect.
[0086] Optionally, after obtaining the trained video encoder, the following may also be included:
[0087] The parameters of the trained audio encoder and the trained video encoder are fine-tuned to obtain the fine-tuned audio encoder and the fine-tuned video encoder.
[0088] Accordingly, the target audio is input into a trained audio encoder to obtain target audio features, and the target video is input into a trained video encoder to obtain target image features, including:
[0089] The target audio is input into a finely tuned audio encoder to obtain the target audio features;
[0090] The target video is input into a finely tuned video encoder to obtain the target image features.
[0091] The video encoder is an optimization based on the audio encoder. In order to further improve the feature extraction effect of the audio encoder and the video encoder, the parameters of the trained audio encoder and the trained video encoder are fine-tuned to obtain fine-tuned audio encoder and fine-tuned video encoder, which are used to extract features from target audio and target video to improve speech recognition performance.
[0092] Optionally, the parameters of the trained audio encoder and the trained video encoder are fine-tuned to obtain fine-tuned audio encoder and fine-tuned video encoder, including:
[0093] The audio is input into the trained audio encoder to obtain the second audio feature, and the video is input into the trained video encoder to obtain the second image feature;
[0094] The second audio feature is added to the second image feature to obtain the second audiovisual feature. The second audio feature, the second image feature and the second audiovisual feature are fused. The feature fusion result is input into the trained feature clustering model to obtain the second prediction clustering result.
[0095] The feature fusion results are clustered using a pre-defined clustering algorithm to obtain a second reference clustering result;
[0096] The second clustering loss is calculated based on the second predicted clustering result and the second reference clustering result. The parameters of the trained audio encoder and the trained video encoder are updated based on the second clustering loss until the second clustering loss converges, resulting in the fine-tuned audio encoder and the fine-tuned video encoder.
[0097] The training samples consist of a large number of corresponding audio and video samples. The audio is input into the trained audio encoder to obtain the second audio feature, and the video is input into the trained video encoder to obtain the second image feature. The second audio feature and the second image feature are then added together to obtain the second audiovisual feature. The second audiovisual feature combines the speech information represented by the second audio feature and the second image feature, thereby improving the ability to represent speech information.
[0098] Then, the feature fusion result is input into the trained feature clustering model to obtain the second predicted clustering result. The feature distribution information representing the feature fusion result is used to predict and represent speech information. A preset clustering algorithm is used to cluster the feature fusion result to obtain the second reference clustering result to represent speech information. The second reference clustering result is used as the training label. The second clustering loss is calculated based on the second predicted clustering result and the second reference clustering result. Based on the second clustering loss, the parameters of the trained audio encoder and the trained video encoder are updated until the second clustering loss converges, so that the second predicted clustering result is closer to the second reference clustering result, and the second predicted clustering result and the second reference clustering result can represent the same speech information, so as to obtain the fine-tuned audio encoder and the fine-tuned video encoder.
[0099] The steps described above—acquiring the target audio and target video, inputting the target audio into a trained audio encoder to obtain M target audio features, and inputting the target video into a trained video encoder to obtain M target image features—effectively characterize the speech information contained in the audio and video, serving as the basis for speech recognition and significantly improving speech recognition performance.
[0100] Step S205: Add each target audio feature to each target image feature one by one to obtain M target audiovisual features. Input the M target fusion features of the target audio features, target image features and target audiovisual features into the trained fully connected layer to obtain the speech recognition result.
[0101] In order to improve the accuracy of speech recognition, the target audio features are combined with the target image features to obtain the target audiovisual features, and the target audio features, target image features and target audiovisual features are fused to obtain the target fusion features, so as to improve the ability to represent speech information.
[0102] Specifically, for M target audio features and M target image features, each target audio feature and each target image feature is added one by one to obtain M target audiovisual features. These M target audio features, M target image features, and M target audiovisual features are then concatenated and fused in the same order to obtain M target fused features. Finally, the M target fused features are input into a trained fully connected layer to obtain the speech recognition result.
[0103] The above steps involve adding each target audio feature to each target image feature to obtain M target audiovisual features, and then inputting the M target fusion features of the target audio features, target image features, and target audiovisual features into a trained fully connected layer to obtain the speech recognition result. By combining the target audio features, target image features, and target audiovisual features, the ability to represent the speech information contained in the target audio and target video is improved, thus enhancing the speech recognition performance.
[0104] This invention provides an embodiment that obtains M audio features by inputting audio into a trained audio encoder, M image features by inputting video into a video encoder, and M audiovisual features and M fused features by inputting video into a trained feature clustering model. The M fused features are then input into a trained feature clustering model to obtain a predicted clustering result. The M audio features are also input into the trained feature clustering model to obtain a reference clustering result. This reference clustering result is used as the training label for the video encoder to obtain a trained video encoder. Next, the target audio is input into the trained audio encoder to obtain M target audio features, and the target video is input into the trained video encoder to obtain M target image features. This results in M target audiovisual features and M target fused features by inputting the M target fused features into a trained fully connected layer to obtain a speech recognition result. In the absence of labels, a preset clustering algorithm is used to obtain a reference clustering result, which is then used as the training label for the video encoder. This ensures that the speech information represented by the predicted clustering result is close to the speech information represented by the reference clustering result, resulting in a trained video encoder and effectively improving the speech recognition performance.
[0105] Corresponding to the speech recognition method in the above embodiments, Figure 3 A structural block diagram of an artificial intelligence-based speech recognition device provided in Embodiment 2 of the present invention is given. For ease of explanation, only the parts related to the embodiments of the present invention are shown.
[0106] See Figure 3 The voice recognition device includes:
[0107] The first feature extraction module 31 is used to acquire preset audio and video, input the audio into a trained audio encoder to obtain M audio features, and input the video into a video encoder to obtain M image features, where M is an integer greater than 1;
[0108] The feature clustering module 32 is used to add each audio feature to each image feature one by one to obtain M audiovisual features. The M fused features of audio features, image features and audiovisual features are input into the trained feature clustering model to obtain the predicted clustering result. The predicted clustering result includes M1 feature sets, where M1 is a positive integer less than M.
[0109] The parameter update module 33 is used to input M audio features into the trained feature clustering model to obtain a reference clustering result. The reference clustering result is used as the training label of the video encoder. The parameters of the video encoder are updated using the training label and the predicted clustering result to obtain the trained video encoder. The reference clustering result includes M2 feature sets, where M2 is a positive integer less than M.
[0110] The second feature extraction module 34 is used to acquire target audio and target video, input the target audio into the trained audio encoder to obtain M target audio features, and input the target video into the trained video encoder to obtain M target image features;
[0111] The speech recognition module 35 is used to add each target audio feature to each target image feature one by one to obtain M target audiovisual features. The M target fusion features of the target audio features, target image features and target audiovisual features are input into the trained fully connected layer to obtain the speech recognition result.
[0112] Optionally, the feature clustering module 32 mentioned above includes:
[0113] The first feature clustering submodule is used to input sample audio into the audio encoder to obtain M sample audio features, and input the M sample audio features into the feature clustering model to obtain the sample prediction clustering result;
[0114] The second feature clustering submodule is used to cluster the sample audio according to a preset clustering algorithm to obtain the sample reference clustering result;
[0115] The first parameter correction submodule is used to calculate the loss function based on the sample prediction clustering results and the sample reference clustering results, and to correct the parameters of the audio encoder and the feature clustering model in reverse according to the gradient descent method until the loss function converges, thus obtaining the trained audio encoder and the trained feature clustering model.
[0116] Optionally, the above-mentioned second feature clustering submodule includes:
[0117] The first feature extraction unit is used to perform frame segmentation on the sample audio to obtain M sub-audio, and to perform Mel-Cepstral Analysis on the sub-audio to obtain the Mel-Cepstral Coefficients of the M sub-audio.
[0118] The first feature clustering unit is used to obtain N preset values, and then use the mean clustering algorithm to cluster M Mel-spectrum cepstral coefficients under each preset value to obtain N sample clustering results. The sample clustering results include M3 feature sets, where N is an integer greater than 1 and M3 is a positive integer less than M.
[0119] The clustering result filtering unit is used to calculate the dispersion of the clustering results of each sample and determine the sample clustering results corresponding to the dispersion that meets the preset conditions as the sample reference clustering results.
[0120] Optionally, the above clustering result filtering unit includes:
[0121] The discreteness calculation subunit is used to determine the cluster center of any sample clustering result, calculate the distance between each Mel-spectrum cepstral coefficient in the sample clustering result and the cluster center, and calculate the discreteness of the sample clustering result based on the distance.
[0122] The clustering result filtering subunit is used to compare the dispersion of the clustering results of N samples and determine the sample clustering result corresponding to the minimum dispersion as the sample reference clustering result.
[0123] Optionally, the above parameter update module 33 includes:
[0124] The loss calculation submodule is used to substitute the predicted clustering results and reference clustering results into a preset loss function to calculate the clustering loss;
[0125] The second parameter correction submodule is used to adjust the parameters of the video encoder based on the clustering loss until the clustering loss converges, thus obtaining a trained video encoder.
[0126] Optionally, the second feature extraction module 34 mentioned above includes:
[0127] The parameter fine-tuning submodule is used to fine-tune the parameters of the trained audio encoder and the trained video encoder to obtain the fine-tuned audio encoder and the fine-tuned video encoder.
[0128] The audio feature extraction submodule is used to input the target audio into a finely tuned audio encoder to obtain the target audio features;
[0129] The image feature extraction submodule is used to input the target video into a finely tuned video encoder to obtain the target image features.
[0130] Optionally, the above parameter fine-tuning submodule includes:
[0131] The second feature extraction unit is used to input audio into a trained audio encoder to obtain second audio features, and to input video into a trained video encoder to obtain second image features.
[0132] The second feature clustering unit is used to add the second audio feature and the second image feature to obtain the second audiovisual feature, perform feature fusion on the second audio feature, the second image feature and the second audiovisual feature, and input the feature fusion result into the trained feature clustering model to obtain the second predicted clustering result;
[0133] The third feature clustering unit is used to cluster the feature fusion results using a preset clustering algorithm to obtain the second reference clustering result;
[0134] The parameter fine-tuning unit is used to calculate the second clustering loss based on the second predicted clustering result and the second reference clustering result. Based on the second clustering loss, the parameters of the trained audio encoder and the trained video encoder are updated until the second clustering loss converges, resulting in the fine-tuned audio encoder and the fine-tuned video encoder.
[0135] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0136] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, which, when executed by the processor, implements the steps in any of the above-described speech recognition method embodiments.
[0137] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0138] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0139] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of a computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0140] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0141] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0142] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0143] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0144] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0145] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0146] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech recognition method based on artificial intelligence, characterized in that, The speech recognition method includes: Obtain preset audio and video, input the audio into a trained audio encoder to obtain M audio features, and input the video into a video encoder to obtain M image features, where M is an integer greater than 1; Each of the audio features and each of the image features are added one by one to obtain M audiovisual features. The audio features, the image features and the M fused features of the audiovisual features are input into the trained feature clustering model to obtain the predicted clustering result. The predicted clustering result includes M1 feature sets, where M1 is a positive integer less than M. The M audio features are input into the trained feature clustering model to obtain a reference clustering result. The reference clustering result is used as the training label of the video encoder. The parameters of the video encoder are updated using the training label and the predicted clustering result to obtain a trained video encoder. The reference clustering result includes M2 feature sets, where M2 is a positive integer less than M. Acquire target audio and target video, input the target audio into the trained audio encoder to obtain M target audio features, and input the target video into the trained video encoder to obtain M target image features; Each of the target audio features and each of the target image features are added one by one to obtain M target audiovisual features. The M target fusion features of the target audio features, the target image features and the target audiovisual features are then input into the trained fully connected layer to obtain the speech recognition result.
2. The speech recognition method according to claim 1, characterized in that, The training process of the audio encoder and the feature clustering model includes: The sample audio is input into the audio encoder to obtain M sample audio features. The M sample audio features are then input into the feature clustering model to obtain the sample prediction clustering result. The sample audio is clustered according to a preset clustering algorithm to obtain sample reference clustering results; The loss function is calculated based on the sample prediction clustering result and the sample reference clustering result. The parameters of the audio encoder and the feature clustering model are then corrected in reverse using the gradient descent method until the loss function converges, thus obtaining the trained audio encoder and the trained feature clustering model.
3. The speech recognition method according to claim 2, characterized in that, The step of clustering the sample audio according to a preset clustering algorithm to obtain sample reference clustering results includes: The sample audio is segmented into frames to obtain M sub-audio files. Mel-frequency cepstral analysis is then performed on the sub-audio files to obtain the Mel-frequency cepstral coefficients of the M sub-audio files. N preset values are obtained, and the mean clustering algorithm is used to cluster the M Mel-spectrum cepstral coefficients under each preset value to obtain N sample clustering results. The sample clustering results include M3 feature sets, where N is an integer greater than 1 and M3 is a positive integer less than M. Calculate the dispersion of each sample clustering result, and determine the sample clustering result corresponding to the dispersion that meets the preset conditions as the sample reference clustering result.
4. The speech recognition method according to claim 3, characterized in that, The step of calculating the dispersion of each of the sample clustering results and determining the sample clustering result corresponding to the dispersion that meets the preset conditions as the sample reference clustering result includes: For any of the sample clustering results, determine the cluster center of the sample clustering result, calculate the distance between each of the Mel-spectral cepstral coefficients in the sample clustering result and the cluster center, and calculate the dispersion of the sample clustering result based on the distance; Compare the dispersion of the clustering results of the N samples, and determine the clustering result of the sample with the smallest dispersion as the sample reference clustering result.
5. The speech recognition method according to claim 1, characterized in that, The step of using the reference clustering results as training labels for the video encoder, and updating the parameters of the video encoder using the training labels and the predicted clustering results to obtain a trained video encoder includes: Substitute the predicted clustering results and the reference clustering results into a preset loss function to calculate the clustering loss; Based on the clustering loss, the parameters of the video encoder are adjusted until the clustering loss converges, thus obtaining a trained video encoder.
6. The speech recognition method according to any one of claims 1 to 5, characterized in that, After obtaining the trained video encoder, the following is also included: The parameters of the trained audio encoder and the trained video encoder are fine-tuned to obtain the fine-tuned audio encoder and the fine-tuned video encoder. Accordingly, the step of inputting the target audio into the trained audio encoder to obtain target audio features, and inputting the target video into the trained video encoder to obtain target image features, includes: The target audio is input into the finely tuned audio encoder to obtain the target audio features; The target video is input into the finely tuned video encoder to obtain the target image features.
7. The speech recognition method according to claim 6, characterized in that, The step of fine-tuning the parameters of the trained audio encoder and the trained video encoder to obtain fine-tuned audio encoder and fine-tuned video encoder includes: The audio is input into the trained audio encoder to obtain the second audio feature, and the video is input into the trained video encoder to obtain the second image feature; The second audio feature is added to the second image feature to obtain the second audiovisual feature. The second audio feature, the second image feature and the second audiovisual feature are fused. The feature fusion result is input into the trained feature clustering model to obtain the second predicted clustering result. The feature fusion results are clustered using a preset clustering algorithm to obtain a second reference clustering result; The second clustering loss is calculated based on the second predicted clustering result and the second reference clustering result. The parameters of the trained audio encoder and the trained video encoder are updated based on the second clustering loss until the second clustering loss converges, resulting in the fine-tuned audio encoder and the fine-tuned video encoder.
8. A speech recognition device based on artificial intelligence, characterized in that, The speech recognition device includes: The first feature extraction module is used to acquire preset audio and video, input the audio into a trained audio encoder to obtain M audio features, and input the video into a video encoder to obtain M image features, where M is an integer greater than 1; The feature clustering module is used to add each of the audio features to each of the image features one by one to obtain M audiovisual features. The audio features, the image features and the M fused features of the audiovisual features are input into the trained feature clustering model to obtain the predicted clustering result. The predicted clustering result includes M1 feature sets, where M1 is a positive integer less than M. The parameter update module is used to input the M audio features into the trained feature clustering model to obtain a reference clustering result, use the reference clustering result as the training label of the video encoder, and update the parameters of the video encoder using the training label and the predicted clustering result to obtain a trained video encoder. The reference clustering result includes M2 feature sets, where M2 is a positive integer less than M. The second feature extraction module is used to acquire target audio and target video, input the target audio into the trained audio encoder to obtain M target audio features, and input the target video into the trained video encoder to obtain M target image features; The speech recognition module is used to add each of the target audio features and each of the target image features one by one to obtain M target audiovisual features, and input the target audio features, the target image features and the M target audiovisual features into the trained fully connected layer to obtain the speech recognition result.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for controlling multidimensional film-watching system based on intelligent home device
CN103970892A
Receiver detection framework and method based on audio and face input
CN115620356A