A speech recognition method, apparatus, device, and storage medium

By combining acoustic and visual features in a fusion prediction method, the problem of low recognition accuracy of general speech recognition solutions in narration speech is solved, achieving efficient recognition of narration speech in specific fields and improving the user experience in narration scenarios.

CN116758912BActive Publication Date: 2026-06-02IFLYTEK CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2023-05-31
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing general speech recognition solutions perform poorly in recognizing narration, especially in specific fields such as game commentary, where the recognition accuracy is low.

Method used

By extracting the acoustic features of the target speech and combining them with the visual features of the target video, a pre-trained speech recognition model is used for fusion prediction, supplemented by visual content information from the target video for narration speech recognition.

Benefits of technology

It significantly improves the recognition accuracy of narration voice in specific fields, especially the recognition effect of proper words and rare words, thus improving the user experience in narration scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758912B_ABST
    Figure CN116758912B_ABST
Patent Text Reader

Abstract

The application provides a speech recognition method, device and equipment and a storage medium. The speech recognition method comprises: obtaining target speech and target video, wherein the target speech is commentary speech of video content of the target video; extracting acoustic features of the target speech to obtain acoustic features of the target speech, and extracting visual features containing video content information of the target video to obtain visual features of the target video; and determining a speech recognition result of the target speech according to the acoustic features of the target speech and the visual features of the target video. Considering that the target speech is commentary speech of video content of the target video, the target speech has a certain correlation with the video content of the target video. The application extracts visual features containing video content information of the target video, and performs speech recognition on the commentary speech with the aid of the visual features. When speech recognition is performed on the target speech, i.e. the commentary speech, the visual features containing video content information are used as an aid, and a more accurate speech recognition result can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and in particular to a speech recognition method, apparatus, device, and storage medium. Background Technology

[0002] Commentary voice generally refers to the audio provided by a commentator explaining the content of a video. For example, game commentary voice is the audio from a game commentator explaining the content of a game video. In certain fields, commentary voice recognition services are becoming increasingly important, with growing demand and increasingly higher quality requirements.

[0003] Most current speech recognition solutions are designed for general speech, and because they focus on recognizing general speech, they perform well in this area. However, narration is usually speech specific to a particular domain. Therefore, general speech recognition solutions are not very applicable to narration; that is, using general speech recognition solutions to recognize narration results in poor performance. Summary of the Invention

[0004] In view of this, the present invention provides a speech recognition method, apparatus, device, and storage medium to solve the problem that the recognition effect is poor when using a recognition scheme for general speech to recognize narration speech. The technical solution is as follows:

[0005] A speech recognition method, comprising:

[0006] Acquire target audio and target video, wherein the target audio is the narration audio of the video content of the target video;

[0007] Acoustic features are extracted from the target speech to obtain the acoustic features of the target speech, and visual features containing video content information are extracted from the target video to obtain the visual features of the target video;

[0008] The speech recognition result of the target speech is determined based on the acoustic features of the target speech and the visual features of the target video.

[0009] Optionally, the step of extracting acoustic features from the target speech to obtain the acoustic features of the target speech, and extracting visual features containing video content information from the target video to obtain the visual features of the target video; determining the speech recognition result of the target speech based on the acoustic features of the target speech and the visual features of the target video, includes:

[0010] The target speech and the target video are processed using a pre-trained speech recognition model to obtain the speech recognition result of the target speech; wherein:

[0011] The speech recognition model is trained using the first training data in the first training set. The first training data includes the first training video and the first training speech with text labeled with speech content. The first training speech is the narration of the video content of the first training video.

[0012] The training objective of the speech recognition model includes: making the speech recognition result predicted based on the acoustic features of the first training speech and supplemented by the visual features of the first training video consistent with the speech content text annotated by the first training speech.

[0013] Optionally, the step of processing the target speech and the target video using a pre-trained speech recognition model to obtain the speech recognition result of the target speech includes:

[0014] Using the speech recognition model, the acoustic features of the target speech and the visual features of the target video are obtained;

[0015] The acoustic features of the target speech are fused with the visual features of the target video using the speech recognition model.

[0016] Using the aforementioned speech recognition model and based on the fused features, the speech recognition result of the target speech is predicted.

[0017] Optionally, the first training video is annotated with video content description text;

[0018] The training objective of the speech recognition model also includes: making the video content description text predicted based on the visual features of the first training video more consistent with the video content description text annotated in the first training video.

[0019] Optionally, the training process of the speech recognition model includes:

[0020] The acoustic features of the first training speech and the visual features of the first training video are obtained by using a speech recognition model. Based on the acoustic features of the first training speech and supplemented by the visual features of the first training video, the speech recognition result of the first training speech is predicted to obtain the first prediction result.

[0021] Based on the visual features of the first training video, predict the video content description text of the first training video to obtain a second prediction result;

[0022] Based on the first prediction result and the speech content text of the first training speech annotation, as well as the second prediction result and the video content description text of the first training video annotation, the parameters of the speech recognition model are updated.

[0023] Optionally, predicting the speech recognition result of the first training speech based on the acoustic features of the first training speech, supplemented by the visual features of the first training video, includes:

[0024] Modality discarding processing is performed on the acoustic features of the first training speech and the visual features of the first training video to obtain the features after modality discarding processing;

[0025] The features after the modality discarding process are fused;

[0026] Predict the speech recognition result of the first training speech based on the fused features.

[0027] Optionally, the step of predicting the video content description text of the first training video based on its visual features to obtain a second prediction result includes:

[0028] Decode the visual features of the first training video: at each decoding time, obtain the visual context features required for decoding at that decoding time based on the visual features of the first training video, decode the visual context features required for decoding at that decoding time, and obtain the video content description text prediction result at that decoding time.

[0029] The second prediction result is composed of the video content description text prediction results at each decoding time.

[0030] Optionally, the speech recognition model includes: an encoding module for extracting acoustic features from input speech, extracting visual features from input video, and fusing the extracted acoustic features with the extracted visual features; and a decoding module for decoding the fused features output by the encoding module.

[0031] The encoding module in the initial speech recognition model is pre-trained using the second training data in the second training set. The second training data includes unlabeled second training speech and unlabeled second training video. The second training speech is the narration of the video content of the second training video.

[0032] Optionally, the process of training the encoding module using the second training data in the second training set includes:

[0033] For each piece of second training data in the second training set, the acoustic features of the second training speech in the second training data are obtained based on the pre-trained general speech recognition model, and used as the acoustic features corresponding to the second training data.

[0034] Cluster the acoustic features corresponding to each second training data in the second training set to obtain multiple acoustic features, and assign a category label to each acoustic feature. The category label of each acoustic feature is then determined as the category label of the corresponding second training data.

[0035] Each piece of second training data with a category label is used as third training data, and the obtained third training data forms a third training set.

[0036] The third training set is used to train the encoding module in conjunction with the data classification task.

[0037] Optionally, the step of using the third training set, combined with the data classification task, to train the encoding module includes:

[0038] Construct a data classification model that includes an encoding module and a classification module;

[0039] Obtain third training data from the third training set;

[0040] The acquired third training data is input into the encoding module of the data classification model to obtain the fusion features corresponding to the third training data;

[0041] The fusion features corresponding to the third training data are input into the classification module of the data classification model to predict the category, thus obtaining the category prediction result of the third training data.

[0042] Based on the category prediction results and category labels of the third training data, the parameters of the data classification model are updated.

[0043] A speech recognition device includes: a data acquisition module, a feature acquisition module, and a speech recognition result determination module;

[0044] The data acquisition module is used to acquire target speech and target video, wherein the target speech is the narration speech of the video content of the target video;

[0045] The feature acquisition module is used to extract acoustic features from the target speech to obtain the acoustic features of the target speech, and to extract visual features containing video content information from the target video to obtain the visual features of the target video.

[0046] The speech recognition result determination module is used to determine the speech recognition result of the target speech based on the acoustic features of the target speech and the visual features of the target video.

[0047] A voice recognition device includes: a memory and a processor;

[0048] The memory is used to store programs;

[0049] The processor is configured to execute the program to implement each step of the speech recognition method described above.

[0050] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the speech recognition method described in any of the preceding claims.

[0051] The speech recognition method, apparatus, device, and storage medium provided by this invention first acquire a target video and target speech (the target speech is the narration of the video content of the target video). Then, acoustic features are extracted from the target speech, and visual features containing video content information are extracted from the target video. Finally, based on the acoustic features of the target speech and supplemented by the visual features of the target video, the speech recognition result of the target speech is determined. Considering that the target speech is the narration of the video content of the target video, and that it has a certain correlation with the video content, this invention, based on this characteristic of the target speech being the narration, proposes to extract visual features containing video content information from the target video, and then use these visual features to perform speech recognition on the narration. By supplementing the speech recognition of the target speech (i.e., the narration) with visual features containing video content information, a more accurate speech recognition result can be obtained. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0053] Figure 1 This is a schematic diagram of the hardware architecture involved in the present invention;

[0054] Figure 2 A schematic flowchart of the speech recognition method provided in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the process for training a speech recognition model according to an embodiment of the present invention;

[0056] Figure 4 This is a schematic diagram of the structure of the speech recognition model provided in an embodiment of the present invention;

[0057] Figure 5 This is a schematic diagram illustrating how acoustic and visual features are processed sequentially through a modality discarding module and a feature fusion module, as provided in an embodiment of the present invention.

[0058] Figure 6This is a schematic diagram illustrating the decoding of visual features via an attention module and a visual description decoding module, as provided in an embodiment of the present invention.

[0059] Figure 7 A flowchart illustrating the process of training an encoding module using second training data from a second training set, provided in an embodiment of the present invention;

[0060] Figure 8 This is a schematic diagram of the structure of the speech recognition device provided in an embodiment of the present invention;

[0061] Figure 9 This is a schematic diagram of the structure of a speech recognition device provided in an embodiment of the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] Most current speech recognition solutions are based on general recognition models. This involves pre-training a general speech recognition model and then using the trained model to recognize the input speech. While these general recognition model-based solutions perform well for general speech, their accuracy is lower when applied to recognizing narration in specific domains.

[0064] Given that recognition schemes based on general recognition models have low accuracy in recognizing narration speech in specific fields (such as game commentary), the inventors of this case conducted research. The initial idea was to first train a general speech recognition model, and then fine-tune the general speech recognition model using training speech and labeled text of the training speech in a specific field through transfer learning. The fine-tuned model then serves as the speech recognition model for the specific field. Subsequently, the narration speech in the specific field is input into the speech recognition model for speech recognition.

[0065] The inventors of this case studied the above-mentioned solution and found that although the solution improved the recognition accuracy of narration speech in a specific field to a certain extent, it still had a large gap compared with the recognition accuracy of speech in a general field.

[0066] The inventors of this case, through research on speech recognition tasks in specific domains, discovered that speech recognition tasks in specific domains are significantly more difficult than speech recognition tasks in general domains. Taking game commentary speech recognition tasks in the gaming field as an example: many factors affect the accuracy of game commentary speech recognition, such as the commentator's Mandarin proficiency, regional accent, game-specific terminology, and game scene noise. In addition, many game commentators may also be game players, meaning that different game commentators may have different levels of education, ages, and accents. In other words, the accuracy of game commentary speech recognition is also affected by complex factors such as the commentator's education level, age, and accent. All of these factors significantly reduce the accuracy of game commentary speech recognition.

[0067] In order to significantly improve the recognition accuracy of narration speech in specific fields, the inventors of this case continued their research. During the research process, they discovered that the current speech recognition scheme is a single-modal speech recognition scheme. Based on this discovery, the inventors of this case conceived of using a multimodal speech recognition scheme, that is, using the speaker's lip reading content to assist in the recognition of the speaker's speech. Specifically, the speech recognition model is trained using the speaker's training speech and the corresponding video (the video corresponding to the training speech is a video of the speaker of the training speech, which includes the face of the speaker of the training speech). Then, the speech to be recognized and the corresponding video (the video corresponding to the speech to be recognized includes the face of the speaker of the speech to be recognized) are input into the speech recognition model.

[0068] This study investigated the aforementioned multimodal speech recognition scheme and found that it has several shortcomings: First, it does not fully consider the characteristics of the narration; second, the scheme requires capturing the speaker's (narrator's) facial information via a camera, which, due to user privacy concerns, renders the multimodal speech recognition model unsuitable for many situations.

[0069] Given the numerous shortcomings of the aforementioned multimodal speech recognition schemes, the inventors of this case continued their research. Through this research, the inventors discovered that narration is the voice of a commentator describing and analyzing the content of a video screen. This means that the content of the narration is related to the content of the video screen being described. Taking game commentary as an example, since game commentary is the voice of a game commentator describing and analyzing the content of a game video screen, the content of the game video screen and the content of the narration are related. For example, if the content of the game video screen is "The robot is coming from the middle lane..." The commentary states, "Huo Wu is walking along the river, and at this moment, Huo Wu gets the first blood." Based on this discovery, the inventor continued research and eventually proposed a more effective speech recognition method. The basic concept of this speech recognition method is to use the content of the video footage narrated by the commentator to assist in the recognition of the commentary's speech. This speech recognition method fully considers the characteristics of the commentary's speech and can significantly improve the recognition accuracy of commentary speech in specific fields. In addition, since this speech recognition method does not require capturing the commentator's facial information through a camera, it does not infringe on the commentator's privacy.

[0070] Before introducing the speech recognition method provided by this invention, the hardware architecture involved in this invention will be described first.

[0071] In one possible implementation, such as Figure 1 As shown, the hardware architecture involved in this invention may include: electronic device 101 and server 102.

[0072] For example, electronic device 101 can be any electronic product that can interact with a user, such as PC, laptop, tablet, mobile phone, learning machine, smart TV, etc.

[0073] It should be noted that, Figure 1 This is just one example; there can be many types of electronic devices, not limited to... Figure 1 The laptop in the middle.

[0074] For example, server 102 can be a single server, a server cluster consisting of multiple servers, or a cloud computing server center. Server 102 may include processors, memory, and network interfaces, etc.

[0075] For example, electronic device 101 can establish a connection and communicate with server 102 through a wireless communication network; for example, electronic device 101 can establish a connection and communicate with server 102 through a wired communication network.

[0076] Electronic device 101 can acquire target speech and target video (target speech is the narration of the video content of the target video), and send the target speech and target video to server 102. Server 102 performs speech recognition on the target speech according to the speech recognition method provided by the present invention.

[0077] In another possible implementation, the hardware architecture involved in this invention may include: an electronic device.

[0078] The electronic device is an electronic product with strong data processing capabilities. The electronic device can acquire target speech and target video (the target speech is the narration of the video content of the target video), and perform speech recognition on the target speech according to the speech recognition method provided by this invention.

[0079] Those skilled in the art should understand that the above-described electronic devices and servers are merely examples, and other existing or future electronic devices or servers that are applicable to this invention should also be included within the scope of protection of this invention, and are hereby incorporated by reference.

[0080] The speech recognition method provided by the present invention will be described in the following embodiments.

[0081] Please see Figure 2 The diagram illustrates a flowchart of a speech recognition method provided in an embodiment of the present invention. This speech recognition method may include:

[0082] Step S201: Obtain the target audio and target video.

[0083] Among them, target speech and target video are speech and video in a specific field. Target speech is the speech to be recognized, which is the narration of the video content of the target video, that is, the narrator's narration of the video content of the target video.

[0084] For example, the target speech and target video are speech and video in the field of games. Specifically, the target video is a game video, and the target speech is the narration of the video content of the game video, that is, the speech of the game commentator narrating the video content of the game video.

[0085] Besides being applicable to the gaming industry, target audio and target video can also be used in other fields, such as sports. Specifically, target video can be sports event videos, and target audio can be the narration of the video content of sports event videos, that is, the audio of sports commentators explaining the content of sports event videos.

[0086] Step S202: Extract acoustic features from the target speech to obtain the acoustic features of the target speech, and extract visual features containing video content information from the target video to obtain the visual features of the target video.

[0087] This embodiment extracts features from both the target speech and target video modalities to obtain the acoustic features of the target speech and the visual features of the target video. The visual features of the target video include video content information.

[0088] Step S203: Determine the speech recognition result of the target speech based on the acoustic features of the target speech and the visual features of the target video.

[0089] Specifically, the acoustic features of the target speech are fused with the visual features of the target video, and the speech recognition result of the target speech is determined based on the fused features.

[0090] Since the target speech is the narration of the target video content, the speech content of the target speech is related to the video content of the target video. In view of this, the present invention extracts visual features containing video content information from the target video and uses the visual features of the target video to identify the target speech.

[0091] The speech recognition method provided in this invention first acquires the target speech and the target video, then extracts acoustic features from the target speech and visual features containing video content information from the target video, and finally determines the speech recognition result of the target speech based on the acoustic features of the target speech and the visual features of the target video. Considering that the target speech is the narration of the target video content, and that it has a certain correlation with the video content, this invention, based on this characteristic of the target speech being the narration, proposes to extract visual features containing video content information from the target video and use these visual features to recognize the narration. When recognizing the narration, supplementing it with visual features containing video content information can obtain more accurate recognition results, especially for some proper nouns and rare words in a specific domain (the domain to which the target speech belongs). The speech recognition method provided in this invention improves the user experience of speech recognition services in narration scenarios.

[0092] In one possible implementation, the speech recognition method provided in the above embodiments can be implemented based on a pre-trained speech recognition model. It should be noted that the model-based implementation of the speech recognition method provided in the above embodiments is merely an example, and the present invention does not limit the specific implementation form.

[0093] Specifically, the process of implementing speech recognition based on a speech recognition model can include:

[0094] Step a1: Obtain the target audio and target video.

[0095] The target speech is the speech to be recognized, which is the narration of the video content of the target video. The target video is used to assist in the recognition of the target speech.

[0096] Step a2: Use a speech recognition model to obtain the acoustic features of the target speech and the visual features of the target video.

[0097] The target speech and target video are input into the speech recognition model. The speech recognition model extracts acoustic features from the input target speech to obtain the acoustic features of the target speech, and extracts visual features containing video content information from the input target video to obtain the visual features of the target video.

[0098] Step a3: Using a speech recognition model, fuse the acoustic features of the target speech with the visual features of the target video.

[0099] After obtaining the acoustic features of the target speech and the visual features of the target video, the speech recognition model fuses the acoustic features of the target speech and the visual features of the target video to obtain the fused features.

[0100] Step a4: Using a speech recognition model, based on the fused features, predict the speech recognition result of the target speech.

[0101] The speech recognition model predicts the speech recognition result of the target speech based on the fused features.

[0102] The aforementioned speech recognition model was trained using the first training data in the first training set.

[0103] In one possible implementation, the first training data includes a first training video and first training speech with annotated text. The first training speech is the narration of the video content of the first training video. The speech recognition model is trained with the training objective of making the speech recognition result predicted based on the acoustic features of the first training speech and supplemented by the visual features of the first training video consistent with the annotated text of the first training speech. It should be noted that the annotated text of the first training speech can be obtained by: firstly, inputting the first training speech into a pre-trained general speech recognition model for recognition to obtain the recognition result, and then manually correcting the recognition result. This method has high annotation efficiency. Of course, this embodiment is not limited to this, and the first pair of training speech can also be directly annotated manually.

[0104] In another possible implementation, to train a better-performing speech recognition model, the first training data includes a first training video labeled with video content description text and a first training speech labeled with speech content text. The first training speech is the narration of the video content of the first training video. The speech recognition model is trained to make the speech recognition result predicted based on the acoustic features of the first training speech and supplemented by the visual features of the first training video consistent with the speech content text labeled on the first training speech as the first training objective, and to make the video content description text predicted based on the visual features of the first training video consistent with the video content description text labeled on the first training video as the second training objective.

[0105] It should be noted that when training the speech recognition model, training it in conjunction with the second training objective mentioned above enables the speech recognition model to extract visual features containing more and richer video content information from the input video. In turn, by using visual features containing more and richer video content information for speech recognition, more accurate recognition results can be obtained.

[0106] In another embodiment of the present invention, taking the first training data in the first training set, which includes first training speech with text labeled with speech content and first training video with text labeled with video content description, as an example, the process of training a speech recognition model using the first training data in the first training set is described.

[0107] Please see Figure 3 The diagram illustrates the process of training a speech recognition model, which may include:

[0108] Step S301: Obtain the first training data from the first training set.

[0109] The first training set includes multiple sets of first training data, each set of first training data including first training audio with text labeled with audio content and first training video with text labeled with video content description.

[0110] Step S302a: Use a speech recognition model to obtain the acoustic features of the first training speech in the first training data.

[0111] Acoustic features are extracted from the first training speech in the first training data using a speech recognition model to obtain the acoustic features of the first training speech.

[0112] A speech recognition model may include an encoding module, for example, such as Figure 4 As shown, the encoding module can use the acoustic feature extraction module 401a to obtain the first training speech from the first training data and input the acoustic feature extraction module 401a of the speech recognition model. The acoustic feature extraction module 401a extracts acoustic features from the input first training speech and outputs them.

[0113] Optionally, the acoustic feature extraction 401 may include a speech encoder, which encodes the first training speech and outputs the acoustic features of the first training speech. The speech encoder may, but is not limited to, include downsampling convolution and a Conformer module.

[0114] Step S302b: Use a speech recognition model to obtain the visual features of the first training video in the first training data.

[0115] Visual features of the first training video in the first training data are extracted using a speech recognition model.

[0116] like Figure 4 As shown, the encoding module of the speech recognition model may include a visual feature extraction module 401b. The first training video in the acquired first training data is input into the visual feature extraction module 401b of the speech recognition model. The visual feature extraction module 401b extracts visual features from the input first training video and outputs them.

[0117] Optionally, the visual feature extraction module 401b may include a video encoder and a feature processing module. The video encoder encodes the input first training video and outputs visual features representing the visual information of the first training video. The visual features output by the video encoder are further input into the feature processing module for processing. The feature processing module outputs visual features that can better represent the visual information of the first training video. The video encoder may, but is not limited to, using 3DConvNet, and the feature processing module may, but is not limited to, using a bidirectional long short-term memory network (Bidirectional LSTM). It should be noted that, in another possible implementation, the visual feature extraction module 401b may also include only a video encoder.

[0118] Step S303a: Using a speech recognition model, based on the acoustic features of the first training speech and supplemented by the visual features of the first training video, predict the speech recognition result of the first training speech to obtain the first prediction result.

[0119] Specifically, the acoustic features of the first training speech are fused with the visual features of the first training video using a speech recognition model, and the speech recognition model is used to predict the speech recognition result of the first training speech based on the fused features, thus obtaining the first prediction result.

[0120] like Figure 4As shown, the encoding module of the speech recognition model may include a feature fusion module 402. In one possible implementation, the acoustic features of the first training speech and the visual features of the first training video can be directly input into the feature fusion module 402 for feature fusion; in another possible implementation, such as... Figure 5 As shown, the acoustic features of the first training speech and the visual features of the first training video can be input into the modality discarding module for modality discarding processing to obtain the features after modality discarding processing. Then, the features after modality discarding processing are input into the feature fusion module 402 for feature fusion.

[0121] It should be noted that after the acoustic features of the first training speech and the visual features of the first training video are input into the modality discarding module, the modality discarding module performs modality discarding processing on the acoustic features of the first training speech and the visual features of the first training video based on the set modality discarding probabilities. The set modality discarding probabilities include the probability of discarding features of each modality (the probability of discarding acoustic features, the probability of discarding visual features) and the probability of not performing modality discarding (not discarding any modality features). The sum of all probabilities is 1. The result of the modality discarding module's modality discarding processing may be the discarding of features of a certain modality, or it may be the not-discarding of any modality features.

[0122] After obtaining the features after modal discarding, the features after modal discarding are fused based on the feature fusion module 402. It should be noted that if the features after modal discarding include acoustic features and visual features, the acoustic features are fused with the visual features. If the features after discarding only include features of one modality, the features of that one modality are fused with 0. For example, if the features after modal discarding are acoustic features, the visual features are treated as 0 and the acoustic features are fused with 0.

[0123] Assume the acoustic features of the first training speech are F a The visual features of the first training video are F v The modal discarding module selects acoustic features F based on a set probability. a and visual features F v The result of modal discarding is one of the following three results: acoustic feature F a Discarded, visual feature F v Features that were discarded and those that were not. After obtaining the features after modality discarding, the feature fusion module 402 performs feature fusion on the features after modality discarding in the channel dimension to obtain the fused feature F. av , fusion feature F av It can be represented as:

[0124]

[0125] Where, f(F) a ,F v ) indicates that the acoustic feature F a With visual feature F v Integration, and the same applies to other aspects.

[0126] When fusing acoustic and visual features, the feature fusion module 402 determines the weights of each feature and then sums them according to these weights to obtain the fused features. It's important to note that the sum of the weights for each feature is 1. A higher weight for a visual feature indicates a greater correlation between the video and audio content, while a lower weight indicates a weaker correlation. This fusion method allows the model to prioritize acoustic features in determining the speech recognition result when the video and audio content are weakly or unrelated. Optionally, the feature fusion module 402 can employ a Conformer module.

[0127] It should also be noted that the modality discarding module mentioned above is only used during the training phase. The purpose of introducing the modality discarding module during the training phase is to enable the speech recognition model to have the following ability: when the input is only speech, it can also output relatively accurate speech recognition results.

[0128] Speech recognition models also include decoding modules, such as Figure 4 As shown, the decoding module of the speech recognition model may include a speech recognition decoder 403. The fused features output by the feature fusion module 402 are input to the speech recognition decoder 403 for decoding, and the resulting decoding result is the first prediction result.

[0129] Step S303b: Based on the visual features of the first training video, predict the video content description text of the first training video to obtain the second prediction result.

[0130] The visual features of the first training video can be input into the video content description text prediction module. The video content description text prediction module predicts the video content description text of the first training video based on the input visual features to obtain the second prediction result.

[0131] Specifically, the process of predicting the video content description text of the first training video based on the visual features of the first training video to obtain the second prediction result includes: decoding the visual features of the first training video: at each decoding time, obtaining the visual context features required for decoding at that decoding time based on the visual features of the first training video, decoding the visual context features required for decoding at that decoding time, and obtaining the video content description text prediction result at that decoding time; the second prediction result is composed of the video content description text prediction results at each decoding time.

[0132] More specifically, such as Figure 6 As shown, the video content description text prediction module includes an attention module and a visual description decoder. The visual features of the first training video are input into the attention module. At each decoding time, the attention module obtains the visual context features required for decoding at that decoding time based on the visual features of the first training video. The visual context features required for decoding at that decoding time are input into the visual description decoder. The visual description decoder decodes the visual context features required for decoding at that decoding time and outputs the video content description text prediction result for that decoding time.

[0133] It should be noted that the attention module and visual description decoder mentioned above are only used during the training phase. That is, the attention module and visual description decoder are auxiliary training modules for the speech recognition model. The purpose of introducing the attention module and visual description decoder during the training phase is to improve the feature representation capability of the model, that is, to enable the visual feature extraction module to extract visual features containing more and richer video content information.

[0134] Step S304: Update the parameters of the speech recognition model based on the first prediction result and the speech content text of the first training speech annotation, as well as the second prediction result and the video content description text of the first training video annotation.

[0135] Specifically, based on the first prediction result and the speech content text annotated by the first training speech, and the second prediction result and the video content description text annotated by the first training video, the parameters of the speech recognition model are updated, including:

[0136] Step S3041a: Determine the first prediction loss based on the first prediction result and the speech content text of the first training speech annotation.

[0137] Optionally, the first prediction loss can be the cross-entropy loss. The calculation method of the cross-entropy loss is a prior art, and will not be described in detail in this embodiment.

[0138] Step S3041b: Determine the second prediction loss based on the second prediction result and the video content description text labeled in the first training video.

[0139] Optionally, the second prediction loss can be the cross-entropy loss. The calculation method of the cross-entropy loss is a prior art, and will not be described in detail in this embodiment.

[0140] Step S3042: Update the parameters of the speech recognition model based on the first prediction loss and the second prediction loss.

[0141] Specifically, the process of updating the parameters of the speech recognition model based on the first prediction loss and the second prediction loss may include: fusing the first prediction loss and the second prediction loss, and updating the parameters of the speech recognition model based on the fused loss.

[0142] There are several ways to fuse the first prediction loss and the second prediction loss. In one possible implementation, the first prediction loss and the second prediction loss can be directly summed. In another possible implementation, the first prediction loss and the second prediction loss can be weighted and summed. The fusion method of weighted summation is shown in the following formula:

[0143] Loss total =λLoss asr +θLoss vc (2)

[0144] Among them, Loss asr This represents the first prediction loss, i.e., the prediction loss on the speech recognition task. vc The second prediction loss represents the prediction loss for the visual description task. λ is the weight corresponding to the first prediction loss, and θ is the weight corresponding to the second prediction loss. The specific values ​​of λ and θ can be set according to the specific scenario. Loss total This refers to the loss after fusion.

[0145] Repeat steps S301 to S304 until the training termination conditions are met, such as model convergence or reaching the set number of training iterations.

[0146] To fully utilize video content information and improve the recognition effect of narration, this invention trains a speech recognition model using data from both the narrator's voice and the video content they narrate. Therefore, the speech recognition model in this invention is a multimodal speech recognition model. Compared to multimodal speech recognition models that supplement lip reading for speech recognition, this invention's model, by using the video content narrated by the narrator, achieves better recognition results for the narration. Furthermore, since this invention's multimodal speech recognition model does not require capturing the speaker's (narrator's) facial information through a camera, it does not infringe on the narrator's privacy. Thus, the multimodal speech recognition model in this invention has a wider range of applications.

[0147] In one possible implementation, a large amount of initial training data (labeled training data) can be used to train the speech recognition model to achieve better performance. Considering the difficulty of obtaining a large amount of labeled training data, another possible implementation is to pre-train a large amount of unlabeled training data (e.g., 100 hours of unlabeled training data) to train encoding modules (e.g., the acoustic feature extraction module, visual feature extraction module, and feature fusion module mentioned above) for extracting acoustic features from input speech, extracting visual features from input video, and fusing the extracted acoustic and visual features. Then, a model including the encoding module (trained encoding module) and the decoding module (e.g., the speech recognition decoder mentioned above) is constructed as the initial speech recognition model. Finally, a small amount of labeled training data (e.g., 10 hours of initial training data) is used to fine-tune the initial speech recognition model to obtain a speech recognition model with better performance.

[0148] For the second implementation method mentioned above, in addition to obtaining the first training set, a second training set is also required. The second training set includes multiple second training data, and each second training data includes unlabeled second training audio and unlabeled second training video. The second training audio is the narration audio of the video content of the second training video in the second training data.

[0149] The following section describes the process of training the encoding modules (such as the acoustic feature extraction module, visual feature extraction module, and feature fusion module mentioned above) using the second training data in the second training set.

[0150] Please see Figure 7 The diagram illustrates a flowchart of training an encoding module using second training data from a second training set, which may include:

[0151] Step S701: For each piece of second training data in the second training set, obtain the acoustic features of the second training speech in the second training data based on the pre-trained general speech recognition model, and use it as the acoustic features corresponding to the second training data.

[0152] The acoustic features corresponding to each second training data in the second training set can be obtained through step S701.

[0153] Step S702: Cluster the acoustic features corresponding to each second training data in the second training set to obtain multiple acoustic features, and assign a category label to each acoustic feature. The category label of each acoustic feature is determined as the category label of the corresponding second training data.

[0154] Existing clustering methods (such as K-Means clustering) can be used to cluster the acoustic features corresponding to each piece of second training data in the second training set. Through clustering, multiple types of acoustic features can be obtained. After obtaining multiple types of acoustic features, a category label can be assigned to each type of acoustic feature. For example, if four types of acoustic features are obtained through clustering, category labels "1", "2", "3", and "4" can be assigned to the four types of acoustic features respectively. After setting the category label for each type of acoustic feature, the category label of each acoustic feature can be used as the category label of the corresponding second training data. For example, if the category label of the acoustic feature corresponding to a piece of second training data is "1", then the category label "1" is used as the category label corresponding to that piece of second training data.

[0155] The category label of each piece of second training data in the second training set can be obtained through step S702.

[0156] Step S703: Use each piece of second training data with a category label as third training data, and form a third training set from the obtained third training data.

[0157] Training data with category labels can be obtained through steps S701 to S703.

[0158] Step S704: Using the third training set and combining it with the data classification task, train the encoding module.

[0159] Specifically, using the third training set and combining it with the data classification task, the process of training the encoding module can include:

[0160] Step S7041: Construct a data classification model that includes an encoding module and a classification module.

[0161] Step S7042: Obtain the third training data from the third training set.

[0162] Step S7043: Obtain the fusion features corresponding to the third training data by encoding the coding module of the training speech input data classification model in the acquired third training data.

[0163] Specifically, after the acquired third training data is input into the encoding module of the data classification model, the encoding module extracts acoustic features from the training speech in the third training data, extracts visual features from the training video in the third training data, and fuses the extracted acoustic features with the extracted visual features to output the fused features corresponding to the third training data.

[0164] Step S7044: Input the fusion features corresponding to the third training data into the classification module of the data classification model to obtain the category prediction result of the third training data.

[0165] The classification module of the data classification model predicts the category of the third training data based on the input fusion features and outputs the category prediction result of the third training data.

[0166] Step S7045: Update the parameters of the data classification model based on the category prediction results and category labels of the third training data.

[0167] Specifically, based on the category prediction results and category labels of the third training data, the category prediction loss is determined, and the parameters of the data classification model are updated according to the category prediction loss.

[0168] Optionally, the category prediction loss can be the cross-entropy loss. The calculation method of the cross-entropy loss is a prior art, and will not be described in detail in this embodiment.

[0169] Repeat steps S7042 to S7045 until the training termination condition is met (such as model convergence, or reaching the set number of training iterations).

[0170] After training, the encoded modules (such as visual feature extraction module, acoustic feature extraction module and feature fusion module) in the trained data classification model can be used, along with the decoding module (such as speech recognition decoder) to build an initial speech recognition model. After building the initial speech recognition model, the initial speech recognition model is trained using the first training data in the first training dataset. During training, a modality discarding module and a visual description decoder can be used as supplementary tools.

[0171] After obtaining the trained speech recognition model, speech recognition can be performed using it. Specifically, the target speech and target video are first acquired. Then, the target speech is input into the encoding module of the speech recognition model. The encoding module extracts acoustic features from the target speech and visual features from the target video. The extracted acoustic features and visual features are then fused together to output the fused features corresponding to the target speech. Finally, the fused features corresponding to the target speech are input into the decoding module of the speech recognition model for decoding to obtain the speech recognition result of the target speech.

[0172] This invention also provides a speech recognition device. The speech recognition device provided in this invention will be described below. The speech recognition device described below can be referred to in correspondence with the speech recognition method described above.

[0173] Please see Figure 8 The diagram shows a schematic of the structure of a speech recognition device provided in an embodiment of the present invention, which may include: a data acquisition module 801, a feature acquisition module 802, and a speech recognition result determination module 803.

[0174] The data acquisition module 801 is used to acquire target speech and target video, wherein the target speech is the narration of the video content of the target video.

[0175] The feature acquisition module 802 is used to extract acoustic features from the target speech to obtain the acoustic features of the target speech, and to extract visual features containing video content information from the target video to obtain the visual features of the target video.

[0176] The speech recognition result determination module 803 is used to determine the speech recognition result of the target speech based on the acoustic features of the target speech and the visual features of the target video.

[0177] Optionally, the feature acquisition module 802 and the speech recognition result determination module 803 can be implemented through a speech recognition model. Specifically, the target speech and the target video are processed using a pre-trained speech recognition model to obtain the speech recognition result of the target speech.

[0178] The speech recognition model is trained using first training data in a first training set. The first training data includes a first training video and a first training speech with text labeled with speech content. The first training speech is the narration of the video content of the first training video. The training objective of the speech recognition model is to make the speech recognition result predicted based on the acoustic features of the first training speech and supplemented by the visual features of the first training video consistent with the text labeled with the speech content of the first training speech.

[0179] Optionally, the target speech and the target video are processed using a pre-trained speech recognition model to obtain the speech recognition result of the target speech, including:

[0180] Using the speech recognition model, the acoustic features of the target speech and the visual features of the target video are obtained;

[0181] The acoustic features of the target speech are fused with the visual features of the target video using the speech recognition model.

[0182] Using the aforementioned speech recognition model and based on the fused features, the speech recognition result of the target speech is predicted.

[0183] Optionally, the first training video is annotated with video content description text; the training objective of the speech recognition model further includes: making the video content description text predicted based on the visual features of the first training video consistent with the video content description text annotated in the first training video.

[0184] Optionally, the speech recognition device provided in this embodiment of the invention may further include: a first training module for training a speech recognition model. The first training module, when training the speech recognition model, is specifically used for:

[0185] The acoustic features of the first training speech and the visual features of the first training video are obtained by using a speech recognition model. Based on the acoustic features of the first training speech and supplemented by the visual features of the first training video, the speech recognition result of the first training speech is predicted to obtain the first prediction result.

[0186] Based on the visual features of the first training video, predict the video content description text of the first training video to obtain a second prediction result;

[0187] Based on the first prediction result and the speech content text of the first training speech annotation, as well as the second prediction result and the video content description text of the first training video annotation, the parameters of the speech recognition model are updated.

[0188] Optionally, when the first training module uses a speech recognition model to predict the speech recognition result of the first training speech based on the acoustic features of the first training speech and supplemented by the visual features of the first training video, it is specifically used for:

[0189] Modality discarding processing is performed on the acoustic features of the first training speech and the visual features of the first training video to obtain the features after modality discarding processing;

[0190] The features after modality discarding are fused using a speech recognition model;

[0191] Using a speech recognition model, based on the fused features, the speech recognition result of the first training speech is predicted.

[0192] Optionally, when the first training module predicts the video content description text of the first training video based on the visual features of the first training video to obtain the second prediction result, it is specifically used for:

[0193] Decode the visual features of the first training video: at each decoding time, obtain the visual context features required for decoding at that decoding time based on the visual features of the first training video, decode the visual context features required for decoding at that decoding time, and obtain the video content description text prediction result at that decoding time.

[0194] The second prediction result is composed of the video content description text prediction results at each decoding time.

[0195] Optionally, the speech recognition model includes: an encoding module for extracting acoustic features from input speech, extracting visual features from input video, and fusing the extracted acoustic features with the extracted visual features; and a decoding module for decoding the fused features output by the encoding module.

[0196] The encoding module in the initial speech recognition model is pre-trained using the second training data in the second training set. The second training data includes unlabeled second training speech and unlabeled second training video. The second training speech is the narration of the video content of the second training video.

[0197] Optionally, the speech recognition device provided in this embodiment of the invention may further include: a second training module for training the encoding module using second training data in the second training set.

[0198] When the second training module trains the encoding module using the second training data from the second training set, it is specifically used for:

[0199] For each piece of second training data in the second training set, the acoustic features of the second training speech in the second training data are obtained based on the pre-trained general speech recognition model, and used as the acoustic features corresponding to the second training data.

[0200] Cluster the acoustic features corresponding to each second training data in the second training set to obtain multiple acoustic features, and assign a category label to each acoustic feature. The category label of each acoustic feature is then determined as the category label of the corresponding second training data.

[0201] Each piece of second training data with a category label is used as third training data, and the obtained third training data forms a third training set.

[0202] The third training set is used to train the encoding module in conjunction with the data classification task.

[0203] Optionally, when the second training module uses the third training set and combines it with the data classification task to train the encoding module, it is specifically used for:

[0204] Construct a data classification model that includes an encoding module and a classification module;

[0205] Obtain third training data from the third training set;

[0206] The acquired third training data is input into the encoding module of the data classification model to obtain the fusion features corresponding to the third training data;

[0207] The fusion features corresponding to the third training data are input into the classification module of the data classification model to predict the category, thus obtaining the category prediction result of the third training data.

[0208] Based on the category prediction results and category labels of the third training data, the parameters of the data classification model are updated.

[0209] The speech recognition device provided in this embodiment of the invention first acquires target speech and target video, then extracts acoustic features from the target speech and visual features containing video content information from the target video, and finally determines the speech recognition result of the target speech based on the acoustic features of the target speech supplemented by the visual features of the target video. Considering that the target speech is the narration of the target video content, and that it has a certain correlation with the video content, the speech recognition device provided in this embodiment of the invention extracts visual features containing video content information from the target video and uses these visual features to recognize the narration. By supplementing the narration recognition with visual features containing video content information, a more accurate recognition result can be obtained.

[0210] This invention also provides a voice recognition device; please refer to [link to relevant documentation]. Figure 9 The diagram shows the structure of the voice recognition device, which may include: at least one processor 901, at least one communication interface 902, at least one memory 903 and at least one communication bus 904.

[0211] In this embodiment of the invention, the number of processor 901, communication interface 902, memory 903, and communication bus 904 is at least one, and processor 901, communication interface 902, and memory 903 communicate with each other through communication bus 904.

[0212] The processor 901 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0213] The memory 903 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0214] The memory stores a program, which the processor can call. The program is used for:

[0215] Acquire target audio and target video, wherein the target audio is the narration audio of the video content of the target video;

[0216] Acoustic features are extracted from the target speech to obtain the acoustic features of the target speech, and visual features containing video content information are extracted from the target video to obtain the visual features of the target video;

[0217] The speech recognition result of the target speech is determined based on the acoustic features of the target speech and the visual features of the target video.

[0218] Optionally, the refined and extended functions of the program can be found in the description above.

[0219] This invention also provides a computer-readable storage medium that stores a program suitable for execution by a processor, the program being used for:

[0220] Acquire target audio and target video, wherein the target audio is the narration audio of the video content of the target video;

[0221] Acoustic features are extracted from the target speech to obtain the acoustic features of the target speech, and visual features containing video content information are extracted from the target video to obtain the visual features of the target video;

[0222] The speech recognition result of the target speech is determined based on the acoustic features of the target speech and the visual features of the target video.

[0223] Optionally, the refined and extended functions of the program can be found in the description above.

[0224] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0225] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0226] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech recognition method, characterized in that, include: Acquire target audio and target video, wherein the target audio is the narration audio of the video content of the target video; Using a pre-trained speech recognition model, acoustic features are extracted from the target speech to obtain the acoustic features of the target speech, and visual features containing video content information are extracted from the target video to obtain the visual features of the target video. Using the speech recognition model, the speech recognition result of the target speech is determined based on the acoustic features of the target speech and the visual features of the target video. The speech recognition model is trained using the first training data in the first training set. The first training data includes a first training video labeled with video content description text and a first training speech labeled with speech content text. The training objectives of the speech recognition model include: making the speech recognition result predicted based on the acoustic features of the first training speech and supplemented by the visual features of the first training video consistent with the speech content text annotated by the first training speech; and making the video content description text predicted based on the visual features of the first training video consistent with the video content description text annotated by the first training video.

2. The speech recognition method according to claim 1, characterized in that, The step of using the speech recognition model to determine the speech recognition result of the target speech based on the acoustic features of the target speech and supplemented by the visual features of the target video includes: The acoustic features of the target speech are fused with the visual features of the target video using the speech recognition model. Using the aforementioned speech recognition model and based on the fused features, the speech recognition result of the target speech is predicted.

3. The speech recognition method according to claim 1, characterized in that, The training process of the speech recognition model includes: Using a speech recognition model, the acoustic features of the first training speech and the visual features of the first training video are obtained. Based on the acoustic features of the first training speech and supplemented by the visual features of the first training video, the speech recognition result of the first training speech is predicted to obtain the first prediction result. Based on the visual features of the first training video, predict the video content description text of the first training video to obtain a second prediction result; Based on the first prediction result and the speech content text of the first training speech annotation, as well as the second prediction result and the video content description text of the first training video annotation, the parameters of the speech recognition model are updated.

4. The speech recognition method according to claim 3, characterized in that, The step of predicting the speech recognition result of the first training speech based on the acoustic features of the first training speech and supplemented by the visual features of the first training video includes: Modality discarding processing is performed on the acoustic features of the first training speech and the visual features of the first training video to obtain the features after modality discarding processing; The features after the modality discarding process are fused; Predict the speech recognition result of the first training speech based on the fused features.

5. The speech recognition method according to claim 3, characterized in that, The step of predicting the video content description text of the first training video based on its visual features to obtain a second prediction result includes: Decode the visual features of the first training video: at each decoding time, obtain the visual context features required for decoding at that decoding time based on the visual features of the first training video, decode the visual context features required for decoding at that decoding time, and obtain the video content description text prediction result at that decoding time. The second prediction result is composed of the video content description text prediction results at each decoding time.

6. The speech recognition method according to any one of claims 1 to 5, characterized in that, The speech recognition model includes: an encoding module for extracting acoustic features from input speech, extracting visual features from input video, and fusing the extracted acoustic features with the extracted visual features; and a decoding module for decoding the fused features output by the encoding module. The encoding module in the initial speech recognition model is pre-trained using the second training data in the second training set. The second training data includes unlabeled second training speech and unlabeled second training video. The second training speech is the narration of the video content of the second training video.

7. The speech recognition method according to claim 6, characterized in that, The process of training the encoding module using the second training data in the second training set includes: For each piece of second training data in the second training set, the acoustic features of the second training speech in the second training data are obtained based on the pre-trained general speech recognition model, and used as the acoustic features corresponding to the second training data. Cluster the acoustic features corresponding to each second training data in the second training set to obtain multiple acoustic features, and assign a category label to each acoustic feature. The category label of each acoustic feature is then determined as the category label of the corresponding second training data. Each piece of second training data with a category label is used as third training data, and the obtained third training data forms a third training set. The third training set is used to train the encoding module in conjunction with the data classification task.

8. The speech recognition method according to claim 7, characterized in that, The step of using the third training set, combined with the data classification task, to train the encoding module includes: Construct a data classification model that includes an encoding module and a classification module; Obtain third training data from the third training set; The acquired third training data is input into the encoding module of the data classification model to obtain the fusion features corresponding to the third training data; The fusion features corresponding to the third training data are input into the classification module of the data classification model to predict the category, thus obtaining the category prediction result of the third training data. Based on the category prediction results and category labels of the third training data, the parameters of the data classification model are updated.

9. A voice recognition device, characterized in that, include: The module consists of a data acquisition module, a feature acquisition module, and a speech recognition result determination module. The data acquisition module is used to acquire target speech and target video, wherein the target speech is the narration speech of the video content of the target video; The feature acquisition module is used to extract acoustic features from the target speech using a pre-trained speech recognition model to obtain the acoustic features of the target speech, and to extract visual features containing video content information from the target video to obtain the visual features of the target video. The speech recognition result determination module is used to determine the speech recognition result of the target speech by using the speech recognition model, based on the acoustic features of the target speech and supplemented by the visual features of the target video. The speech recognition model is trained using the first training data in the first training set. The first training data includes a first training video labeled with video content description text and a first training speech labeled with speech content text. The training objectives of the speech recognition model include: making the speech recognition result predicted based on the acoustic features of the first training speech and supplemented by the visual features of the first training video consistent with the speech content text annotated by the first training speech; and making the video content description text predicted based on the visual features of the first training video consistent with the video content description text annotated by the first training video.

10. A voice recognition device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the speech recognition method as described in any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the speech recognition method as described in any one of claims 1 to 8.