Method and device for generating speech recognition training set

By combining audio and video text information to generate speech recognition training sets, the problem of low efficiency in training set construction in existing technologies is solved, more efficient and flexible training set generation is achieved, and the generalization performance of the speech recognition model is improved.

CN115312032BActive Publication Date: 2025-09-12JINGDONG TECH HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110514350.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-08
Publication Date
2025-09-12
Estimated Expiration
2041-05-08

AI Technical Summary

Technical Problem

In existing technologies, improving the generalization performance of speech recognition models relies on a large amount of manually labeled speech data, and constructing training sets is inefficient and inflexible.

Method used

By obtaining the audio and video to be processed, the automatic speech recognition model and OCR technology are used to extract the text information of the audio and video respectively, and a speech recognition training set is generated based on the consistency of the audio text and the video text, including silence detection, audio segmentation, video frame text splicing and edit distance optimization.

Benefits of technology

The flexibility and efficiency of constructing speech recognition training sets have been improved, and the generated training sets are of higher quality, which can better optimize the speech recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115312032B_ABST
    Figure CN115312032B_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for generating a speech recognition training set. A specific implementation of the method includes: obtaining audio and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; identifying the audio to be processed to obtain audio text; identifying text information in the video to be processed to obtain video text; and, based on the consistency between the audio text and the video text, using the audio to be processed as a speech sample and the video text as a label to obtain a speech recognition training set. This application provides a method for automatically obtaining a speech recognition training set, improving the flexibility and efficiency of constructing a speech recognition training set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and specifically to a method and device for generating a speech recognition training set. Background Art

[0002] In recent years, with the rapid development of deep learning technology, the use of automatic speech recognition (ASR) models based on deep neural networks has become a mainstream trend in speech recognition technology. To improve the generalization performance of speech recognition models, it is necessary to collect a large amount of speech data and optimize the speech recognition models through training sets constructed through manual annotation. Summary of the Invention

[0003] The embodiments of the present application provide a method and apparatus for generating a speech recognition training set.

[0004] In the first aspect, an embodiment of the present application provides a method for generating a speech recognition training set, comprising: obtaining audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; identifying the audio to be processed to obtain audio text; identifying the text information in the video to be processed to obtain video text; based on the consistency between the audio text and the video text, using the audio to be processed as a speech sample and the video text as a label to obtain a speech recognition training set.

[0005] In some embodiments, the above-mentioned identification of the audio to be processed and obtaining the audio text includes: deleting the silent part in the audio to be processed according to the silence detection algorithm to obtain multiple non-silent audio segments; identifying multiple audio segments to obtain multiple audio segment texts included in the audio text.

[0006] In some embodiments, the above-mentioned identification of text information in the video to be processed to obtain video text includes: determining multiple video frame sequences corresponding one-to-one to multiple audio clips from the video to be processed; identifying text information in each video frame in the multiple video frame sequences, and obtaining video frame text included in the video text.

[0007] In some embodiments, the above-mentioned consistency between audio text and video text is based on the audio to be processed as a speech sample and the video text as a label to obtain a speech recognition training set, including: for each video frame sequence in a plurality of video frame sequences, performing the following operations: taking one video frame text of at least one video frame text recognized in the video frame in the video frame sequence as a unit, splicing the text information included in each video frame in the video frame sequence to obtain a plurality of video frame sequence texts corresponding to the video frame sequence; determining the target video frame sequence text according to the edit distance between each video frame sequence text in the plurality of video frame sequence texts and the target audio segment text, wherein the target audio segment text is the audio segment text corresponding to the audio segment corresponding to the video frame sequence; taking each audio segment in the plurality of audio segments as a speech sample and the target video frame sequence text corresponding to the audio segment as a label to obtain a speech recognition training set.

[0008] In some embodiments, the above-mentioned method takes one video frame text of at least one video frame text identified in the video frame in the video frame sequence as a unit, splicing the text information included in each video frame in the video frame sequence to obtain multiple video frame sequence texts corresponding to the video frame sequence, including: for each video frame including text information in the video frame sequence, performing the following operations: determining multiple texts to be spliced ​​corresponding to the video frame, and splicing the multiple texts to be spliced ​​with at least one video frame text in the video frame to obtain multiple spliced ​​texts; based on the editing distance between the multiple spliced ​​texts and the target audio segment text, selecting a preset number of spliced ​​texts from the multiple spliced ​​texts as the multiple texts to be spliced ​​corresponding to the next video frame of the video frame.

[0009] In some embodiments, the above method also includes: for each video frame sequence in multiple video frame sequences, in response to determining that the edit distance between the target video frame sequence text corresponding to the video frame sequence and the target audio segment text is greater than a preset distance threshold, deleting the training sample corresponding to the video frame sequence in the speech recognition training set.

[0010] In the second aspect, an embodiment of the present application provides a device for generating a speech recognition training set, comprising: an acquisition unit, configured to acquire audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; a first recognition unit, configured to recognize the audio to be processed and obtain audio text; a second recognition unit, configured to recognize text information in the video to be processed and obtain video text; an acquisition unit, configured to obtain a speech recognition training set based on the consistency between the audio text and the video text, with the audio to be processed as a speech sample and the video text as a label.

[0011] In some embodiments, the first recognition unit is further configured to: delete the silent part in the audio to be processed according to the silence detection algorithm to obtain multiple non-silent audio segments; recognize multiple audio segments to obtain multiple audio segment texts included in the audio text.

[0012] In some embodiments, the second recognition unit is further configured to: determine multiple video frame sequences corresponding one-to-one to multiple audio clips from the video to be processed; recognize text information in each video frame in the multiple video frame sequences, and obtain the video frame text included in the video text.

[0013] In some embodiments, the obtaining unit is further configured to: for each video frame sequence in a plurality of video frame sequences, perform the following operations: taking one video frame text among at least one video frame text identified in the video frames in the video frame sequence as a unit, splicing the text information included in each video frame in the video frame sequence to obtain a plurality of video frame sequence texts corresponding to the video frame sequence; determining the target video frame sequence text according to the editing distance between each video frame sequence text in the plurality of video frame sequence texts and the target audio segment text, wherein the target audio segment text is the audio segment text corresponding to the audio segment corresponding to the video frame sequence; taking each audio segment in the plurality of audio segments as a speech sample and the target video frame sequence text corresponding to the audio segment as a label to obtain a speech recognition training set.

[0014] In some embodiments, the obtaining unit is further configured to: for each video frame including text information in the video frame sequence, perform the following operations: determine multiple texts to be spliced ​​corresponding to the video frame, and splice the multiple texts to be spliced ​​with at least one video frame text in the video frame to obtain multiple spliced ​​texts; based on the editing distance between the multiple spliced ​​texts and the target audio segment text, select a preset number of spliced ​​texts from the multiple spliced ​​texts as the multiple texts to be spliced ​​corresponding to the next video frame of the video frame.

[0015] In some embodiments, the above-mentioned device also includes: a deletion unit, which is configured to delete the training samples corresponding to the video frame sequence in the speech recognition training set for each video frame sequence in a plurality of video frame sequences in response to determining that the editing distance between the target video frame sequence text corresponding to the video frame sequence and the target audio segment text is greater than a preset distance threshold.

[0016] In a third aspect, an embodiment of the present application provides a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present application provides an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0018] The method and device for generating a speech recognition training set provided in an embodiment of the present application obtain audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; recognize the audio to be processed to obtain audio text; recognize the text information in the video to be processed to obtain video text; based on the consistency between the audio text and the video text, use the audio to be processed as a speech sample and the video text as a label to obtain a speech recognition training set, thereby providing a method for automatically obtaining a speech recognition training set and improving the flexibility and efficiency of constructing a speech recognition training set. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0020] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present application may be applied;

[0021] Figure 2 is a flowchart of an embodiment of a method for generating a speech recognition training set according to the present application;

[0022] Figure 3 is a schematic diagram of the text splicing process according to this embodiment;

[0023] Figure 4 is a schematic diagram of an application scenario of the method for generating a speech recognition training set according to this embodiment;

[0024] Figure 5 is a flowchart of another embodiment of a method for generating a speech recognition training set according to the present application;

[0025] Figure 6 is a structural diagram of an embodiment of a device for generating a speech recognition training set according to the present application;

[0026] Figure 7 It is a structural diagram of a computer system suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION

[0027] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0028] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0029] Figure 1 An exemplary architecture 100 is shown to which the method and apparatus for generating a speech recognition training set of the present application can be applied.

[0030] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, and 103 constitute a topological network, and network 104 is used to provide a medium for communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0031] Terminal devices 101, 102, and 103 can be hardware devices or software that support network connection for data interaction and data processing. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connection, information acquisition, interaction, display, processing, and other functions, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, for example, to provide distributed services, or they can be implemented as a single software or software module. No specific limitations are given here.

[0032] Server 105 can be a server that provides various services, such as a backend processing server that obtains the corresponding video and audio to be processed sent by users through terminal devices 101, 102, and 103, processes the information, and automatically constructs a speech recognition training set. Furthermore, the server can also train an initial speech recognition model based on the speech recognition training set or optimize a pre-trained speech recognition model. For example, server 105 can be a cloud server.

[0033] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., software or software modules for providing distributed services), or as a single software or software module. No specific limitations are given here.

[0034] It should also be noted that the method for generating a speech recognition training set provided in the embodiments of the present application can be executed by a server, or by a terminal device, or by a server and a terminal device in cooperation with each other. Accordingly, the various parts (e.g., various units) included in the apparatus for generating a speech recognition training set can be all provided in the server, or all provided in the terminal device, or provided separately in the server and the terminal device.

[0035] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system is merely illustrative. Any number of terminal devices, networks, and servers may be provided as needed. When the electronic device on which the method for generating a speech recognition training set is running does not need to transmit data with other electronic devices, the system architecture may only include the electronic device (e.g., a server or terminal device) on which the method for generating a speech recognition training set is running.

[0036] Continue to refer Figure 2 , shows a process 200 of an embodiment of a method for generating a speech recognition training set, comprising the following steps:

[0037] Step 201: Obtain the audio and video to be processed.

[0038] In this embodiment, the execution subject of the method for generating a speech recognition training set (eg Figure 1 The server in the embodiment can obtain the audio and video to be processed remotely or locally through a wired network connection or a wireless network connection. The video to be processed includes text information corresponding to the audio to be processed.

[0039] As an example, the data including the corresponding audio and video to be processed may be various audio and video data such as movies, TV series, short videos, etc. The text information in the video to be processed is subtitle information, and the audio to be processed is voice information corresponding to the subtitle information.

[0040] In this embodiment, the speech data represented by the audio to be processed can be various types of speech, including but not limited to foreign language audio, Mandarin audio, and dialect audio. The audio to be processed and the video to be processed can be data of a long duration or data of a short duration.

[0041] Step 202: Identify the audio to be processed and obtain the audio text.

[0042] In this embodiment, the execution entity can identify the audio to be processed and obtain the audio text.

[0043] As an example, the execution subject may process the audio to be processed based on an automatic speech recognition model to obtain the audio text. The automatic speech recognition model is used to characterize the correspondence between the audio to be processed and the text.

[0044] In some optional implementations of this embodiment, the execution entity may perform step 202 as follows:

[0045] First, according to the silence detection algorithm, the silent parts in the audio to be processed are deleted to obtain multiple non-silent audio clips.

[0046] In this implementation, the execution subject may use the silent portion in the audio to be processed as a segmentation point, and segment the audio to be processed after deleting the silent portion to obtain multiple audio segments.

[0047] In the case where the obtained audio clips are long, the above-mentioned execution entity can set a duration threshold, further cut the audio clips whose duration is greater than the duration threshold into units of the duration represented by the duration threshold, and record the start and end time of each audio clip.

[0048] For example, to prevent background music and other factors from preventing the silence detection algorithm from completely truncating the audio, resulting in longer audio segments, a duration threshold T is set to forcibly split audio segments whose duration exceeds the duration threshold T into multiple segments of duration T. The duration threshold can be set based on actual conditions, for example, T = 10a.

[0049] Second, multiple audio segments are identified to obtain multiple audio segment texts included in the audio text.

[0050] In this implementation, the execution entity may input each of the multiple audio segments into an automatic speech recognition model to obtain multiple audio segment texts, wherein the multiple audio segments correspond to the multiple audio segment texts one by one, and the multiple audio segment texts constitute the audio text.

[0051] Step 203: Identify text information in the video to be processed and obtain video text.

[0052] In this embodiment, the execution subject can identify text information in the video to be processed and obtain the video text.

[0053] As an example, the execution entity may utilize OCR (Optical Character Recognition) technology to identify the text information contained in each video frame of the video to be processed, and then, according to the playback order of the video frames in the video to be processed, splice the text information corresponding to each video frame to obtain the video text. OCR technology is currently a relatively mature technology and will not be described in detail here.

[0054] In some optional implementations of this embodiment, the execution entity may perform step 203 as follows:

[0055] First, multiple video frame sequences corresponding to multiple audio clips are determined from the video to be processed.

[0056] In this implementation, for each audio segment of the multiple audio segments, the execution entity extracts multiple video frames corresponding to the audio segment from the video to be processed to obtain a video frame sequence.

[0057] As an example, the start and end times of the audio clip are t sk and t ek , the above execution entities can The start and end video frames of the video frame sequence corresponding to the audio clip are determined in sequence. and Represents rounding up and rounding down, f P Characterizes the frame rate of the video to be processed. The above execution entity can pre-set the sampling rate and based on the sampling rate, start from the starting frame termination Extract video frames between the two audio clips to obtain the video frame sequence corresponding to the audio clip.

[0058] Second, text information in each video frame in a plurality of video frame sequences is identified to obtain video frame text included in the video text.

[0059] In this implementation, the execution subject may utilize OCR technology to identify text information in each video frame in a plurality of video frame sequences, and obtain the video frame text included in the video text.

[0060] It is understood that for each video frame, the execution entity may not recognize text information, that is, the video frame does not contain text information; or it may recognize multiple text information, resulting in multiple video frame texts. For example, the multiple video frame texts include subtitle information in the video frame and text information in the video frame (such as store name information on a store sign, road name information on a road sign, advertising slogan information, etc.).

[0061] In other cases, the text information included in adjacent frames contains the same text information, for example, the subtitle information included in adjacent video frames is the same.

[0062] In this embodiment, a preset identifier can be added to the video frame to indicate that the video frame does not include text information, or includes the same text information as an adjacent video frame. The preset identifier can be any pre-set identifier, such as "Blank".

[0063] In step 204 , based on the consistency between the audio text and the video text, the audio to be processed is used as a speech sample and the video text is used as a label to obtain a speech recognition training set.

[0064] In this embodiment, the execution subject can obtain a speech recognition training set based on the consistency between the audio text and the video text, using the audio to be processed as a speech sample and the video text as a label.

[0065] As an example, when the audio text is consistent with the video text, the above execution subject uses the audio to be processed as a voice sample and the video text as a label to obtain a speech recognition training set.

[0066] In some optional implementations of this embodiment, the execution entity may perform step 204 as follows:

[0067] First, for each video frame sequence in the multiple video frame sequences, perform the following operations:

[0068] First, taking one video frame text among at least one video frame text identified in the video frames in the video frame sequence as a unit, the text information included in each video frame in the video frame sequence is spliced ​​to obtain multiple video frame sequence texts corresponding to the video frame sequence.

[0069] As an example, the video frame sequence includes 3 video frames, and the number of video frame texts corresponding to the 3 video frames is 3, 4, and 3 respectively. Then, there are 36 (3×4×3) multiple video frame sequence texts corresponding to the video frame sequence.

[0070] In some optional implementations, the video frame text set corresponding to each video frame includes, in addition to the video frame text identified from the video frame, a preset identifier indicating that the video frame includes the same text information as adjacent video frames. The preset identifier may also indicate a situation where the video frame does not include text information.

[0071] Continuing with the above example, after adding a preset identifier to the video frame text combination corresponding to each video frame, the number of video frame texts corresponding to the three video frames is 4, 5, and 4 respectively, and the number of multiple video frame sequence texts corresponding to the video frame sequence is 80 (4×5×4).

[0072] like Figure 3 As shown, the preset identifier is "Blank". The recognition results corresponding to video frames 301, 302, and 303 are 304, 305, and 306, respectively. Each video frame text in each video frame can be combined with the video frame text in other video frames to obtain multiple video frame sequence texts. For example, the recognition results of video frames 301, 302, and 303 can be combined into "How is the weather today?"

[0073] Second, a target video frame sequence text is determined according to an edit distance between each video frame sequence text in the plurality of video frame sequence texts and the target audio segment text.

[0074] The target audio segment text is the audio segment text corresponding to the audio segment of the video frame sequence. The edit distance refers to the minimum number of edit operations required to convert two strings from one to the other.

[0075] As an example, the execution entity may determine, among the multiple video frame sequence texts, a video frame sequence text having the smallest edit distance with the target audio segment text as the target video frame sequence text.

[0076] Then, each audio segment in the multiple audio segments is used as a speech sample, and the target video frame sequence text corresponding to the audio segment is used as a label to obtain a speech recognition training set.

[0077] In some optional implementations of this embodiment, the execution entity may perform the first step as follows:

[0078] For each video frame in the video frame sequence that includes text information, perform the following operations:

[0079] First, a plurality of texts to be spliced ​​corresponding to the video frame are determined, and the plurality of texts to be spliced ​​are spliced ​​with at least one video frame text in the video frame to obtain a plurality of spliced ​​texts.

[0080] Then, based on the edit distances between the multiple spliced ​​texts and the target audio segment text, a preset number of spliced ​​texts are selected from the multiple spliced ​​texts as the multiple texts to be spliced ​​corresponding to the next video frame of the video frame.

[0081] As an example, the execution entity may sort the edit distances from small to large and select a preset number of spliced ​​texts as the multiple texts to be spliced ​​corresponding to the next video frame of the video frame. The preset number may be set specifically according to actual conditions, for example, 10.

[0082] When the number of spliced ​​texts obtained is small (for example, the number of spliced ​​texts is less than a preset number), the execution entity may set a preset distance threshold to delete the spliced ​​texts whose editing distance is less than the preset distance threshold.

[0083] It can be understood that when the number of texts after splicing is large, the above-mentioned execution entity can also combine the method of selecting a preset number of texts and deleting texts with an editing distance less than a preset distance threshold to determine the multiple texts to be spliced ​​corresponding to the next video frame of the video frame.

[0084] As another example, the execution entity may determine the matching degree between the retained multiple concatenated texts and the audio text using the following formula:

[0085] Q i =-|d(p ki ,S k )-|‖S k ‖-‖p ki ‖||

[0086] Among them, d(·,·) represents the edit distance calculation function of two texts, ‖·‖ represents the length of the text, and p ki Represents the concatenated text, S k Indicates audio text, Q i Indicates the matching degree between two texts. In order to further reduce the number of spliced ​​texts obtained after each splicing, a matching threshold T is designed. h , when Q i <T h , then delete the concatenated text. For example, T h =-3.

[0087] In some optional implementations of this embodiment, the above-mentioned execution entity may also, for each video frame sequence in multiple video frame sequences, delete the training samples corresponding to the video frame sequence in the speech recognition training set in response to determining that the edit distance between the target video frame sequence text and the target audio segment text corresponding to the video frame sequence is greater than a preset distance threshold, thereby filtering out low-quality training samples.

[0088] Continue to see Figure 4 , Figure 4 FIG4 is a schematic diagram 400 of an application scenario of the method for generating a speech recognition training set according to this embodiment. Figure 4In an application scenario, server 401 first obtains audio 402 to be processed and video 403 to be processed. Video 403 to be processed includes text information corresponding to audio 402. Server 401 then identifies audio 402 to be processed, obtaining audio text 404; and identifies text information in video 403 to be processed, obtaining video text 405. Finally, server 401 determines the consistency between audio text 404 and video text 405, uses audio 402 to be processed as a speech sample, and video text 405 as a label, to obtain speech recognition training set 406.

[0089] The method provided by the above-mentioned embodiment of the present application obtains the audio to be processed and the video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; identifies the audio to be processed to obtain the audio text; identifies the text information in the video to be processed to obtain the video text; based on the consistency between the audio text and the video text, uses the audio to be processed as a voice sample and the video text as a label to obtain a speech recognition training set, thereby providing a method for automatically obtaining a speech recognition training set, and improving the flexibility and efficiency of constructing a speech recognition training set.

[0090] In some optional implementations of this embodiment, the above-mentioned execution entity can train an untrained initial speech recognition model or optimize a pre-trained speech recognition model based on the speech recognition training set.

[0091] Specifically, the above-mentioned execution entity adopts a machine learning algorithm, takes the audio to be processed in the training sample as input, and takes the input audio to be processed as the expected output, trains an untrained initial speech recognition model, or optimizes the pre-trained speech recognition model to obtain the final speech recognition model.

[0092] Continue to refer Figure 5 , shows a schematic process 500 of an embodiment of a method for generating a speech recognition training set according to the present application, comprising the following steps:

[0093] Step 501: Obtain the audio and video to be processed.

[0094] The video to be processed includes text information corresponding to the audio to be processed;

[0095] Step 502: Delete the silent parts in the audio to be processed according to the silence detection algorithm to obtain multiple non-silent audio segments.

[0096] Step 503: Identify multiple audio segments and obtain multiple audio segment texts included in the audio text.

[0097] Step 504: Determine a plurality of video frame sequences corresponding one-to-one to the plurality of audio segments from the video to be processed.

[0098] Step 505 : Identify text information in each video frame in a plurality of video frame sequences to obtain video frame text included in the video text.

[0099] Step 506: For each video frame sequence in the plurality of video frame sequences, perform the following operations:

[0100] Step 5061: Taking one video frame text among at least one video frame text identified in the video frames in the video frame sequence as a unit, splicing the text information included in each video frame in the video frame sequence to obtain multiple video frame sequence texts corresponding to the video frame sequence.

[0101] Step 5062: Determine the target video frame sequence text based on the edit distance between each video frame sequence text in the multiple video frame sequence texts and the target audio segment text, wherein the target audio segment text is the audio segment text corresponding to the audio segment corresponding to the video frame sequence.

[0102] Step 507: Each audio segment in the plurality of audio segments is used as a speech sample, and the target video frame sequence text corresponding to the audio segment is used as a label to obtain a speech recognition training set.

[0103] It can be seen from this embodiment that Figure 2 Compared with the corresponding embodiment, the process 400 of the method for generating a speech recognition training set in this embodiment specifically describes the segmentation process of the audio to be processed and the video to be processed, as well as the splicing process of the video frame text, thereby improving the accuracy of the training samples in the speech recognition training set.

[0104] Continue to refer Figure 6 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a device for generating a speech recognition training set. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0105] like Figure 6 As shown, the device for generating a speech recognition training set includes: an acquisition unit 601, configured to acquire audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; a first recognition unit 602, configured to recognize the audio to be processed and obtain audio text; a second recognition unit 603, configured to recognize text information in the video to be processed and obtain video text; an acquisition unit 604, configured to obtain a speech recognition training set based on the consistency between the audio text and the video text, with the audio to be processed as a speech sample and the video text as a label.

[0106] In some optional implementations of this embodiment, the first recognition unit 602 is further configured to: delete the silent part in the audio to be processed according to the silence detection algorithm to obtain multiple non-silent audio segments; recognize multiple audio segments to obtain multiple audio segment texts included in the audio text.

[0107] In some optional implementations of this embodiment, the second recognition unit 603 is further configured to: determine multiple video frame sequences corresponding one-to-one to multiple audio clips from the video to be processed; identify text information in each video frame in the multiple video frame sequences, and obtain the video frame text included in the video text.

[0108] In some optional implementations of this embodiment, the obtaining unit 604 is further configured to: for each video frame sequence in a plurality of video frame sequences, perform the following operations: taking one video frame text among at least one video frame text identified in the video frames in the video frame sequence as a unit, splicing the text information included in each video frame in the video frame sequence to obtain a plurality of video frame sequence texts corresponding to the video frame sequence; determining the target video frame sequence text according to the edit distance between each video frame sequence text in the plurality of video frame sequence texts and the target audio segment text, wherein the target audio segment text is the audio segment text corresponding to the audio segment corresponding to the video frame sequence; taking each audio segment in the plurality of audio segments as a speech sample and the target video frame sequence text corresponding to the audio segment as a label to obtain a speech recognition training set.

[0109] In some optional implementations of this embodiment, unit 604 is obtained and is further configured to: for each video frame including text information in the video frame sequence, perform the following operations: determine multiple texts to be spliced ​​corresponding to the video frame, and splice the multiple texts to be spliced ​​with at least one video frame text in the video frame to obtain multiple spliced ​​texts; based on the editing distance between the multiple spliced ​​texts and the target audio segment text, select a preset number of spliced ​​texts from the multiple spliced ​​texts as the multiple texts to be spliced ​​corresponding to the next video frame of the video frame.

[0110] In some optional implementations of this embodiment, the above-mentioned device also includes: a deletion unit (not shown in the figure), which is configured to delete the training samples corresponding to the video frame sequence in the speech recognition training set for each video frame sequence in a plurality of video frame sequences in response to determining that the editing distance between the target video frame sequence text corresponding to the video frame sequence and the target audio segment text is greater than a preset distance threshold.

[0111] In this embodiment, an acquisition unit in a speech recognition training set generation device acquires audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; a first recognition unit recognizes the audio to be processed to obtain an audio text; a second recognition unit recognizes the text information in the video to be processed to obtain a video text; the acquisition unit obtains a speech recognition training set based on the consistency between the audio text and the video text, using the audio to be processed as a speech sample and the video text as a label, thereby providing an automatic acquisition device for a speech recognition training set, which improves the flexibility and efficiency of constructing a speech recognition training set.

[0112] Reference below Figure 7 , which shows a device suitable for implementing the embodiments of the present application (eg Figure 1 Schematic diagram of the structure of the computer system 700 of the devices 101, 102, 103, 105 shown. Figure 7 The device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.

[0113] like Figure 7 As shown, the computer system 700 includes a processor (e.g., CPU, central processing unit) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the system 700 are also stored in the RAM 703. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0114] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read therefrom can be installed into the storage section 708 as needed.

[0115] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from a removable medium 711. When the computer program is executed by the processor 701, the above-mentioned functions defined in the method of the present application are performed.

[0116] It should be noted that the computer-readable medium of the present application may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0117] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the client computer, partially on the client computer, as a stand-alone software package, partially on the client computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the client computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0119] The units involved in the embodiments described in the present application can be implemented by software or by hardware. The described units can also be set in a processor. For example, they can be described as: a processor comprising an acquisition unit, a first recognition unit, a second recognition unit, and a obtaining unit. Among them, the names of these units do not constitute a limitation on the unit itself under certain circumstances. For example, the obtaining unit can also be described as "a unit that obtains a speech recognition training set based on the consistency between the audio text and the video text, with the audio to be processed as a voice sample and the video text as a label."

[0120] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the device, the computer device: obtains audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; recognizes the audio to be processed to obtain audio text; recognizes text information in the video to be processed to obtain video text; and based on the consistency between the audio text and the video text, uses the audio to be processed as a speech sample and the video text as a label to obtain a speech recognition training set.

[0121] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for generating a speech recognition training set, comprising: Acquire audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; Identify the audio to be processed and obtain an audio text; Identifying text information in the video to be processed to obtain video text; Based on the consistency between the audio text and the video text, the audio to be processed is used as a voice sample and the video text is used as a label to obtain the speech recognition training set, including: For each video frame sequence in the video to be processed, perform the following operations: For each video frame including text information in the video frame sequence, a plurality of texts to be spliced ​​corresponding to the video frame are determined, the plurality of texts to be spliced ​​are spliced ​​with at least one video frame text in the video frame to obtain a plurality of spliced ​​texts, and based on the edit distance between the plurality of spliced ​​texts and the target audio segment text, a preset number of spliced ​​texts are selected from the plurality of spliced ​​texts as the plurality of texts to be spliced ​​corresponding to the next video frame of the video frame, until the plurality of texts to be spliced ​​corresponding to the last video frame are obtained and determined as the plurality of video frame sequence texts corresponding to the video frame sequence, wherein the target audio segment text is the audio segment text corresponding to the audio segment corresponding to the video frame sequence in the audio to be processed; determining a target video frame sequence text according to an edit distance between each video frame sequence text in the plurality of video frame sequence texts and the target audio segment text; And also includes: Each audio segment in the audio to be processed is used as a speech sample, and the target video frame sequence text corresponding to the audio segment is used as a label to obtain the speech recognition training set.

2. The method according to claim 1, wherein The step of identifying the audio to be processed and obtaining an audio text includes: Deleting silent parts of the audio to be processed according to a silence detection algorithm to obtain multiple non-silent audio segments; The multiple audio segments are identified to obtain multiple audio segment texts included in the audio text.

3. The method according to claim 2, wherein: The identifying text information in the video to be processed to obtain the video text includes: Determining, from the video to be processed, a plurality of video frame sequences corresponding one-to-one to the plurality of audio clips; The text information in each video frame in the plurality of video frame sequences is identified to obtain the video frame text included in the video text.

4. The method according to claim 1, wherein Also includes: For each video frame sequence in the multiple video frame sequences, in response to determining that the edit distance between the target video frame sequence text and the target audio segment text corresponding to the video frame sequence is greater than a preset distance threshold, the training samples corresponding to the video frame sequence in the speech recognition training set are deleted.

5. A device for generating a speech recognition training set, comprising: an acquisition unit configured to acquire audio to be processed and video to be processed, wherein the video to be processed includes text information corresponding to the audio to be processed; A first recognition unit is configured to recognize the audio to be processed and obtain an audio text; A second recognition unit is configured to recognize text information in the video to be processed to obtain a video text; The obtaining unit is configured to obtain the speech recognition training set based on the consistency between the audio text and the video text, using the audio to be processed as a speech sample and the video text as a label, including: For each video frame sequence in the video to be processed, perform the following operations: For each video frame including text information in the video frame sequence, a plurality of texts to be spliced ​​corresponding to the video frame are determined, the plurality of texts to be spliced ​​are spliced ​​with at least one video frame text in the video frame to obtain a plurality of spliced ​​texts, and based on the edit distance between the plurality of spliced ​​texts and the target audio segment text, a preset number of spliced ​​texts are selected from the plurality of spliced ​​texts as the plurality of texts to be spliced ​​corresponding to the next video frame of the video frame, until the plurality of texts to be spliced ​​corresponding to the last video frame are obtained and determined as the plurality of video frame sequence texts corresponding to the video frame sequence, wherein the target audio segment text is the audio segment text corresponding to the audio segment corresponding to the video frame sequence in the audio to be processed; determining a target video frame sequence text according to an edit distance between each video frame sequence text in the plurality of video frame sequence texts and the target audio segment text; And also includes: Each audio segment in the audio to be processed is used as a speech sample, and the target video frame sequence text corresponding to the audio segment is used as a label to obtain the speech recognition training set.

6. The device according to claim 5, wherein The first recognition unit is further configured to: According to the silence detection algorithm, the silent parts in the audio to be processed are deleted to obtain multiple non-silent audio segments; the multiple audio segments are identified to obtain multiple audio segment texts included in the audio text.

7. The device according to claim 6, wherein The second recognition unit is further configured to: Determine a plurality of video frame sequences corresponding one-to-one to the plurality of audio segments from the video to be processed; identify text information in each video frame in the plurality of video frame sequences, and obtain the video frame text included in the video text.

8. The device according to claim 5, wherein Also includes: The deletion unit is configured to delete, for each video frame sequence in the multiple video frame sequences, a training sample corresponding to the video frame sequence in the speech recognition training set in response to determining that the edit distance between the target video frame sequence text and the target audio segment text corresponding to the video frame sequence is greater than a preset distance threshold.

9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

10. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Video subtitle determining method and video subtitle determining device

    CN106604125A

  • Learning video caption adding method and device, terminal equipment and storage medium

    CN111639233A