Method and apparatus for generating an audio recognition training set

The method generates an audio recognition training set by recognizing audio and video text from corresponding audio and video data, addressing the challenge of constructing diverse training sets for improving audio recognition model performance.

JP7700270B2Active Publication Date: 2025-06-30JINGDONG TECH HLDG CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023568632
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-08
Filing Date
2022-04-15
Publication Date
2025-06-30
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

Existing audio recognition technologies face challenges in constructing a large and diverse training set for improving the generalization performance of audio recognition models, which requires extensive manual labeling of audio data.

Method used

A method and apparatus for generating an audio recognition training set by obtaining processing target audio and video with corresponding text information, recognizing the audio and video text, and creating a training set based on the consistency between the audio and video text.

Benefits of technology

This approach enables the automatic acquisition of a speech recognition training set, improving the efficiency and flexibility in constructing the training set, thereby enhancing the performance of audio recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700270000004
    Figure 0007700270000004
  • Figure 0007700270000005
    Figure 0007700270000005
  • Figure 0007700270000006
    Figure 0007700270000006
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for generating a speech recognition training set. A specific embodiment of the method includes the steps of obtaining a target audio and a target video including text information corresponding to the target audio, recognizing the target audio to obtain an audio text, recognizing the text information in the target video to obtain a video text, and obtaining a speech recognition training set with the target audio as a speech sample and the video text as a label based on the correspondence between the audio text and the video text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] <Cross - Reference to Related Applications> This disclosure claims priority based on a Chinese patent application with application number 202110514350.X, filed on May 8, 2021, and titled "Method and Apparatus for Generating an Audio Recognition Training Set", and incorporates the entire text of the patent application by reference into this disclosure.

[0002] Embodiments of the present disclosure relate to the field of computer technology, and more specifically, to a method and apparatus for generating an audio recognition training set.

Background Art

[0003] In recent years, with the rapid development of deep learning technology, audio recognition using an automatic speech recognition (ASR) model based on a deep neural network has already become the mainstream in the current field of audio recognition technology. In order to improve the generalization performance of the audio recognition model, it is necessary to widely collect a large amount of audio data and optimize the audio recognition model with a training set constructed by manually labeling.

Summary of the Invention

[0004] Embodiments of the present disclosure provide a method and apparatus for generating an audio recognition training set.

[0005] In one or more embodiments, the present disclosure provides a method for generating an audio recognition training set, the method including: obtaining processing target audio and processing target video including text information corresponding to the processing target audio; recognizing the processing target audio to obtain audio text; recognizing the text information in the processing target video to obtain video text; and obtaining an audio recognition training set with the processing target audio as an audio sample and the video text as a label based on the consistency between the audio text and the video text.

[0006] In one or more embodiments, the present disclosure provides an apparatus for generating an audio recognition training set, the apparatus including: an acquisition unit configured to obtain processing target audio and processing target video including text information corresponding to the processing target audio; a first recognition unit configured to recognize the processing target audio to obtain audio text; a second recognition unit configured to recognize the text information in the processing target video to obtain video text; and an acquisition unit configured to obtain an audio recognition training set with the processing target audio as an audio sample and the video text as a label based on the consistency between the audio text and the video text.

[0007] In one or more embodiments, the present disclosure provides a computer-readable medium storing a computer program which, when executed by a processor, the above one or more embodiments realizes the method according to any of the above embodiments.

[0008] In one or more embodiments, the present disclosure provides an electronic device including one or more processors and a storage device storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors the above one or more embodimentsProvide an electronic device that implements the method according to any of the embodiments.

[0009] In one or more embodiments, the present disclosure, when executed by a processor, the above one or more embodiments a computer that implements the method according to any of the embodiments program is provided.

Brief Description of the Drawings

[0010] Other features, objects, and advantages of the present disclosure will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0011] Hereinafter, the present disclosure will be described in more detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining the related invention and do not limit the invention. For convenience of explanation, only the parts related to the invention are shown in the drawings.

[0012] In addition, the embodiments of the present disclosure and the features in the embodiments can be combined with each other as long as no contradiction occurs. Hereinafter, the present disclosure will be described in detail with reference to the drawings and embodiments.

[0013] FIG. 1 shows an exemplary architecture 100 to which a method and apparatus for generating an audio recognition training set according to the present disclosure can be applied.

[0014] As shown in FIG. 1, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The terminal devices 101, 102, 103 are communicatively connected to form a topology network, and the network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various types of connections such as wired, wireless communication links, or optical fiber cables.

[0015] The terminal devices 101, 102, and 103 may be hardware devices or software that support network connections for data exchange and data processing. When the terminal devices 101, 102, and 103 are hardware, they may be various electronic devices that support functions such as network connection, information acquisition, interaction, display, and processing, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When the terminal devices 101, 102, and 103 are software, they may be installed on the above-exemplified electronic devices. It may be implemented, for example, as a plurality of software or software modules for providing distributed services, or as a single software or software module. It is not particularly limited here.

[0016] The server 105 may be a server that provides various services (for example, a back-end processing server that acquires corresponding video and audio to be processed transmitted by a user via the terminal devices 101, 102, 103, performs information processing, and automatically constructs a speech recognition training set). Furthermore, the server can also train an initial speech recognition model based on the speech recognition training set or optimize a pre-trained speech recognition model. As an example, the server 105 may be a cloud server.

[0017] Note that the server may be hardware or software. When the server is hardware, it may be implemented as a distributed server cluster composed of a plurality of servers or as a single server. When the server is software, it may be implemented as a plurality of software or software modules (for example, software or software modules for providing distributed services) or as a single software or software module. It is not particularly limited here.

[0018] Note that the method for generating the speech recognition training set provided by the embodiments of the present disclosure may be executed by a server, or may be executed by a terminal device, or may be executed by the cooperation of the server and the terminal device. Correspondingly, each part (for example, each unit) included in the device for generating the speech recognition training set may all be provided in the server, may all be provided in the terminal device, or may be provided in the server and the terminal device respectively.

[0019] It should be understood that the numbers of the terminal device, the network, and the server in FIG. 1 are merely exemplary. According to the needs of implementation, the numbers of the terminal device, the network, and the server may be arbitrarily increased or decreased. When the electronic device on which the method for generating the speech recognition training set operates does not need to perform data transmission with other electronic devices, the system architecture may include only the electronic device (for example, a server or a terminal device) on which the method for generating the speech recognition training set operates. Next, referring to FIG. 2, FIG. 2 shows a flow 200 of an embodiment of the method for generating the speech recognition training set, including the following steps.

[0020] In step 201, the audio to be processed and the video to be processed are acquired.

[0021] In this embodiment, the execution entity of the method for generating the speech recognition training set (for example, the server shown in FIG. 1) may acquire the audio to be processed and the video to be processed remotely or locally in a wired network connection form or a wireless network connection form. The video to be processed includes the text information corresponding to the audio to be processed.

[0022] As an example, the data including the corresponding audio to be processed and the video to be processed may be various audio and video data such as movies, TV dramas, and short videos. The text information of the video to be processed is subtitle information, and the audio to be processed is the audio information corresponding to the subtitle information.

[0023] In this embodiment, the audio data represented by the audio to be processed can be various types of audio including, but not limited to, foreign language audio, Chinese audio, and dialect audio. The audio to be processed and the video to be processed may be data of a longer time or data of a shorter time.

[0024] In step 202, the audio to be processed is recognized to obtain audio text.

[0025] In this embodiment, the above execution entity can recognize the audio to be processed to obtain audio text.

[0026] As an example, the above execution entity can process the audio to be processed by an automatic speech recognition model to obtain audio text. The automatic speech recognition model is used to represent the correspondence between the audio to be processed and the text.

[0027] In some optional embodiments of this embodiment, the above execution entity can execute step 202 in the following manner.

[0028] In the first step, a mute detection algorithm is used to delete the muted parts in the audio to be processed to obtain a plurality of non-muted audio segments.

[0029] In this embodiment, the above execution entity can use the muted part of the audio to be processed as a split point to split the audio to be processed with the muted part deleted to obtain a plurality of audio segments.

[0030] When the obtained audio segment is long, the above execution entity may set a time threshold, further cut out the audio segment longer than the time threshold in time duration units represented by the time threshold, and record the start time and end time of each audio segment.

[0031] As an example, in order to prevent the audio from being completely segmented by the mute detection algorithm due to factors such as background music and to prevent the obtained audio segments from being long, a time threshold T is set, and an audio segment with a duration longer than the time threshold T is forcibly segmented into a plurality of segments with a duration of T. Here, the time threshold can be specifically set according to the actual situation. For example, T = 10s.

[0032] In the second step, a plurality of audio segments are recognized to obtain a plurality of audio segment texts included in the audio text.

[0033] In this embodiment, the above execution entity can input each of the plurality of audio segments into an automatic speech recognition model to obtain a plurality of audio segment texts. Here, the plurality of audio segments correspond one-to-one with the plurality of audio segment texts, and the plurality of audio segment texts constitute the audio text.

[0034] In step 203, the text information in the video to be processed is recognized to obtain video text.

[0035] In this embodiment, the above execution entity can recognize the text information in the video to be processed to obtain video text.

[0036] As an example, for each video frame constituting the video to be processed, the above execution entity uses OCR (Optical Character Recognition) technology to recognize the text information included in the video frame, and splices (connects) the text information corresponding to each video frame according to the playback order of the video frames in the video to be processed to obtain video text. Since the OCR technology is currently a relatively mature technology, it will not be elaborated here.

[0037] In some optional embodiments of the present embodiment, the above execution entity can execute the above step 203 in the following manner.

[0038] In the first step, a plurality of video frame sequences that correspond one-to-one to a plurality of audio segments are determined from the video to be processed.

[0039] In the present embodiment, for each of the plurality of audio segments, the above execution entity extracts a plurality of video frames corresponding to the audio segment from the video to be processed to obtain a video frame sequence.

[0040] JPEG0007700270000001.jpg74170

[0041] In the second step, text information in each video frame of the plurality of video frame sequences is recognized to obtain video frame text included in the video text.

[0042] In the present embodiment, the above execution entity can recognize text information in each video frame of the plurality of video frame sequences by means of OCR technology to obtain video frame text included in the video text.

[0043] As can be understood, for each video frame, the above execution entity may not be able to recognize text information, that is, the video frame may not include text information, and there may also be a case where a plurality of text information is recognized to obtain a plurality of video frame texts. For example, the plurality of video frame texts include subtitle information within the video frame and text information within the video frame screen (such as store name information on a store signboard, road name information on a road sign, advertising term information, etc.).

[0044] In other cases, the same text information exists in the text information included in adjacent frames. For example, the subtitle information included in adjacent video frames is the same.

[0045] In this embodiment, when the video frame does not contain text information, or when the video frame contains the same text information as the adjacent video frames, a preset identifier for representing such a case may be added to the video frame. Here, the preset identifier may be any preset identifier (for example, "Blank").

[0046] In step 204, based on the consistency between the audio text and the video text, a voice recognition training set is obtained by using the audio to be processed as an audio sample and the video text as a label.

[0047] In this embodiment, based on the consistency between the audio text and the video text, the execution entity can obtain a voice recognition training set by using the audio to be processed as an audio sample and the video text as a label.

[0048] As an example, when the audio text and the video text match, the execution entity obtains a voice recognition training set by using the audio to be processed as an audio sample and the video text as a label.

[0049] In some optional embodiments of this embodiment, the execution entity can execute step 204 as follows.

[0050] First, for each of a plurality of video frame sequences, the following operations are performed.

[0051] In the first step, in the video frames of the video frame sequence, taking one of the at least one recognized video frame text in the video frame as a unit, the text information included in each video frame in the video frame sequence is spliced to obtain a plurality of video frame sequence texts corresponding to the video frame sequence.

[0052] As an example, if the video frame sequence includes three video frames and the numbers of video frame texts corresponding to the three video frames are 3, 4, and 3 in sequence, then the number of multiple video frame sequence texts corresponding to the video frame sequence is 36 (3 × 4 × 3) in total.

[0053] In some optional embodiments, the set of video frame texts corresponding to each video frame includes, in addition to the video frame texts recognized from the video frame, a preset identifier indicating that the video frame includes the same text information as the adjacent video frames. The preset identifier may also represent a situation where no text information is included in the video frame.

[0054] Continuing to refer to the above example, after adding the preset identifier to the combination of video frame texts corresponding to each video frame, if the numbers of video frame texts corresponding to the three video frames are 4, 5, and 4 in sequence, then the number of multiple video frame sequence texts corresponding to the video frame sequence is 80 (4 × 5 × 4) in total.

[0055] As shown in FIG. 3, the preset identifier is "Blank". The recognition results corresponding to video frames 301, 302, and 303 are 304, 305, and 306 in sequence. Each of the video frame texts for each video frame can be combined with the video frame texts of other video frames to obtain multiple video frame sequence texts. For example, the recognition results of video frames 301, 302, and 303 can be combined as "What's the weather like today".

[0056] In the second step, based on the edit distance between each of the multiple video frame sequence texts and the target audio segment text, a target video frame sequence text is determined.

[0057] Here, the target audio segment text is the audio segment text of the audio segment corresponding to the video frame sequence. The edit distance is the minimum number of editing operations required to convert one string to the other between two strings.

[0058] As an example, the executing entity may determine, among a plurality of video frame sequence texts, the video frame sequence text with the smallest edit distance from the target audio segment text as the target video frame sequence text.

[0059] Then, each of the plurality of audio segments is used as an audio sample, and the target video frame sequence text corresponding to the audio segment is used as a label to obtain an audio recognition training set.

[0060] In some optional embodiments of this embodiment, the executing entity may execute the first step as follows.

[0061] For each video frame in the video frame sequence that contains text information, the following operations are performed.

[0062] First, determine a plurality of splicing target texts corresponding to the video frame, and splice the plurality of splicing target texts with at least one video frame text in the video frame to obtain a plurality of spliced texts.

[0063] Next, based on the edit distance between the plurality of spliced texts and the target audio segment text, select a preset number of spliced texts from the plurality of spliced texts as the plurality of splicing target texts corresponding to the next video frame of the video frame.

[0064] As an example, the above execution entity may sort the edited distances from small to large, and select the top preset number of spliced texts from the small ones as a plurality of text to be spliced corresponding to the next video frame of the video frame. Here, the preset number may be specifically set according to the actual situation (for example, it may be 10).

[0065] When the number of obtained spliced texts is small (for example, when the number of spliced texts is less than the preset number), the above execution entity may set a preset distance threshold and delete the spliced texts whose edited distance is smaller than the preset distance threshold.

[0066] It is understood that when the number of spliced texts is relatively large, the above execution entity can also determine a plurality of text to be spliced corresponding to the next video frame of the video frame by combining selecting a preset number of texts and deleting the texts whose edited distance is smaller than the preset distance threshold.

[0067] As yet another example, the above execution entity can determine the matching degree between the remaining plurality of spliced texts and the audio text according to the following formula. JPEG0007700270000002.jpg13127

[0068] JPEG0007700270000003.jpg38170

[0069] In some optional embodiments of the present embodiment, for each of the plurality of video frame sequences, when it is determined that the edit distance between the target video frame sequence text and the target audio segment text corresponding to the video frame sequence is greater than a preset distance threshold, the training sample corresponding to the video frame sequence in the speech recognition training set may be deleted to filter out and remove low-quality training samples.

[0070] Next, refer to FIG. 4, which is a schematic diagram 400 of an application scenario of a method for generating a speech recognition training set according to the present embodiment. In the application scenario of FIG. 4, first, the server 401 acquires the audio 402 to be processed and the video 403 to be processed. Here, the video 403 to be processed includes text information corresponding to the audio 402 to be processed. Then, the server 401 recognizes the audio 402 to be processed to obtain the audio text 404, and recognizes the text information of the video 403 to be processed to obtain the video text 405. Finally, based on the consistency between the audio text 404 and the video text 405, the server 401 uses the audio 402 to be processed as a speech sample and the video text 405 as a label to obtain a speech recognition training set 406.

[0071] The method provided by the above embodiment of the present disclosure acquires the audio to be processed and the video to be processed including the text information corresponding to the audio to be processed, recognizes the audio to be processed to obtain the audio text, recognizes the text information in the video to be processed to obtain the video text, and based on the consistency between the audio text and the video text, uses the audio to be processed as a speech sample and the video text as a label to obtain a speech recognition training set, thereby providing an automatic acquisition method for the speech recognition training set and improving the flexibility and efficiency of constructing the speech recognition training set.

[0072] In some optional embodiments of the present embodiment, the above execution entity may also train an untrained initial speech recognition model or optimize a pre-trained speech recognition model based on the speech recognition training set.

[0073] Specifically, the above execution entity adopts a machine learning algorithm, uses the audio to be processed in the training sample as the input, the input audio to be processed as the desired output, trains an untrained initial speech recognition model, or optimizes a pre-trained speech recognition model to obtain a final speech recognition model.

[0074] Next, refer to FIG. 5 showing a schematic flow 500 of an embodiment of a method for generating a speech recognition training set according to the present disclosure.

[0075] In step 501, the audio to be processed and the video to be processed are acquired.

[0076] Here, the video to be processed includes text information corresponding to the audio to be processed.

[0077] In step 502, the muted parts in the audio to be processed are removed by a mute detection algorithm to obtain a plurality of non-muted audio segments.

[0078] In step 503, the plurality of audio segments are recognized to obtain a plurality of audio segment texts included in the audio text.

[0079] In step 504, a plurality of video frame sequences corresponding one-to-one to the plurality of audio segments are determined from the video to be processed.

[0080] In step 505, the text information of each video frame in the plurality of video frame sequences is recognized to obtain the video frame text included in the video text.

[0081] In step 506, for each of the plurality of video frame sequences, the following processing is performed.

[0082] In step 5061, for each video frame in the video frame sequence, taking one of the at least one recognized video frame text in the video frame as a unit, the text information included in each video frame in the video frame sequence is spliced to obtain a plurality of video frame sequence texts corresponding to the video frame sequence.

[0083] In step 5062, based on the edit distance between each of the plurality of video frame sequence texts and the target audio segment text which is the audio segment text of the audio segment corresponding to the video frame sequence, the target video frame sequence text is determined.

[0084] In step 507, each of the plurality of audio segments is used as an audio sample, and the target video frame sequence text corresponding to the audio segment is used as a label to obtain an audio recognition training set.

[0085] From the perspective of this embodiment, compared with the embodiment corresponding to FIG. 2, the flow 400 of the method for generating the audio recognition training set in this embodiment specifically describes the splitting process of the audio to be processed and the video to be processed, and the splicing process of the video frame text, and it can be seen that the accuracy of the training samples in the audio recognition training set can be improved.

[0086] Continuing to refer to FIG. 6, as an implementation manner of the method shown in each of the above figures, the present disclosure provides an embodiment of an apparatus for generating an audio recognition training set. The embodiment of the apparatus corresponds to the embodiment of the method shown in FIG. 2, and the apparatus can be specifically applied to various electronic devices.

[0087] As shown in FIG. 6, an apparatus for generating an audio recognition training set includes an acquisition unit 601 configured to acquire audio to be processed and processed video including text information corresponding to the audio to be processed, a first recognition unit 602 configured to recognize the audio to be processed to obtain audio text, a second recognition unit 603 configured to recognize text information in the processed video to obtain video text, and an acquisition unit 604 configured to obtain an audio recognition training set with the audio to be processed as an audio sample and the video text as a label based on the consistency between the audio text and the video text.

[0088] In some optional embodiments of the present embodiment, the first recognition unit 602 is further configured to delete the muted parts in the audio to be processed by a mute detection algorithm to obtain a plurality of non-muted audio segments, and recognize the plurality of audio segments to obtain a plurality of audio segment texts included in the audio text.

[0089] In some optional embodiments of the present embodiment, the second recognition unit 603 is further configured to determine a plurality of video frame sequences corresponding one-to-one to the plurality of audio segments from the processed video, and recognize the text information of each video frame in the plurality of video frame sequences to obtain video frame texts included in the video text.

[0090] In some optional embodiments of the present embodiment, the acquisition unit 604 is further configured to For each of a plurality of video frame sequences, in the video frames of the video frame sequence, splicing the text information included in each video frame of the video frame sequence in units of one video frame text among at least one recognized video frame text, and obtaining a plurality of video frame sequence texts corresponding to the video frame sequence; and determining a target video frame sequence text according to the edit distance between each of the plurality of video frame sequence texts and a target audio segment text which is the audio segment text of the audio segment corresponding to the video frame sequence. Regarding each of a plurality of audio segments as an audio sample, and obtaining an audio recognition training set with the target video frame sequence text corresponding to the audio segment as a label. It is configured to perform the following.

[0091] In some optional embodiments of the present embodiment, the acquisition unit 604 For each video frame including text information in the video frame sequence determine a plurality of splicing target texts corresponding to the video frame, splice the plurality of splicing target texts and at least one video frame text in the video frame, and obtain a plurality of spliced texts; and further configured to select a preset number of spliced texts from the plurality of spliced texts based on the edit distance between the plurality of spliced texts and the target audio segment text, and use the selected spliced texts as the plurality of splicing target texts corresponding to the next video frame of the video frame.

[0092] In some optional embodiments of the present embodiment, for each of a plurality of video frame sequences, in response to determining that the edit distance between the target video frame sequence text corresponding to the video frame sequence and the target audio segment text is greater than a preset distance threshold, the apparatus further includes a deletion unit (not shown) configured to delete the training sample corresponding to the video frame sequence in the speech recognition training set.

[0093] In the present embodiment, in an apparatus for generating a speech recognition training set, an acquisition unit acquires a processing target audio and a processing target video including text information corresponding to the processing target audio. A first recognition unit recognizes the processing target audio to obtain audio text, and a second recognition unit recognizes the text information in the processing target video to obtain video text. An acquisition unit obtains a speech recognition training set with the processing target audio as a speech sample and the video text as a label based on the consistency between the audio text and the video text. Thereby, an automatic acquisition apparatus for a speech recognition training set is provided, and the flexibility and efficiency in constructing the speech recognition training set are improved.

[0094] Next, refer to FIG. 7, which shows a schematic structure of a computer system 700 suitable for devices (for example, devices 101, 102, 103, 105 shown in FIG. 1) for implementing the embodiments of the present disclosure. The devices shown in FIG. 7 are merely examples and do not impose any limitations on the functions and usage ranges of the embodiments of the present disclosure.

[0095] As shown in FIG. 7, the computer system 700 includes a processor (e.g., a CPU, a central processing unit) 701 that can execute various appropriate operations and processes by a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data necessary for the operation of the system 700 are further stored in the RAM 703. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0096] An input unit 706 including a keyboard, a mouse, etc., an output unit 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc., a storage unit 708 including a hard disk, etc., and a communication unit 709 including a network interface card such as a LAN card, a modem, etc. are connected to the I / O interface 705. The communication unit 709 executes communication processing via a network such as the Internet. A driver 710 is connected to the I / O interface 705 as necessary. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in a drive 710 as necessary so that a computer program read therefrom is installed in the storage unit 708 as necessary.

[0097] In particular, according to an embodiment of the present disclosure, the process described with reference to the above-described flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure is a computer embodied in a computer-readable medium programcomprises, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 709, and / or can also be installed from the removable media 711. When the computer program is executed by the processor 701, the above-described functions limited by the method of the present disclosure are executed.

[0098] Note that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, electrical connections by one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program that can be used by or incorporated into a command execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may further be any computer-readable medium other than the computer-readable storage medium. The computer-readable medium can transmit, propagate, or transmit a program used by or incorporated into a command execution system, apparatus, or device. The program code included in the computer-readable medium can be transmitted via any suitable medium, which includes, but is not limited to, wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0099] The computer program code for performing the operations of the present disclosure can be created in one or more programming languages, or combinations thereof, and the programming languages include object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, executed partially on the user's computer while being executed partially on a remote computer, or executed entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet connection service provided by an Internet service provider).

[0100] The flowcharts and block diagrams among the drawings relate to the apparatuses, methods, and computers according to various embodiments of the present disclosure programIllustrated are system architectures, functions, and operations that can be realized thereby. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code. One or more executable instructions for realizing a given logic function are included in the module, the program segment, or the part of code. Note that in some alternative embodiments, the functions shown in the blocks can be executed in an order different from that shown in the drawings. For example, two consecutively shown blocks may actually be executed substantially in parallel or sometimes in a reverse order depending on the functions involved. Further to be noted is that all blocks in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented in a dedicated hardware-based system for executing a given function or operation, or may be implemented in a combination of dedicated hardware and computer instructions.

[0101] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. The described units may be provided in a processor, and for example, may be described as "a processor including an acquisition unit, a first recognition unit, a second recognition unit, and an acquisition unit". Here, the names of these units do not limit the unit itself in some cases. For example, the acquisition unit may be described as "a unit that acquires an audio recognition training set using the processing target audio as an audio sample and the video text as a label based on the consistency between the audio text and the video text".

[0102] On the one hand, the present disclosure further provides a computer-readable medium, which may be included in the devices described in the above embodiments or may exist separately without being implemented in the devices. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the device, it causes the computer device to perform steps of: obtaining a processing target audio and a processing target video including text information corresponding to the processing target audio; recognizing the processing target audio to obtain audio text; recognizing the text information in the processing target video to obtain video text; and obtaining an audio recognition training set with the processing target audio as an audio sample and the video text as a label based on the consistency between the audio text and the video text.

[0103] The above description is only an explanation of the preferred embodiments of the present disclosure and the applicable technical principles. Those skilled in the art should understand that the scope of the invention according to the present disclosure is not limited to the technical solutions consisting of specific combinations of the above technical features, and other technical solutions consisting of any combination of the above technical features or their equivalent features should also be included without departing from the spirit of the present disclosure. For example, a technical solution formed by replacing the above features with technical features having similar functions (not limited thereto) disclosed in the present disclosure can be cited.

Claims

1. A method for generating an audio recognition training set, comprising: obtaining a target audio and a target video including text information corresponding to the target audio; recognizing the target audio to obtain audio segment texts of a plurality of audio segments included in the target audio; determining, from the target video, a plurality of video frame sequences that correspond one-to-one to the plurality of audio segments, and recognizing text information for each video frame in the plurality of video frame sequences to obtain video frame texts; for each of the plurality of video frame sequences, splicing, in units of one video frame text out of at least one recognized video frame text, the text information included in each video frame of the video frame sequence to obtain a plurality of video frame sequence texts corresponding to the video frame sequence; and determining target video frame sequence texts according to edit distances between each of the plurality of video frame sequence texts and target audio segment texts that are audio segment texts of the audio segments corresponding to the video frame sequence; obtaining an audio recognition training set by using, as audio samples, each of the plurality of audio segments based on the consistency between the audio segment texts and the video frame texts, and using the target video frame sequence texts corresponding to the audio segments as labels; A method for generating an audio recognition training set, comprising the above steps.

2. The step of recognizing the target audio includes: The method according to claim 1, further comprising obtaining, by a mute detection algorithm, a plurality of non-muted audio segments by removing muted portions in the target audio.

3. In the video frames of the video frame sequence, splicing the text information included in each video frame of the video frame sequence in units of one of the at least one video frame text recognized, and obtaining a plurality of video frame sequence texts corresponding to the video frame sequence, which is In the video frame sequence, for each video frame in which text information is included, determining a plurality of splicing target texts corresponding to the video frame, and splicing the plurality of splicing target texts and at least one video frame text in the video frame to obtain a plurality of spliced texts, and selecting a preset number of spliced texts from the plurality of spliced texts based on the edit distance between the plurality of spliced texts and the target audio segment text to be the plurality of splicing target texts corresponding to the next video frame of the video frame. The method according to claim 2, comprising performing

4. For each of the plurality of video frame sequences, in response to determining that the edit distance between the target video frame sequence text corresponding to the video frame sequence and the target audio segment text is greater than a preset distance threshold, further comprising the step of deleting the training sample corresponding to the video frame sequence in the speech recognition training set. The method according to claim 3

5. An apparatus for generating a speech recognition training set, comprising an acquisition unit configured to acquire a processing target audio and a processing target video including text information corresponding to the processing target audio, and a first recognition unit configured to recognize the processing target audio and obtain audio segment texts of a plurality of audio segments included in the processing target audio, and A second recognition unit configured to determine, from the video to be processed, a plurality of video frame sequences that correspond one-to-one to the plurality of audio segments, and to recognize text information for each video frame in the plurality of video frame sequences to obtain video frame text. For each of the plurality of video frame sequences, in the video frames of the video frame sequence, splicing the text information included in each video frame of the video frame sequence in units of one video frame text among at least one recognized video frame text, and obtaining a plurality of video frame sequence texts corresponding to the video frame sequence; and determining a target video frame sequence text according to the edit distance between each of the plurality of video frame sequence texts and a target audio segment text which is the audio segment text of the audio segment corresponding to the video frame sequence. An apparatus for generating an audio recognition training set, including an acquisition unit configured to obtain an audio recognition training set by using each of the plurality of audio segments as an audio sample and the target video frame sequence text corresponding to the audio segment as a label based on the consistency between the audio segment text and the video frame text.

6. The first recognition unit further The apparatus according to claim 5, wherein the first recognition unit is configured to delete muted portions in the audio to be processed by a mute detection algorithm to obtain a plurality of non-muted audio segments.

7. The acquisition unit further For each video frame including text information in the video frame sequence, determining a plurality of splicing target texts corresponding to the video frame, and splicing the plurality of splicing target texts and at least one video frame text in the video frame to obtain a plurality of spliced texts. Based on the edit distance between the plurality of spliced texts and the target audio segment text, select a preset number of spliced texts from the plurality of spliced texts, and use them as a plurality of texts to be spliced corresponding to the next video frame of the video frame, the apparatus according to claim 6, which is configured to perform the above.

8. For each of the plurality of video frame sequences, in response to determining that the edit distance between the target video frame sequence text and the target audio segment text corresponding to the video frame sequence is greater than a preset distance threshold, a deletion unit configured to delete the training sample corresponding to the video frame sequence in the speech recognition training set is further included in the apparatus according to claim 7.

9. A computer-readable medium storing a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 4.

10. An electronic device comprising one or more processors and a storage device storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

11. A computer program that, when executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Meta information generating method and device, retrieval method and device

    JP2005092295A

  • Scene meta information generation apparatus and scene meta information generating method

    JP2019198074A

  • Voice recognition result formatted model learning device and its program

    JP2020030367A

  • Training speech recognition using captions

    US20150088508A1