Method, apparatus, device, and storage medium for obtaining speech-to-text training corpus
By obtaining video keyframes and processing subtitle recognition intervals, the problem of low subtitle extraction efficiency is solved, and a high-quality phon-to-text training corpus is generated, which is suitable for the phon-to-text model.
Patent Information
- Application Number
- CN202110282332.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-03-16
AI Technical Summary
In the prior art, subtitle extraction efficiency is low, requires manual annotation, and the identified subtitles have not been processed, so they are not suitable as training corpus for speech to text models.
By obtaining the keyframes of the target video, determining the subtitle recognition interval, and identifying and processing subtitles, including deduplication, merging, filtering, eliminating uncommon characters, and generating suitable pronunciation-to-text training corpus.
It improves the efficiency of subtitle extraction and the convenience of obtaining corpus from pronunciation to text training, improves the quality of corpus, and reduces the typo rate.
Smart Images

Figure CN113705300B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer software technology, and in particular, to a method, apparatus, device, and storage medium for obtaining speech-to-text training corpus. Background Art
[0002] With the development of computer software technology, the video resources on the Internet have increased significantly. Subtitle extraction technology for videos is required in various scenarios. For example, in the training process of a speech-to-text model, in order to obtain training corpus, it is necessary to extract subtitles from videos. The inventors of this application found in the research and practice process that in the prior art, subtitle extraction technology requires manual annotation of subtitle intervals in videos so as to perform text recognition within the subtitle intervals of the videos. For example, methods such as Optical Character Recognition (OCR) technology are used to recognize the text within the manually selected subtitle intervals to obtain the subtitles of the videos, which consumes a lot of manpower and has a low subtitle recognition efficiency. In the prior art, the subtitles recognized by methods such as OCR technology are not processed after recognition, and the subtitle extraction method is rough and not suitable as training corpus for speech-to-text models. Summary of the Invention
[0003] Embodiments of this application provide a method, apparatus, device, and storage medium for obtaining speech-to-text training corpus, which can improve the subtitle extraction efficiency of videos, improve the convenience of obtaining speech-to-text training corpus, have simple operations, and high applicability.
[0004] In a first aspect, embodiments of this application provide a method for obtaining speech-to-text training corpus, and the method includes:
[0005] Obtain multiple target video key frames of a target video, and determine the text positions and text contents from each of the target video key frames;
[0006] Determine a subtitle recognition interval of the target video according to the text positions and text contents in each of the target video key frames, where different target video key frames correspond to different text contents at the same position in the subtitle recognition interval;
[0007] Recognize the subtitles of the target video according to the subtitle recognition interval to obtain subtitles to be processed from the target video, and perform character processing on the subtitles to be processed according to a preset corpus acquisition rule to obtain the target subtitles of the target video;
[0008] Generate speech-to-text training corpus for video speech recognition according to the target video and the target subtitles.
[0009] In combination with the first aspect, in a possible implementation manner, obtaining multiple target video key frames of a target video includes:
[0010] Obtain the video to be processed, and determine the video key frames of the video to be processed and the number of frames of the video key frames;
[0011] When the number of frames of the video key frames is greater than or equal to the frame number threshold, determine the video to be processed as the target video, and determine multiple video key frames of the video to be processed as multiple target video key frames.
[0012] Combined with the first aspect, in a possible implementation manner, obtaining multiple target video key frames of the target video includes:
[0013] Obtain the video to be processed, and determine the video key frames of the video to be processed and the Chinese character appearance rate of the video key frames;
[0014] When the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, determine the video to be processed as the target video, and determine multiple video key frames of the video to be processed as multiple target video key frames.
[0015] Combined with the first aspect, in a possible implementation manner, determining the subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame includes:
[0016] Determine at least one text position from the text positions of each target video key frame as at least one candidate text recognition interval, and the number of times the text content in the candidate text recognition interval appears repeatedly is less than the number threshold;
[0017] Determine the jitter degree of each candidate text recognition interval, and determine the candidate text recognition interval with the jitter degree less than or equal to the jitter degree threshold as the subtitle recognition interval of the target video.
[0018] Combined with the first aspect, in a possible implementation manner, the above method further includes:
[0019] Determine the text similarity of the text contents that appear in the text positions of each target video key frame;
[0020] When the text similarity of any two occurrences of the text content in any text position is greater than the threshold, determine the text content of any two occurrences as the repeatedly appearing text content, and determine that the above-mentioned any text position is not used as the candidate text recognition interval.
[0021] Combined with the first aspect, in a possible implementation manner, performing character processing on the subtitle to be processed according to the preset corpus acquisition rule includes:
[0022] Divide the subtitle to be processed into multiple subtitle clauses according to a preset time interval, and deduplicate each subtitle clause. Merge the subtitle clauses with a character length less than the character length threshold in the deduplicated subtitle clauses to obtain the subtitle to be processed after merging;
[0023] Based on the characters in the subtitle to be processed after merging, perform subtitle clause screening to determine the target subtitle of the target video.
[0024] Combined with the first aspect, in a possible implementation manner, performing subtitle clause screening based on the characters in the subtitle to be processed after merging includes:
[0025] Eliminate the subtitle clauses containing rare characters in the subtitle to be processed after merging to screen out the subtitle clauses without rare characters, where the rare characters include at least one of letters, numbers, and rare radicals.
[0026] In the embodiments of the present application, by obtaining multiple target video key frames of the target video, further, determine the text position and text content from each target video key frame, so that the subtitle recognition interval of the target video can be determined according to the text position and text content in each target video key frame. Among them, it can be understood that the text content corresponding to the same position in the subtitle recognition interval is different for different target video key frames. Perform text recognition on the subtitle of the target video according to the subtitle recognition interval, and the subtitle to be processed can be obtained from the target video. Perform post-processing such as division, deduplication, and merging on the subtitle to be processed according to the corpus acquisition rule. Further, the subtitle clauses containing rare characters in the subtitle to be processed after merging can be eliminated to obtain the target subtitle, so that the error rate of the target subtitle is reduced to within the standard. Furthermore, the audio-to-text training corpus for video speech recognition is generated according to the target video and the target subtitle. Thus, it is possible to automatically obtain the video, extract and screen the subtitle of the video, and then obtain the training corpus suitable for audio-to-text, improve the acquisition efficiency of the audio-to-text training corpus, and at the same time improve the corpus quality of the audio-to-text training corpus.
[0027] In a second aspect, an apparatus for obtaining an audio-to-text training corpus according to an embodiment of the present application includes:
[0028] A video acquisition module, configured to obtain multiple target video key frames of a target video, and determine a text position and text content from each target video key frame;
[0029] An interval division module, configured to determine a subtitle recognition interval of the target video according to the text position and text content in each target video key frame, where the text content corresponding to the same position in the subtitle recognition interval is different for different target video key frames;
[0030] A subtitle extraction module, configured to recognize the subtitles of a target video according to subtitle recognition intervals, so as to obtain subtitles to be processed from the target video, and perform character processing on the subtitles to be processed according to a preset corpus acquisition rule to obtain the target subtitles of the target video;
[0031] A corpus generation module, configured to generate a speech-to-text training corpus for video speech recognition according to the target video and the target subtitles.
[0032] Combined with the second aspect, in a possible implementation manner, the above video acquisition module includes:
[0033] A frame number determination unit, configured to obtain a video to be processed, determine the video key frames of the video to be processed and the number of frames of the video key frames. When the number of frames of the video key frames is greater than or equal to a frame number threshold, determine the video to be processed as the target video, and determine multiple video key frames of the video to be processed as multiple target video key frames.
[0034] Combined with the second aspect, in a possible implementation manner, the above video acquisition module includes:
[0035] A Chinese character determination unit, configured to obtain a video to be processed, determine the video key frames of the video to be processed and the Chinese character appearance rate of the video key frames. When the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to an appearance rate threshold, determine the video to be processed as the target video, and determine multiple video key frames of the video to be processed as multiple target video key frames.
[0036] Combined with the second aspect, in a possible implementation manner, the above interval division module includes:
[0037] An interval de-duplication unit, configured to determine at least one text position from the text positions of each target video key frame as at least one candidate text recognition interval, and the number of times the text content of the candidate text recognition interval appears repeatedly is less than a number threshold;
[0038] An interval determination unit, configured to determine the jitter degree of each candidate text recognition interval, and determine the candidate text recognition interval with a jitter degree less than or equal to a jitter degree threshold as the subtitle recognition interval of the target video.
[0039] Combined with the second aspect, in a possible implementation manner, the above interval division module further includes:
[0040] A text recognition module, configured to determine the text similarity of the text content appearing in the text positions of each target video key frame. When the text similarity of any two occurrences of the text content in any text position is greater than a threshold, determine the text content of any two occurrences as repeatedly appearing text content, and determine that the above any text position is not used as a candidate text recognition interval.
[0041] In combination with the second aspect, in a possible implementation, the above subtitle extraction module includes:
[0042] A subtitle division unit, configured to divide the subtitle to be processed into multiple subtitle clauses according to a preset time interval, remove duplicates from each subtitle clause, and merge subtitle clauses with a character length less than a character length threshold in the de-duplicated subtitle clauses, so as to obtain the subtitle to be processed after merging;
[0043] A subtitle screening unit, configured to screen subtitle clauses based on the characters in the subtitle to be processed after merging, so as to determine the target subtitle of the target video.
[0044] In combination with the second aspect, in a possible implementation, the above subtitle screening unit includes:
[0045] A clause elimination subunit, configured to eliminate subtitle clauses containing rare characters in the subtitle to be processed after merging, so as to screen out subtitle clauses that do not contain rare characters, where the rare characters include at least one of letters, numbers, and rare radical components.
[0046] In the embodiments of the present application, by obtaining multiple target video key frames of the target video, further, determining the text positions and text contents from each target video key frame, the subtitle recognition interval of the target video can be determined according to the text positions and text contents in each target video key frame. It can be understood that the text contents corresponding to the same position in the subtitle recognition interval are different for different target video key frames. Performing text recognition on the subtitle of the target video according to the subtitle recognition interval can obtain the subtitle to be processed from the target video. After post-processing such as dividing, de-duplicating, and merging the subtitle to be processed according to the corpus acquisition rules, further, subtitle clauses containing rare characters in the subtitle to be processed after merging can be eliminated to obtain the target subtitle, so as to reduce the error rate of the target subtitle to within the standard. Furthermore, a speech-to-text training corpus for video speech recognition is generated according to the target video and the target subtitle. Thus, it is possible to automatically obtain a video, extract and screen the subtitle of the video, and further obtain a training corpus suitable for speech-to-text, improve the acquisition efficiency of the speech-to-text training corpus, and at the same time improve the corpus quality of the speech-to-text training corpus.
[0047] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor and a memory, and the processor and the memory are connected to each other. The memory is used to store a computer program that supports the terminal to execute the method provided in the second aspect and / or any possible implementation manner of the second aspect. The computer program includes program instructions, and the processor is configured to call the above program instructions to execute the method provided in the second aspect and / or any possible implementation manner of the second aspect.
[0048] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is executed by a processor to implement the method provided in the second aspect and / or any possible implementation manner of the second aspect.
[0049] In the embodiment of the present application, by obtaining multiple target video key frames of a target video, and further determining the text positions and text contents from each target video key frame, the subtitle recognition intervals of the target video can be determined according to the text positions and text contents in each target video key frame. It can be understood that the text contents corresponding to the same position in the subtitle recognition intervals are different for different target video key frames. Performing text recognition on the subtitles of the target video according to the subtitle recognition intervals, the subtitles to be processed can be obtained from the target video. After post-processing such as dividing, de-duplicating, and merging the subtitles to be processed according to the corpus acquisition rules, further, the subtitle clauses containing rare characters in the merged subtitles to be processed can be removed to obtain the target subtitles, and the error rate of the target subtitles can be reduced to within the standard. Furthermore, the audio-to-text training corpus for video speech recognition can be generated according to the target video and the target subtitles. Thus, it can be realized to automatically obtain the video, extract and screen the subtitles of the video, and further obtain the training corpus suitable for audio-to-text, improving the acquisition efficiency of the audio-to-text training corpus and at the same time improving the corpus quality of the audio-to-text training corpus. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0051] Figure 1 is a schematic diagram of the network architecture provided by an embodiment of the present application;
[0052] Figure 2 is a schematic flowchart of a method for obtaining an audio-to-text training corpus provided by an embodiment of the present application;
[0053] Figure 3 is a schematic diagram of the scene for extracting target video key frames provided by an embodiment of the present application;
[0054] Figure 4 is a schematic structural diagram of a method for obtaining an audio-to-text training corpus provided by an embodiment of the present application;
[0055] Figure 5 is a schematic diagram of the scene of the video subtitle interval provided by an embodiment of the present application;
[0056] Figure 6 It is a schematic diagram of the scenario for determining the subtitle recognition interval provided by an embodiment of the present application;
[0057] Figure 7 It is a schematic flowchart of post-processing subtitles provided by an embodiment of the present application;
[0058] Figure 8 It is another schematic flowchart of the method for obtaining the speech-to-text training corpus provided by an embodiment of the present application;
[0059] Figure 9 It is a schematic structural diagram of the device for obtaining the speech-to-text training corpus provided by an embodiment of the present application;
[0060] Figure 10 It is a schematic structural diagram of the terminal device provided by an embodiment of the present application. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0062] Please refer to Figure 1 , which is a schematic structural diagram of the network architecture provided by an embodiment of the present invention. As Figure 1 shown, the network architecture may include a cloud server 2000 and a user terminal cluster; the user terminal cluster may include multiple user terminals. As Figure 1 shown, specifically, it includes user terminals 3000a, 3000b,..., 3000n; as Figure 1 shown, user terminals 3000a, 3000b,..., 3000n can respectively establish a data connection relationship with the cloud server 2000 under certain data interaction conditions, so as to be able to perform data interaction with the cloud server 2000.
[0063] For ease of understanding, an embodiment of the present application may select one user terminal as the target user terminal among the Figure 1 multiple user terminals shown. The target user terminal may include: intelligent terminals such as smartphones, tablets, and desktop computers that can perform subtitle extraction functions on videos (for example, social application functions, short video application functions, film and television application functions, etc.) and then generate speech-to-text training corpora. For example, an embodiment of the present application may use Figure 1The user terminal 3000a shown above is used as the target user terminal, and one or more applications can be integrated in the target user terminal. It should be understood that the target applications integrated in the target user terminal can be collectively referred to as the target application clients. Among them, the above target applications can include social applications (such as WeChat, QQ), short video applications (such as Weishi), film and television applications (such as Tencent Video), etc., which can perform video subtitle extraction, screen the subtitles, and then generate audio-to-text training corpora.
[0064] It can be understood that the method for obtaining the audio-to-text training corpus described in the embodiments of the present application can be applied to all application scenarios for obtaining the audio-to-text training corpus in a web page or an application client (i.e., the aforementioned target application). Among them, when the target application with the function of obtaining the audio-to-text training corpus runs in the target user terminal to extract subtitles from a video, the video for which the target user terminal extracts subtitles can include the videos pre-built in the target application and the videos currently obtained from the server 2000 through the network.
[0065] It should be understood that the application videos pre-built in the target application and the videos currently obtained from the server in the embodiments of the present application can be collectively referred to as target videos, and subtitles need to be extracted to generate an audio-to-text training corpus. Thus, it can be seen that the embodiments of the present application can extract subtitles from the target videos during the operation of the web page or the target application to obtain the target videos and the target subtitles of the target videos, so as to improve the generation quality of the training corpus and reduce the occupation of the system memory during the subtitle extraction process of the videos when generating the audio-to-text corpus in the web page or the application client.
[0066] Optionally, in the embodiments of the present application, before the target user terminal runs the target application, it can be pre-obtained from the above Figure 1The to-be-processed video obtained from the server 2000 shown is screened to obtain a target video and store it in a specified storage space. When the target user terminal runs the target application, it can directly load the target video from the above-mentioned specified storage space, so as to reduce the system performance loss brought to the target user terminal during the operation of the target application (for example, the target user terminal can reduce the occupation of system memory by video data). Optionally, in the embodiment of the present application, before the target user terminal runs the target application, the server 2000 can also be used to extract key frames from the to-be-processed video representation in advance, and screen the to-be-processed video to obtain the target video and the target video key frames. The target user terminal can send a data download instruction (i.e., a data loading instruction) to the server 2000 through the network when running the target application, so that the server can determine whether the target user terminal meets the key frame extraction condition based on the terminal identifier carried in the download instruction. If it is determined by the server 2000 that the target user terminal meets the key frame extraction condition, the target user terminal can obtain the target video and the target video key frames stored after key frame extraction in advance through the server 2000, so that when the target application runs on the target user terminal, the system performance loss can be reduced, and the acquisition efficiency of the speech-to-text corpus can be improved. It can be seen that the target user terminal in the embodiment of the present application can also perform key frame extraction and screening on the to-be-processed video in the server 2000 before running the target application to obtain the target video and the target video key frames.
[0067] Optionally, before the target user terminal runs the target application, it can extract subtitles from the target video obtained from the server 2000 shown above to obtain the aforementioned target video and target subtitles. For example, taking the above-mentioned target application as a short video application (Tencent Weishi) as an example, the target user terminal can load and display the target video and target subtitles through the short video application, and generate a speech-to-text training corpus according to the target video and target subtitles. Figure 1 The to-be-processed video described in the embodiment of the present application may include short videos, movies and TV dramas, music greeting cards, and may also include audio containing subtitle information. For example, taking the above-mentioned target application as a short video application as an example, the target user terminal can capture the videos uploaded, downloaded or browsed by the user through the short video application and perform subtitle extraction to obtain the target video and target subtitles, and then generate a speech-to-text training corpus.
[0068] For details, please refer to
[0069] which Figure 2 is Figure 2 a schematic flowchart of a method for obtaining a speech-to-text training corpus provided by an embodiment of the present application. For the convenience of description, in the embodiment of the present application, the terminal is used as the execution subject of the method for obtaining the speech-to-text training corpus for description. AsFigure 2 As shown, the method for obtaining the speech-to-text training corpus provided by this application includes:
[0070] S101: Obtain multiple target video key frames of the target video, and determine the text positions and text contents from each of the target video key frames.
[0071] In some feasible implementation manners, when the terminal obtains the target video, it may perform key frame extraction on the target video to obtain the target video key frames. For details, please refer to Figure 3 , Figure 3 which is a schematic diagram of the scenario for target video key frame extraction provided by the embodiments of this application. As Figure 3As shown, a video key frame refers to the frame where the key actions of the characters or objects in the video are in motion changes, which is equivalent to the original painting in two-dimensional animation. When the puppy and subtitles in the video are constantly changing, the key information in the video can be completely displayed through the video key frames. Furthermore, the text in the video can be recognized using the key frames to obtain the text position and text content. For example, through the text recognition in the key frames, it can be determined that words unrelated to the subtitles, such as the word "Wang" that appears in the middle position of each target video key frame, and the text serving as subtitles, such as "Toutou is a dog", "Very cute", "Hope you like it too", that appear in the lower position of each target video key frame. The video between key frames can be created and added by software, which is called a transition frame or an intermediate frame. A frame is the smallest unit of a single image in a video, equivalent to each frame of film in a movie. On the time axis of the video, one frame is represented as a grid or a marker. After the terminal obtains the video to be processed, it can extract key frames from the video to be processed. The methods for extracting key frames include, but are not limited to: the key frame extraction method based on shots, the key frame extraction method based on motion analysis, and the key frame extraction method based on video clustering. Among them, the key frame extraction method based on shots is the first developed and currently the most mature general method in the field of video processing. The general implementation process of this method is: first, the video file is segmented according to shot changes, and then the first and last frames in each shot of the video are selected as key frames. The advantage of this method is that it is very simple to implement and has a small amount of calculation. However, this method has great limitations. When the content in the video changes violently and the scene is very complex, selecting the first and last frames in the shot does not represent all the content changes in the video. Therefore, this method can no longer meet the standards and requirements for key frame extraction in today's society. The key frame extraction method based on motion analysis is a method for extracting key frames based on the attributes of object motion characteristics. The implementation process of this method is: analyze the optical flow of object motion in the video shot, and each time select the video frame with the least number of optical flow movements in the video shot as the extracted key frame. This method can extract an appropriate number of key frames from most video shots, and the extracted key frames can also effectively express the motion characteristics of the video. However, this method has the characteristic of poor robustness because it not only depends on the local characteristics of object motion, but also the calculation process is relatively complex, and the algorithm has a large overhead cost in terms of time. The key frame extraction method based on video clustering can divide video frames into several clusters through clustering during the process of extracting key frames, and then select the corresponding frames in each cluster as key frames. The video key frames extracted using this algorithm not only have a small redundancy, but also the key frames can accurately reflect all the content that occurs in the video.However, in the process of dividing clustering clusters, the clustering-based method does not fully consider the sequential change order of time among each frame, and a certain number of clusters need to be preset before clustering. Therefore, the applicability of this method is limited to a certain extent. In the specific implementation manner, there are many methods for extracting key frames from a video. The key frames of the video can be extracted by using a specific key frame extraction method according to the characteristics of the video itself, or the key frames of the video can be extracted by combining several key frame extraction methods. Specifically, it can be determined according to the actual application scenario and will not be limited here. It can be understood that any process of extracting key frames from a video is covered by the protection scope of this application.
[0072] In some feasible implementation manners, in order to ensure the quality of the subsequent speech-to-text training corpus, the terminal can batch obtain the videos to be processed before obtaining multiple target video key frames of the target video, extract the key frames from the videos to be processed, and screen through the video key frames to obtain the target video and the target video key frames suitable as the speech-to-text training corpus. For specific details, please refer to Figure 4 , Figure 4 which is a schematic diagram of data processing of the method for obtaining the speech-to-text training corpus provided by the embodiment of this application. As Figure 4 shown in part a of, the terminal can batch download the videos to be processed during the video acquisition process, and perform acquisition management on the batch-obtained videos to be processed (for example, classify according to video types, classify according to source applications, etc.), and then perform video preprocessing on the videos (for example, image grayscale processing, image stretching (such as stretching the video image to a fixed pixel size), picture sharpening, etc.), which is convenient for subsequent positioning and extraction of subtitle intervals of the videos.
[0073] In some feasible implementation manners, the terminal can determine the video key frames of the video to be processed and the number of frames of the video key frames. When the number of frames of the video key frames is greater than or equal to the frame number threshold (for example, 5 frames), it indicates that the video to be processed is suitable as the speech-to-text training corpus. The terminal can determine the video to be processed as the target video and determine the multiple video key frames of the video to be processed as multiple target video key frames.
[0074] In some feasible embodiments, the terminal can also perform character recognition on the video key frames of the video to be processed, and then determine the Chinese character appearance rate of the video key frames of the video to be processed. When the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, it indicates that the video to be processed includes Chinese and is suitable as a speech-to-text training corpus. The terminal can determine the video to be processed as the target video and determine multiple video key frames of the video to be processed as multiple target video key frames. After obtaining the target video key frames, the terminal can perform character recognition on the target video key frames (for example, use OCR technology to recognize the characters in the target video key frames) to obtain the text positions and text contents in each target video key frame.
[0075] S102: Determine the subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame.
[0076] In some feasible embodiments, the position where subtitles appear in a video is not fixed. To accurately recognize the subtitles of the video and obtain a corpus suitable for speech-to-text training, it is necessary to determine the subtitle recognition interval of the video. For details, please refer to Figure 5 , Figure 5 which is a schematic diagram of the scenario of the video subtitle interval provided by the embodiments of the present application. As Figure 5 shown, 200a is a video key frame without subtitles, 200b is a video key frame with subtitles in the middle of the picture, 200c is a video key frame with subtitles at the top of the picture, and 200d is a video key frame with changing subtitle sizes and vertical arrangements. The ideal video subtitle interval is as Figure 5 shown by 200e in, that is, subtitles appear at relatively fixed positions in each video key frame (for example, subtitles in a movie or TV drama video). Therefore, in order to better recognize the subtitles in the video key frames, the subtitle recognition interval can be an area where the coordinates are fixed at the same position but the text content appearing at this position changes continuously. In other words, in the embodiments of the present application, when recognizing the characters in each video key frame of the target video through this subtitle recognition interval, the text contents displayed at the same position in the subtitle recognition interval of different video key frames of the target video are different, that is, the text contents corresponding to the same position in the subtitle recognition interval of different video key frames are different.
[0077] In some feasible embodiments, the target video may also include some non-caption texts, such as signs, watermarks, etc. However, the text positions and text contents of these texts in the video usually remain unchanged or change less. The terminal can utilize this feature to exclude the text positions irrelevant to the captions, further improving the accuracy of subsequent caption extraction. That is, after determining the text positions and text contents in each key frame of the target video, the terminal can screen the text positions, and then determine the candidate text recognition intervals of the target video.
[0078] In some feasible embodiments, the terminal can determine at least one candidate text recognition interval from the text positions of each key frame of the target video, where the number of times the text content in the candidate text recognition interval appears repeatedly is less than the number threshold, thereby excluding the text positions irrelevant to the captions.
[0079] Specifically, in some feasible embodiments, the terminal can determine the text similarity of the text contents appearing in the text positions of each key frame of the target video. When the text similarity of any two appearances of the text content in any text position is greater than the threshold (for example, 85%), the text contents of any two appearances are determined as the repeatedly appearing text contents, and it is determined that the above-mentioned any text position is not used as a candidate text recognition interval.
[0080] In some feasible embodiments, after excluding the text positions according to the text contents in each target key frame, the terminal can determine the caption recognition interval of the target video. For details, please refer to Figure 6 , Figure 6 which is a schematic diagram of the scenario for determining the caption recognition interval provided by the embodiments of the present application. As Figure 6 shown, the terminal can obtain the candidate text recognition interval according to the coordinates of the text position, and then obtain the caption recognition interval. For example, the terminal can represent the caption recognition interval by the horizontal and vertical coordinates of the boundary pixel points of the caption recognition interval, or only represent the caption recognition interval by the horizontal or vertical coordinate of the boundary pixel points of the caption recognition interval, or represent the caption recognition interval by the horizontal and vertical coordinates of the diagonal pixel points of the caption recognition interval boundary, etc.
[0081] In some feasible embodiments, before determining the caption recognition interval of the target video, in order to further improve the accuracy of subsequent caption extraction, the terminal can exclude some text positions with large jitter in the video. That is, the terminal can determine the jitter degree of each candidate text recognition interval according to at least one candidate text recognition interval, and determine the caption recognition interval of the target video according to the candidate text recognition interval whose jitter degree is less than or equal to the jitter degree threshold (for example, the vertical coordinate jitter of the upper and lower boundaries of the text position does not exceed 6 pixels).
[0082] Specifically, if the jitter degree of the text position in the target video key frame is too large (exceeding the jitter degree threshold), that is, the subtitle quality in the target video is poor and not suitable as the training corpus for speech-to-text conversion. The terminal can also discard the existing target video, re-determine the target video from the to-be-processed videos, and re-execute the foregoing operations until a target subtitle recognition interval suitable as the training corpus for speech-to-text conversion is obtained.
[0083] S103: Recognize the subtitles of the target video according to the subtitle recognition interval, obtain the to-be-processed subtitles from the target video, and perform character processing on the to-be-processed subtitles according to the preset corpus acquisition rule to obtain the target subtitles of the target video.
[0084] In some feasible implementation manners, as Figure 4 shown in part b of, after obtaining the subtitle recognition interval, the terminal can recognize the subtitles of the target video according to the subtitle recognition interval (for example, OCR) to obtain the to-be-processed subtitles from the target video, and perform post-processing on the to-be-processed subtitles. The post-processing of subtitles includes, but is not limited to, duplicate removal processing, short sentence merging, word count recognition, character and screening processing operations, etc. The post-processing process of the to-be-processed subtitles will be illustrated by examples in combination with Figure 7 below.
[0085] For details, please refer to Figure 7 , Figure 7 which is a schematic flowchart of the post-processing of subtitles provided by the embodiments of the present application. As Figure 7 shown, the method for post-processing subtitles provided by the embodiments of the present application includes the steps:
[0086] S201: Divide the to-be-processed subtitles into multiple subtitle clauses according to a preset time interval, perform duplicate removal on each subtitle clause, and merge the subtitle clauses with a character length less than the character length threshold in the de-duplicated subtitle clauses to obtain the merged to-be-processed subtitles.
[0087] In some feasible implementation manners, when the terminal recognizes the target video to obtain the to-be-processed subtitles, it can obtain the time axis information corresponding to the to-be-processed subtitles in the target video. For example, the first subtitle appears in the target video from the nth second to the mth second (or from the jth frame to the kth frame). The terminal can divide the to-be-processed subtitles into multiple subtitle clauses according to a preset time interval (for example, 0.2 seconds), perform duplicate removal on each subtitle clause to avoid repeated subtitles from appearing continuously for multiple times, and then merge the subtitle clauses with a character length less than the character length threshold (for example, 2 characters) in the de-duplicated subtitle clauses to avoid overly short subtitle clauses, so as to obtain the merged to-be-processed subtitles.
[0088] In particular, if the length of the subtitle clauses of the subtitle to be processed is too long or too short (for example, exceeding 4 to 25 characters), that is, the speech rate in the target video is too fast or too slow (for example, more than 6 characters per second or less than 3 characters per second), it indicates that the target video is not suitable as the training corpus for speech-to-text conversion. The terminal can also discard the existing target video, re-determine the target video from the videos to be processed, and re-perform the foregoing operations until a subtitle to be processed suitable as the training corpus for speech-to-text conversion is obtained.
[0089] S202: Remove the subtitle clauses containing rare characters from the combined subtitle to be processed, so as to filter out the subtitle clauses that do not contain rare characters, where the rare characters include at least one of letters, numbers, and rare radicals.
[0090] In some feasible embodiments, due to certain technical defects in the text recognition technology itself, some rare characters may not be accurately recognized. These rare characters usually have a large error from the text in the target video. If these rare characters are used as the target subtitles to generate the training corpus for speech-to-text conversion, it will affect the subsequent training effect. Therefore, the terminal can remove the subtitle clauses containing rare characters from the combined subtitle to be processed, so as to filter out the subtitle clauses that do not contain rare characters. Among them, the rare characters include at least one of letters, numbers, and rare radicals.
[0091] S104: Generate a training corpus for speech recognition for speech-to-text conversion according to the target video and the target subtitles.
[0092] In some feasible embodiments, after obtaining the target subtitles, the terminal can generate a training corpus for speech-to-text conversion according to the target video and the target subtitles, or extract the audio in the target video and generate a training corpus for speech-to-text conversion according to the extracted audio and the target subtitles.
[0093] In the embodiments of the present application, by obtaining multiple target video key frames of a target video, and further determining the text positions and text contents from each target video key frame, the subtitle recognition intervals of the target video can be determined according to the text positions and text contents in each target video key frame. It can be understood that the text contents corresponding to the same position in the subtitle recognition intervals are different for different target video key frames. Performing text recognition on the subtitles of the target video according to the subtitle recognition intervals, the subtitles to be processed can be obtained from the target video. After post-processing such as dividing, de-duplicating, and merging the subtitles to be processed according to the corpus acquisition rules, further, the subtitle sentences containing rare characters in the merged subtitles to be processed can be removed to obtain the target subtitles, so as to reduce the error rate of the target subtitles to within the standard. Furthermore, the audio-to-text training corpus for video speech recognition can be generated according to the target video and the target subtitles. Thus, it is possible to automatically obtain videos, extract and screen the subtitles of the videos, and then obtain the training corpus suitable for audio-to-text, improve the acquisition efficiency of the audio-to-text training corpus, and at the same time improve the corpus quality of the audio-to-text training corpus.
[0094] Please refer to Figure 8 , Figure 8 which is another schematic flowchart of the method for obtaining the audio-to-text training corpus provided by the embodiments of the present application. As Figure 8 shown, another method for obtaining the audio-to-text training corpus provided by the present application includes:
[0095] S301: Obtain the video to be processed, and determine the video key frames of the video to be processed, the number of video key frames, and the Chinese character appearance rate in the video key frames.
[0096] S302: When the number of video key frames is greater than or equal to the frame number threshold, and the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, determine the video to be processed as the target video, and determine the multiple video key frames of the video to be processed as multiple target video key frames.
[0097] In some feasible embodiments, in order to ensure the quality of the subsequent audio-to-text training corpus, the terminal can batch obtain the videos to be processed, extract the key frames of the videos to be processed, and screen through the video key frames to obtain the target video and the target video key frames suitable for use as the audio-to-text training corpus. As shown in part a of Figure 4 , during the process of video acquisition, the terminal can batch download the videos to be processed, and perform acquisition management on the batch-obtained videos to be processed (for example, classify according to video types, classify according to source applications, etc.), and then perform video preprocessing on the videos (for example, image grayscale processing, image stretching (such as stretching the video image to a fixed pixel size), picture sharpening, etc.), which is convenient for subsequent positioning and subtitle extraction of the subtitle intervals of the videos.
[0098] In some feasible embodiments, after the terminal obtains the video to be processed, it can perform key frame extraction on the video to be processed. The key frame extraction methods include, but are not limited to, key frame extraction methods based on shots, key frame extraction methods based on motion analysis, and key frame extraction methods based on video clustering. Among them, the key frame extraction method based on shots is the first developed and currently the most mature general method in the field of video processing. The general implementation process of this method is: first, the video file is segmented according to shot changes, and then the first and last frames are selected as key frames in each shot of the video. The advantage of this method is that it is very simple to implement and has a small amount of calculation. However, this method has great limitations. When the content in the video changes violently and the scene is very complex, selecting the first and last frames in the shot does not represent all the content changes in the video. Therefore, this method can no longer meet the standards and requirements for key frame extraction in today's society. The key frame extraction method based on motion analysis is a method for extracting key frames based on the attributes of object motion characteristics. The implementation process of this method is: analyze the optical flow of object motion in the video shot, and each time select the video frame with the least number of optical flow movements in the video shot as the extracted key frame. This method can extract an appropriate amount of key frames from most video shots, and the extracted key frames can also effectively express the motion characteristics of the video. However, this method has the characteristic of poor robustness because it not only depends on the local characteristics of object motion, but also the calculation process is relatively complex, and the algorithm has a large time overhead cost. The key frame extraction method based on video clustering can divide video frames into several clusters through clustering during the process of extracting key frames, and then select corresponding frames as key frames in each cluster. The video key frames extracted using this algorithm not only have a small redundancy, but also the key frames can accurately reflect all the content that occurs in the video. However, the method based on clustering does not fully consider the chronological order of changes between frames during the process of dividing clustering clusters, and a certain number of clusters need to be preset before clustering. Therefore, the applicability of this method is limited to a certain extent. In the specific implementation manner, there are many methods for extracting key frames from a video. The key frames of the video can be extracted by using a specific key frame extraction method according to the characteristics of the video itself, or the key frames of the video can be extracted by combining several key frame extraction methods. It can be understood that any process of extracting key frames from a video is covered by the protection scope of this application.
[0099] In some feasible embodiments, the terminal can determine the number of video key frames of the video to be processed. When the number of video key frames is greater than or equal to a frame number threshold (for example, 5 frames), it indicates that the video to be processed is suitable as audio-to-text training corpus. The terminal can determine the video to be processed as the target video, and determine multiple video key frames of the video to be processed as multiple target video key frames. After obtaining the target video key frames, the terminal can perform text recognition (such as OCR) on the target video key frames to obtain the text positions and text contents in each target video key frame.
[0100] In some feasible embodiments, before determining the video to be processed as the target video, the terminal can perform text recognition on the video key frames of the video to be processed, and then determine the Chinese character appearance rate in the video key frames of the video to be processed. When the Chinese character appearance rate in any video key frame of the video to be processed is greater than or equal to an appearance rate threshold, it indicates that the video to be processed includes Chinese and is suitable as audio-to-text training corpus. The terminal can determine the video to be processed as the target video, and determine multiple video key frames of the video to be processed as multiple target video key frames.
[0101] S303: Determine at least one text position from the text positions in each target video key frame as at least one candidate text recognition interval, where the number of times the text content in the candidate text recognition interval appears repeatedly is less than a number threshold.
[0102] In some feasible embodiments, the position where subtitles appear in the video is not fixed. To accurately recognize the subtitles of the video and obtain a corpus suitable for audio-to-text training, it is necessary to determine the subtitle recognition interval of the video. As Figure 5 shown, 200a is a video key frame without subtitles, 200b is a video key frame with subtitles in the middle of the screen, 200c is a video key frame with subtitles at the top of the screen, and 200d is a video key frame with changing subtitle sizes and vertical arrangements. The ideal video subtitle interval is as Figure 5 shown in 200e, where subtitles appear at relatively fixed positions in each video key frame (for example, subtitles in a movie or TV drama video).
[0103] In some feasible embodiments, the target video may also include some non-subtitle texts, such as texts in logos, watermarks, etc. However, the text positions and text contents of these texts in the video usually remain unchanged or change less. The terminal can utilize this feature to exclude text positions irrelevant to subtitles, further improving the accuracy of subsequent subtitle extraction. That is, after obtaining the text positions and text contents in each target video key frame, the terminal can screen the text positions to determine the candidate text recognition interval of the target video.
[0104] In some feasible embodiments, the terminal can determine the text similarity of the text content appearing in the text positions of each target video key frame. When the text similarity between any two secondary occurrences of the text content in any text position is greater than a threshold (e.g., 85%), the text content of any two occurrences is determined as the repeatedly occurring text content, and it is determined that the above-mentioned any text position is not used as a candidate text recognition interval.
[0105] In some feasible embodiments, after removing the text positions according to the text content in each target key frame, the terminal can determine the caption recognition interval of the target video. As Figure 6 shown, the terminal can obtain the candidate text recognition interval based on the coordinates of the text position, and then obtain the caption recognition interval. For example, the terminal can use the horizontal and vertical coordinates of the boundary pixel points of the caption recognition interval to represent the caption recognition interval, or only use the horizontal or vertical coordinate of the boundary pixel points of the caption recognition interval to represent the caption recognition interval, or use the horizontal and vertical coordinates of the diagonal pixel points of the caption recognition interval boundary and other ways to represent the caption recognition interval.
[0106] S304: Determine the jitter degree of each candidate text recognition interval, and determine the candidate text recognition interval with the jitter degree less than or equal to the jitter degree threshold as the caption recognition interval of the target video.
[0107] In some feasible embodiments, before determining the caption recognition interval of the target video, in order to further improve the accuracy of subsequent caption extraction, the terminal can remove some text positions with large jitter degrees in the video. That is, the terminal can determine the jitter degree of each candidate text recognition interval according to at least one candidate text recognition interval, and determine the caption recognition interval of the target video according to the candidate text recognition interval with the jitter degree less than or equal to the jitter degree threshold (e.g., the vertical coordinate jitter of the upper and lower boundaries of the text position does not exceed 6 pixels).
[0108] In particular, if the jitter degree of the text position in the target video key frame is too large (exceeding the jitter degree threshold), that is, the caption quality in the target video is poor and not suitable as the training corpus for speech-to-text conversion. The terminal can also discard the existing target video, re-determine the target video from the videos to be processed, and re-execute the foregoing operations until a target caption recognition interval suitable as the training corpus for speech-to-text conversion is obtained.
[0109] S305: Recognize the captions of the target video according to the caption recognition interval to obtain the captions to be processed from the target video.
[0110] In some feasible embodiments, as Figure 4As shown in part b of FIG. 1 , after obtaining the subtitle recognition interval, the terminal can recognize the subtitles of the target video according to the subtitle recognition interval (for example, OCR) to obtain the subtitles to be processed from the target video, and perform subtitle post-processing on the subtitles to be processed. The subtitle post-processing includes but is not limited to de-duplication processing, short sentence merging, word count recognition, and character and screening processing operations.
[0111] S306: Perform character processing on the subtitles to be processed according to a preset corpus acquisition rule to obtain target subtitles for the target video.
[0112] In some feasible implementations, the terminal can obtain the timeline information corresponding to the subtitles to be processed in the target video while identifying the target video to obtain the subtitles to be processed, for example, the first subtitle appears in the target video from the nth second to the mth second (or the jth frame to the kth frame). The terminal can divide the subtitles to be processed into multiple subtitle sentences according to a preset time interval (for example, 0.2 seconds), and remove duplicates from each subtitle sentence to avoid repeated subtitles appearing multiple times in a row, and then merge the subtitle sentences whose character length is less than the character length threshold (for example, 2 characters) in the subtitle sentences after deduplication to avoid the subtitle sentences being too short, so as to obtain the merged subtitles to be processed.
[0113] In particular, if the length of the subtitle sentence of the subtitle to be processed is long or short (for example, more than 4 to 25 characters), that is, the speech speed in the target video is fast or slow (for example, more than 6 words per second or less than 3 words per second), it means that the target video is not suitable as the audio-to-text training corpus. The terminal can also discard the existing target video, re-determine the target video in the video to be processed, and re-perform the above operation until the subtitle to be processed suitable as the audio-to-text training corpus is obtained.
[0114] In some feasible implementations, due to certain technical defects of the text recognition technology itself, some uncommon characters may not be accurately recognized. These uncommon characters usually have large errors with the text in the target video. If these uncommon characters are used as target subtitles to generate audio-to-text training corpus, it will affect the subsequent training effect. Therefore, the terminal can remove the subtitle sentences containing uncommon characters from the merged subtitles to be processed to filter out the subtitle sentences that do not contain uncommon characters. Among them, uncommon characters include at least one of letters, numbers, and uncommon radicals.
[0115] S307: Generate audio-to-text training corpus for video speech recognition based on the target video and the target subtitles.
[0116] In some feasible embodiments, after obtaining the target subtitle, the terminal may generate a speech-to-text training corpus based on the target video and the target subtitle, or may extract the audio in the target video and generate a speech-to-text training corpus based on the extracted audio and the target subtitle.
[0117] In some feasible embodiments, after obtaining the target subtitle, the terminal may verify the error rate of the target subtitle. When the error rate of the target subtitle is less than the error threshold (5%), a speech-to-text training corpus is generated based on the target video and the target subtitle. Among them, the error rate can be calculated by Formula 1, and Formula 1 is specifically as follows:
[0118]
[0119] Among them, CER is the error rate, S is the number of replaced characters, D is the number of deleted characters, I is the number of inserted characters, and N is the total number of characters.
[0120] In the embodiments of the present application, by obtaining multiple target video key frames of the target video, further, the text position and text content are determined from each target video key frame, so that the subtitle recognition interval of the target video can be determined according to the text position and text content in each target video key frame. It can be understood that the text content corresponding to the same position in the subtitle recognition interval is different for different target video key frames. Performing text recognition on the subtitle of the target video according to the subtitle recognition interval, the subtitle to be processed can be obtained from the target video. After post-processing such as partitioning, deduplication, and merging of the subtitle to be processed according to the corpus acquisition rule, further, the subtitle sentences containing rare characters in the merged subtitle to be processed are removed to obtain the target subtitle, so that the error rate of the target subtitle is reduced to within the standard, and then a speech-to-text training corpus for video speech recognition is generated based on the target video and the target subtitle. Thus, it is possible to automatically obtain a video, extract and screen the subtitle of the video, and then obtain a training corpus suitable for speech-to-text, improve the acquisition efficiency of the speech-to-text training corpus, and at the same time improve the corpus quality of the speech-to-text training corpus.
[0121] Further, please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the speech-to-text training corpus acquisition device provided by the embodiments of the present application. As Figure 9 shown, the above device may include:
[0122] A video acquisition module 601, configured to acquire multiple target video key frames of the target video, and determine the text position and text content from each target video key frame.
[0123] In some feasible embodiments, in order to ensure the quality of the subsequent speech-to-text training corpus, the video acquisition module 601 may batch acquire videos to be processed, extract key frames from the videos to be processed, and screen through the video key frames to obtain target videos suitable as speech-to-text training corpus and target video key frames. As Figure 4 shown in part a of Figure 4 , during the video acquisition process, the video acquisition module 601 may batch download the videos to be processed and perform acquisition management on the batch-acquired videos to be processed (for example, classify according to video type, classify according to source application, etc.), and then perform video preprocessing on the videos (for example, image grayscale processing, image stretching (such as stretching the video image to a fixed pixel size), picture sharpening, etc.) to facilitate subsequent positioning and subtitle extraction of the subtitle intervals of the videos.
[0124] In some feasible embodiments, after the video acquisition module 601 acquires the video to be processed, it can extract key frames from the video to be processed. The methods for extracting key frames include, but are not limited to: key frame extraction methods based on shots, key frame extraction methods based on motion analysis, and key frame extraction methods based on video clustering. Among them, the key frame extraction method based on shots is the first developed and currently the most mature general method in the field of video processing. The general implementation process of this method is: first, the video file is segmented according to shot changes, and then the first and last frames in each shot of the video are selected as key frames. The advantage of this method is that it is very simple to implement and has a small amount of calculation. However, this method has great limitations. When the content in the video changes violently and the scene is very complex, selecting the first and last frames in the shot does not represent all the content changes in the video. Therefore, this method can no longer meet the standards and requirements of people for key frame extraction in today's society. The key frame extraction method based on motion analysis is a method for extracting key frames based on the attributes of object motion characteristics. The implementation process of this method is: analyze the optical flow of object motion in the video shot, and each time select the video frame with the least number of optical flow movements in the video shot as the extracted key frame. This method can extract an appropriate number of key frames from most video shots, and the extracted key frames can also effectively express the motion characteristics of the video. However, this method has the characteristic of poor robustness because it not only depends on the local characteristics of object motion, but also the calculation process is relatively complex, and the algorithm has a large time cost. The key frame extraction method based on video clustering can divide video frames into several clusters through clustering during the process of extracting key frames, and then select corresponding frames in each cluster as key frames. The video key frames extracted by using this algorithm not only have a small redundancy, but also the key frames can accurately reflect all the content that occurs in the video. However, the method based on clustering does not fully consider the sequential change order of time between frames during the process of dividing clustering clusters, and a certain number of clusters need to be set in advance before clustering. Therefore, the applicability of this method is limited to a certain extent. In the specific implementation manner, there are many methods for extracting key frames from a video. The key frames of the video can be extracted by using a specific key frame extraction method according to the characteristics of the video itself, or the key frames of the video can be extracted by combining several key frame extraction methods. It can be understood that any process of extracting key frames from a video is covered by the protection scope of this application.
[0125] In some feasible embodiments, the video acquisition module 601 includes:
[0126] The frame number determination unit 6011 is configured to obtain a video to be processed, determine the video key frames of the video to be processed and the number of frames of the video key frames. When the number of frames of the video key frames is greater than or equal to the frame number threshold, the video to be processed is determined as the target video, and multiple video key frames of the video to be processed are determined as multiple target video key frames.
[0127] The Chinese character determination unit 6012 is configured to obtain a video to be processed, determine the video key frames of the video to be processed and the Chinese character appearance rate of the video key frames. When the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, the video to be processed is determined as the target video, and multiple video key frames of the video to be processed are determined as multiple target video key frames.
[0128] In some feasible embodiments, the frame number determination unit 6011 can determine the number of frames of the video key frames of the video to be processed. When the number of frames of the video key frames is greater than or equal to the frame number threshold (for example, 5 frames), it indicates that the video to be processed is suitable as the speech-to-text training corpus. The frame number determination unit 6011 can determine the video to be processed as the target video, and multiple video key frames of the video to be processed are determined as multiple target video key frames.
[0129] In some feasible embodiments, the Chinese character determination unit 6012 can perform character recognition on the video key frames of the video to be processed, and then determine the Chinese character appearance rate of the video key frames of the video to be processed. When the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, it indicates that the video to be processed includes Chinese and is suitable as the speech-to-text training corpus. The Chinese character determination unit 6012 can determine the video to be processed as the target video, and multiple video key frames of the video to be processed are determined as multiple target video key frames.
[0130] In some feasible embodiments, after obtaining the target video key frames, the video acquisition module 601 can perform character recognition (such as OCR) on the target video key frames to obtain the text positions and text contents in each target video key frame.
[0131] The interval division module 602 is configured to determine the subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame, wherein different target video key frames correspond to different text contents at the same position in the subtitle recognition interval.
[0132] In some feasible embodiments, the interval division module 602 includes:
[0133] The interval deduplication unit 6021 is used to determine at least one text position from the text positions of each target video key frame as at least one candidate text recognition interval, and the number of times the text content in the candidate text recognition interval appears repeatedly is less than the number threshold.
[0134] In some feasible implementation manners, the position where subtitles appear in a video is not fixed. To accurately recognize the subtitles of the video and thus obtain a corpus suitable for speech-to-text training, it is necessary to determine the subtitle recognition interval of the video. For example, Figure 5 as shown, 200a is a video key frame without subtitles, 200b is a video key frame with subtitles in the middle of the screen, 200c is a video key frame with subtitles at the top of the screen, and 200d is a video key frame with subtitles of changing sizes and vertical arrangements. And the ideal video subtitle interval is as Figure 5 shown in 200e, where subtitles appear at relatively fixed positions in each video key frame (for example, subtitles in a movie or TV drama video).
[0135] In some feasible implementation manners, the target video may also include some non-subtitle texts, such as texts of logos, watermarks, etc. However, the text positions and text contents of these texts in the video usually remain unchanged or change less. The interval deduplication unit 6021 can utilize this feature to exclude text positions irrelevant to subtitles, further improving the accuracy of subsequent subtitle extraction. That is, after obtaining the text positions and text contents in each target video key frame, the interval deduplication unit 6021 can screen the text positions and then determine the candidate text intervals of the target video.
[0136] In some feasible implementation manners, the interval division module 602 includes:
[0137] The text recognition module 6023 is used to determine the text similarity of the text contents appearing in the text positions of each target video key frame. When the text similarity of any two occurrences of text contents in any text position is greater than a threshold (for example, 85%), the text contents of any two occurrences are determined as repeatedly appearing text contents, and it is determined that the above-mentioned any text position is not used as a candidate text recognition interval.
[0138] In some feasible implementation manners, the text recognition module 6023 can determine the text similarity of the text contents appearing in the text positions of each target video key frame. When the text similarity of any two occurrences of text contents in any text position is greater than a threshold (for example, 85%), the text contents of any two occurrences are determined as repeatedly appearing text contents, and it is determined that the above-mentioned any text position is not used as a candidate text recognition interval.
[0139] In some feasible implementation manners, the interval division module 602 includes:
[0140] An interval determination unit 6022, configured to determine the jitter degree of each candidate text recognition interval, and determine the subtitle recognition interval of the target video as the candidate text recognition interval whose jitter degree is less than or equal to the jitter degree threshold.
[0141] In some feasible implementation manners, after excluding the text positions according to the text contents in each target key frame, the interval determination unit 6022 may determine the subtitle recognition interval of the target video. As Figure 6 shown, the interval determination unit 6022 may obtain the candidate text recognition interval according to the coordinates of the text position, and further obtain the subtitle recognition interval. For example, the interval determination unit 6022 may represent the subtitle recognition interval by using the horizontal and vertical coordinates of the boundary pixel points of the subtitle recognition interval, or may only use the horizontal coordinate or the vertical coordinate of the boundary pixel points of the subtitle recognition interval to represent the subtitle recognition interval, or may also use the horizontal and vertical coordinates of the diagonal pixel points of the boundary of the subtitle recognition interval and other ways to represent the subtitle recognition interval.
[0142] In some feasible implementation manners, before determining the subtitle recognition interval of the target video, in order to further improve the accuracy of subsequent subtitle extraction, the interval determination unit 6022 may exclude some text positions with large jitter degrees in the video. That is, the interval determination unit 6022 may determine the jitter degree of each candidate text recognition interval according to at least one candidate text recognition interval, and determine the subtitle recognition interval of the target video as the candidate text recognition interval whose jitter degree is less than or equal to the jitter degree threshold (for example, the vertical coordinate jitter of the upper and lower boundaries of the text position does not exceed 6 pixels).
[0143] Particularly, if the jitter degree of the text position in the key frame of the target video is too large (exceeding the jitter degree threshold), that is, the subtitle quality in the target video is poor and not suitable as the training corpus for speech-to-text conversion. The interval determination unit 6022 may also discard the existing target video, re-determine the target video in the to-be-processed video through the video acquisition module 601, and re-execute the foregoing operations until a target subtitle recognition interval suitable as the training corpus for speech-to-text conversion is obtained.
[0144] A subtitle extraction module 603, configured to recognize the subtitle of the target video according to the subtitle recognition interval, so as to obtain the to-be-processed subtitle from the target video, and perform character processing on the to-be-processed subtitle according to the preset corpus acquisition rule to obtain the target subtitle of the target video.
[0145] In some feasible implementation manners, as Figure 4As shown in part b), after obtaining the subtitle recognition interval, the subtitle extraction module 603 can recognize the subtitles of the target video according to the subtitle recognition interval (e.g., OCR) to obtain the subtitles to be processed from the target video, and perform post-processing on the subtitles to be processed. Among them, the post-processing of subtitles includes, but is not limited to, duplicate removal, short sentence merging, word count recognition, character and screening, and other processing operations.
[0146] In some feasible implementation manners, the subtitle extraction module 603 includes:
[0147] A subtitle division unit 6031, configured to divide the subtitles to be processed into multiple subtitle clauses according to a preset time interval, perform duplicate removal on each subtitle clause, and merge the subtitle clauses with a character length less than the character length threshold in the duplicate-removed subtitle clauses to obtain the merged subtitles to be processed.
[0148] In some feasible implementation manners, when the subtitle division unit 6031 recognizes the target video to obtain the subtitles to be processed, it can obtain the corresponding time axis information of the subtitles to be processed in the target video. For example, the first subtitle appears in the target video from the nth second to the mth second (or from the jth frame to the kth frame). The subtitle division unit 6031 can divide the subtitles to be processed into multiple subtitle clauses according to a preset time interval (e.g., 0.2 seconds), perform duplicate removal on each subtitle clause to avoid repeated subtitles from appearing continuously for multiple times, and then merge the subtitle clauses with a character length less than the character length threshold (e.g., 2 characters) in the duplicate-removed subtitle clauses to avoid overly short subtitle clauses, so as to obtain the merged subtitles to be processed.
[0149] In some feasible implementation manners, the subtitle extraction module 603 includes:
[0150] A subtitle screening unit 6032, configured to screen subtitle clauses based on the characters in the merged subtitles to be processed to determine the target subtitles of the target video.
[0151] In some feasible implementation manners, if the length of the subtitle clause of the subtitles to be processed is too long or too short (e.g., exceeding 4 to 25 characters), that is, the speech speed in the target video is too fast or too slow (e.g., more than 6 characters per second or less than 3 characters per second), it indicates that the target video is not suitable as a speech-to-text training corpus. The subtitle screening unit 6032 can also discard the existing target video, re-determine the target video in the videos to be processed through the video acquisition module 601 and the interval division module 602, and re-execute the foregoing operations until the subtitles to be processed suitable as a speech-to-text training corpus are obtained.
[0152] In some feasible embodiments, the subtitle screening unit 6032 includes a clause elimination subunit. The clause elimination subunit is used to eliminate subtitle clauses containing rare characters from the merged subtitles to be processed, so as to screen out subtitle clauses that do not contain rare characters. Among them, rare characters include at least one of letters, numbers, and rare radical components.
[0153] In some feasible embodiments, due to certain technical defects in the text recognition technology itself, some rare characters may not be accurately recognized. There are usually large errors between this part of rare characters and the text in the target video. If this part of rare characters is used as the target subtitle to generate the speech-to-text training corpus, it will affect the subsequent training effect. Therefore, the subtitle screening unit 6032 can eliminate subtitle clauses containing rare characters from the merged subtitles to be processed, so as to screen out subtitle clauses that do not contain rare characters. Among them, rare characters include at least one of letters, numbers, and rare radical components.
[0154] The corpus generation module 604 is used to generate a speech-to-text training corpus for video speech recognition according to the target video and the target subtitle.
[0155] In some feasible embodiments, after obtaining the target subtitle, the corpus generation module 604 can generate a speech-to-text training corpus according to the target video and the target subtitle, or extract the audio in the target video and generate a speech-to-text training corpus according to the extracted audio and the target subtitle.
[0156] In some feasible embodiments, after obtaining the target subtitle, the corpus generation module 604 can verify the error rate of the target subtitle. When the error rate of the target subtitle is less than the error threshold (5%), a speech-to-text training corpus is generated according to the target video and the target subtitle. Among them, the error rate can be calculated by Formula 1. The specific Formula 1 is as follows:
[0157]
[0158] Among them, CER is the error rate, S is the number of replaced characters, D is the number of deleted characters, I is the number of inserted characters, and N is the total number of characters.
[0159] In the embodiments of the present application, by obtaining multiple target video key frames of a target video, and further determining the text positions and text contents from each target video key frame, the subtitle recognition intervals of the target video can be determined according to the text positions and text contents in each target video key frame. It can be understood that the text contents corresponding to the same position in the subtitle recognition intervals are different for different target video key frames. Performing text recognition on the subtitles of the target video according to the subtitle recognition intervals, the subtitles to be processed can be obtained from the target video. After performing post-processing such as partitioning, deduplication, and merging on the subtitles to be processed according to the corpus acquisition rules, further, the subtitle sentences containing rare characters in the merged subtitles to be processed can be removed to obtain the target subtitles, so as to reduce the error rate of the target subtitles to within the standard. Furthermore, an audio-to-text training corpus for video speech recognition can be generated according to the target video and the target subtitles. Thus, it is possible to automatically obtain a video, extract and screen the subtitles of the video, and then obtain a training corpus suitable for audio-to-text, improving the acquisition efficiency of the audio-to-text training corpus and at the same time improving the corpus quality of the audio-to-text training corpus.
[0160] See Figure 10 , Figure 10 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. As Figure 10 shown, the terminal device 1000 in this embodiment may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above terminal device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1004 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally be at least one storage device located far from the aforementioned processor 1001. As Figure 10 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0161] In Figure 10 the terminal device 1000 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0162] Obtain multiple target video key frames of the target video, and determine the text positions and text contents from each target video key frame;
[0163] Determine the subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame, where different target video key frames correspond to different text contents at the same position in the subtitle recognition interval;
[0164] Recognize the subtitles of the target video according to the subtitle recognition interval, so as to obtain the subtitles to be processed from the target video, and perform character processing on the subtitles to be processed according to the preset corpus acquisition rule to obtain the target subtitles of the target video;
[0165] Generate the speech-to-text training corpus for video speech recognition according to the target video and the target subtitles.
[0166] In some feasible implementation manners, the above-mentioned processor 1001 is used for:
[0167] Obtain the video to be processed, and determine the video key frames of the video to be processed and the number of frames of the video key frames;
[0168] When the number of frames of the video key frames is greater than or equal to the frame number threshold, determine the video to be processed as the target video, and determine the multiple video key frames of the video to be processed as multiple target video key frames.
[0169] In some feasible implementation manners, the above-mentioned processor 1001 is used for:
[0170] Obtain the video to be processed, and determine the video key frames of the video to be processed and the Chinese character appearance rate of the video key frames;
[0171] When the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, determine the video to be processed as the target video, and determine the multiple video key frames of the video to be processed as multiple target video key frames.
[0172] In some feasible implementation manners, the above-mentioned processor 1001 is used for:
[0173] Determine at least one text position from the text positions of each target video key frame as at least one candidate text recognition interval, and the number of times the text content in the candidate text recognition interval appears repeatedly is less than the number threshold;
[0174] Determine the jitter degree of each candidate text recognition interval, and determine the candidate text recognition interval with the jitter degree less than or equal to the jitter degree threshold as the subtitle recognition interval of the target video.
[0175] In some feasible embodiments, the above-mentioned processor 1001 is further configured to:
[0176] Determine the text similarity of the text content that appears in the text positions of each target video key frame;
[0177] When the text similarity of any two secondary occurrences of the text content in any text position is greater than the threshold, determine the text content of any two occurrences as the repeatedly-occurring text content, and determine that the above-mentioned any text position is not used as a candidate text recognition interval.
[0178] In some feasible embodiments, the above-mentioned processor 1001 is configured to:
[0179] Divide the subtitle to be processed into multiple subtitle clauses according to a preset time interval, remove duplicates from each subtitle clause, and merge the subtitle clauses with a character length less than the character length threshold in the de-duplicated subtitle clauses to obtain the merged subtitle to be processed;
[0180] Based on the characters in the merged subtitle to be processed, perform subtitle clause screening to determine the target subtitle of the target video.
[0181] In some feasible embodiments, the above-mentioned processor 1001 is configured to:
[0182] Remove the subtitle clauses containing rare characters in the merged subtitle to be processed to screen out the subtitle clauses that do not contain rare characters, where the rare characters include at least one of letters, numbers, and rare radical components.
[0183] It should be understood that in some feasible embodiments, the above-mentioned processor 1001 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0184] It should be understood that the device control application program stored in the above-mentioned memory 1005 may include the following functional modules:
[0185] A video acquisition module, configured to acquire multiple target video key frames of a target video, and determine the text positions and text contents from the target video key frames;
[0186] An interval division module, configured to determine a subtitle recognition interval of the target video according to the text positions and text contents in the target video key frames, wherein the text contents at the same position in the subtitle recognition interval corresponding to different target video key frames are different;
[0187] A subtitle extraction module, configured to recognize the subtitles of the target video according to the subtitle recognition interval, so as to obtain subtitles to be processed from the target video, and perform character processing on the subtitles to be processed according to a preset corpus acquisition rule to obtain the target subtitles of the target video;
[0188] A corpus generation module, configured to generate a speech-to-text training corpus for video speech recognition according to the target video and the target subtitles.
[0189] In some feasible embodiments, the above-mentioned video acquisition module includes:
[0190] A frame number determination unit, configured to acquire a video to be processed, determine the video key frames of the video to be processed and the frame numbers of the video key frames, and when the frame number of the video key frame is greater than or equal to a frame number threshold, determine the video to be processed as the target video, and determine the multiple video key frames of the video to be processed as multiple target video key frames.
[0191] In some feasible embodiments, the above-mentioned video acquisition module includes:
[0192] A Chinese character determination unit, configured to acquire a video to be processed, determine the video key frames of the video to be processed and the Chinese character appearance rate of the video key frames, and when the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to an appearance rate threshold, determine the video to be processed as the target video, and determine the multiple video key frames of the video to be processed as multiple target video key frames.
[0193] In some feasible embodiments, the above-mentioned interval division module includes:
[0194] An interval deduplication unit, configured to determine at least one text position from the text positions of the target video key frames as at least one candidate text recognition interval, and the number of times the text content in the candidate text recognition interval appears repeatedly is less than a number threshold;
[0195] An interval determination unit, configured to determine the jitter degree of each candidate text recognition interval, and determine the candidate text recognition interval with the jitter degree less than or equal to a jitter degree threshold as the subtitle recognition interval of the target video.
[0196] In some feasible embodiments, the above interval division module further includes:
[0197] A text recognition module, configured to determine the text similarity of the text content appearing in the text positions of each target video key frame. When the text similarity between any two sub-occurrences of the text content in any text position is greater than a threshold, the text content that appears twice is determined as the repeatedly appearing text content, and it is determined that the above-mentioned any text position is not used as a candidate text recognition interval.
[0198] In some feasible embodiments, the above subtitle extraction module includes:
[0199] A subtitle division unit, configured to divide the subtitle to be processed into multiple subtitle clauses according to a preset time interval, and remove duplicates from each subtitle clause, and merge the subtitle clauses with a character length less than a character length threshold in the de-duplicated subtitle clauses to obtain the subtitle to be processed after merging;
[0200] A subtitle screening unit, configured to screen subtitle clauses based on the characters in the subtitle to be processed after merging to determine the target subtitle of the target video.
[0201] In some feasible embodiments, the above subtitle screening unit includes:
[0202] A clause elimination sub-unit, configured to eliminate the subtitle clauses containing rare characters in the subtitle to be processed after merging, so as to screen out subtitle clauses that do not contain rare characters, where the rare characters include at least one of letters, numbers, and rare radical components.
[0203] In a specific implementation, the above terminal device 1000 can execute the implementation methods provided in each of the above Figure 2 、 Figure 7 and / or Figure 8 by each of its built-in functional modules. For specific details, reference can be made to the implementation methods provided in each of the above steps, which will not be elaborated here.
[0204] In the embodiments of the present application, by obtaining multiple target video key frames of a target video, and further determining the text positions and text contents from each target video key frame, the subtitle recognition intervals of the target video can be determined according to the text positions and text contents in each target video key frame. It can be understood that the text contents corresponding to the same position in the subtitle recognition intervals are different for different target video key frames. By performing text recognition on the subtitles of the target video according to the subtitle recognition intervals, the subtitles to be processed can be obtained from the target video. After post-processing such as dividing, de-duplicating, and merging the subtitles to be processed according to the corpus acquisition rules, further, the subtitle clauses containing rare characters in the merged subtitles to be processed can be removed to obtain the target subtitles, so that the error rate of the target subtitles can be reduced to within the standard. Furthermore, according to the target video and the target subtitles, a speech-to-text training corpus for video speech recognition can be generated. Thus, it is possible to automatically obtain a video, extract and screen the subtitles of the video, and further obtain a training corpus suitable for speech-to-text, improve the acquisition efficiency of the speech-to-text training corpus, and at the same time improve the corpus quality of the speech-to-text training corpus.
[0205] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, which is executed by a processor to implement Figure 2 、 Figure 7 and / or Figure 8 the methods provided by each step in
[0206] For the specific implementation manners, reference may be made to the implementation manners provided by each of the above steps, which will not be elaborated herein.
[0207] In the embodiments of the present application, by obtaining multiple key frames of a target video, and further determining the text positions and text contents from each key frame of the target video, the subtitle recognition intervals of the target video can be determined according to the text positions and text contents in each key frame of the target video. It can be understood that the text contents corresponding to the same position in the subtitle recognition intervals are different for different key frames of the target video. By performing text recognition on the subtitles of the target video according to the subtitle recognition intervals, the subtitles to be processed can be obtained from the target video. After post-processing such as dividing, de-duplicating, and merging the subtitles to be processed according to the corpus acquisition rules, further, the subtitle sentences containing rare characters in the merged subtitles to be processed can be removed to obtain the target subtitles, so that the error rate of the target subtitles can be reduced to within the standard. Furthermore, the audio-to-text training corpus for video speech recognition can be generated according to the target video and the target subtitles. Thus, it is possible to automatically obtain a video, extract and screen the subtitles of the video, and then obtain a training corpus suitable for audio-to-text, improving the acquisition efficiency of the audio-to-text training corpus and at the same time improving the corpus quality of the audio-to-text training corpus.
[0208] In the claims, the description, and the drawings of the present application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. The mention of "embodiment" in this article means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The display of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0209] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0210] The above disclosure is only a preferred embodiment of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for obtaining a speech-to-text training corpus, characterized in that The method includes: Obtaining a video to be processed, and determining the video key frames of the video to be processed, the number of frames of the video key frames, and the Chinese character appearance rate of the video key frames; When the number of frames of the video key frames is greater than or equal to a frame number threshold, and the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to an appearance rate threshold, determining the video to be processed as a target video, determining multiple video key frames of the video to be processed as multiple target video key frames, and determining the text positions and text contents from each target video key frame; Determining a subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame, wherein different target video key frames correspond to different text contents at the same position in the subtitle recognition interval; Recognizing the subtitles of the target video according to the subtitle recognition interval to obtain subtitles to be processed from the target video, and performing character processing on the subtitles to be processed according to a preset corpus acquisition rule to obtain the target subtitles of the target video; Verifying the typo rate of the target subtitles, and when the typo rate of the target subtitles is less than a typo threshold, generating an audio-to-text training corpus for speech recognition according to the target video and the target subtitles; Wherein, the typo rate is calculated by the following formula: Where CER is the typo rate, S is the number of replaced characters, D is the number of deleted characters, I is the number of inserted characters, and N is the total number of characters.
2. The method according to claim 1, characterized in that The determining the subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame includes: Determining at least one text position from the text positions of each target video key frame as at least one candidate text recognition interval, and the number of times the text content of the candidate text recognition interval appears repeatedly is less than a number threshold; Determining the jitter degree of each candidate text recognition interval, and determining the candidate text recognition interval with the jitter degree less than or equal to the jitter degree threshold as the subtitle recognition interval of the target video.
3. The method according to claim 2, wherein The method further includes: Determining the text similarity of the text contents appearing in the text positions of each target video key frame; When the text similarity of the text contents appearing any two times in any text position is greater than a threshold, determining the text contents appearing any two times as repeatedly appearing text contents, and determining that the any text position is not used as a candidate text recognition interval.
4. The method according to claim 3, wherein The performing character processing on the subtitles to be processed according to a preset corpus acquisition rule includes: Dividing the subtitles to be processed into multiple subtitle clauses according to a preset time interval, removing duplicates from each subtitle clause, and merging the subtitle clauses with a character length less than a character length threshold in the deduplicated subtitle clauses to obtain the merged subtitles to be processed; Performing subtitle clause screening based on the characters in the merged subtitles to be processed to determine the target subtitles of the target video.
5. The method according to claim 4, wherein The performing subtitle clause screening based on the characters in the merged subtitles to be processed includes: Removing the subtitle clauses containing rare characters in the merged subtitles to be processed to screen out the subtitle clauses not containing rare characters; Among them, the rare characters at least include one of letters, numbers, and rare radicals.
6. An acquisition device for speech-to-text training corpus, characterized in that, Including: A video acquisition module, configured to acquire a video to be processed, and determine the video key frames of the video to be processed, the number of frames of the video key frames, and the Chinese character appearance rate of the video key frames; When the number of frames of the video key frames is greater than or equal to the frame number threshold, and the Chinese character appearance rate of any video key frame in the video to be processed is greater than or equal to the appearance rate threshold, determine the video to be processed as the target video, determine the multiple video key frames of the video to be processed as multiple target video key frames, and determine the text positions and text contents from each target video key frame; An interval division module, configured to determine the subtitle recognition interval of the target video according to the text positions and text contents in each target video key frame, wherein the text contents corresponding to the same position in the subtitle recognition interval are different for different target video key frames; A subtitle extraction module, configured to recognize the subtitles of the target video according to the subtitle recognition interval, so as to obtain subtitles to be processed from the target video, and perform character processing on the subtitles to be processed according to a preset corpus acquisition rule to obtain the target subtitles of the target video; A corpus generation module, configured to verify the error rate of the target subtitles. When the error rate of the target subtitles is less than the error threshold, generate a speech-to-text training corpus for video speech recognition according to the target video and the target subtitles; Among them, the error rate is calculated by the following formula: Among them, CER is the error rate, S is the number of replaced characters, D is the number of deleted characters, I is the number of inserted characters, and N is the total number of characters.
7. A terminal device, characterized in that, Including a processor and a memory, the processor and the memory are connected to each other; The memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Data collection method and device, storage medium and electronic equipment
CN111445902A
Subtitle area identification method, device and equipment and storage medium
CN112232260A