A subtitle generation method, device, electronic device, and storage medium
By obtaining the speech and text to be recognized, using speech recognition and word segmentation algorithms to divide and correct statement units, accurate subtitles are generated, and the problem of inaccurate subtitles in the prior art is solved and the user experience is improved.
Patent Information
- Application Number
- CN202211590893.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-12-12
AI Technical Summary
In the prior art, the accuracy of subtitles generation is affected by factors such as voice volume and tone, resulting in inaccurate subtitles.
By obtaining the speech to be recognized and the first text, the speech recognition is performed using a preset speech recognition algorithm, the statement units are divided based on punctuation marks, the word segmentation processing is performed using a preset word segmentation algorithm, and the target statement unit is determined through matching, the text statement is corrected, and the target subtitles are generated.
It realizes more accurate subtitle generation, avoids typos in subtitles, and improves user experience.
Smart Images

Figure CN115862631B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly, to a subtitle generation method, apparatus, electronic device, and storage medium. Background Art
[0002] When users watch some audio-visual content, they can understand the audio-visual content by viewing the subtitles displayed on the audio-visual display screen.
[0003] In the prior art, when generating subtitles, it generally relies on manual input or automatically generates text information through speech recognition technology and encodes it into the video file. For the method of manually inputting subtitles, there is a problem of low efficiency. For the method of generating subtitles based on speech recognition, it relies on deep learning neural network technology to convert the input speech information into text information containing timestamp information, and finally converts it into a video frame range according to the timestamp information, and adds the text information to the corresponding video frames to generate subtitles. However, the accuracy of speech recognition is easily affected by factors such as speech volume and tone, resulting in different degrees of error rates in the converted text information, thus causing the generated subtitles to be inaccurate.
[0004] Therefore, how to generate subtitles more accurately is a technical problem to be solved at present.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] Embodiments of the present application provide a subtitle generation method, apparatus, electronic device, and storage medium for generating subtitles more accurately.
[0007] In a first aspect, a subtitle generation method is provided. The method includes: obtaining a speech to be recognized and a first text, performing speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, where the speech to be recognized is generated based on the first text; if the first text has punctuation marks, dividing the first text into multiple sentence units based on the punctuation marks; performing word segmentation processing on each of the sentence units based on a preset word segmentation algorithm to obtain multiple word segments; matching the text sentences with each of the word segments respectively, and determining a target sentence unit corresponding to the text sentence in each of the sentence units according to the matching result; correcting the text sentence based on the target sentence unit to obtain a target text sentence, and generating a target subtitle according to each of the target text sentences.
[0008] Second aspect, a subtitle generation device is provided, and the device includes: a speech recognition module, configured to obtain a speech to be recognized and a first text, perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm, and obtain at least one text sentence, where the speech to be recognized is generated based on the first text; a division module, configured to divide the first text into a plurality of sentence units based on the punctuation marks if the first text has punctuation marks; a word segmentation module, configured to perform word segmentation processing on each of the sentence units based on a preset word segmentation algorithm to obtain a plurality of segmented words; a matching module, configured to match each of the text sentences with each of the segmented words respectively, and determine a target sentence unit corresponding to the text sentence in each of the sentence units according to the matching result; a generation module, configured to correct the text sentence based on the target sentence unit and obtain a target text sentence, and generate a target subtitle according to each of the target text sentences.
[0009] Third aspect, an electronic device is provided, including: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the subtitle generation method described in the first aspect by executing the executable instructions.
[0010] Fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the subtitle generation method described in the first aspect is implemented.
[0011] By applying the above technical solutions, a speech to be recognized and a first text are obtained, speech recognition is performed on the speech to be recognized based on a preset speech recognition algorithm, and at least one text sentence is obtained, where the speech to be recognized is generated based on the first text; if the first text has punctuation marks, the first text is divided into a plurality of sentence units based on the punctuation marks; word segmentation processing is performed on each of the sentence units based on a preset word segmentation algorithm to obtain a plurality of segmented words; the text sentences are respectively matched with each of the segmented words, and a target sentence unit corresponding to the text sentence in each of the sentence units is determined according to the matching result; the text sentence is corrected based on the target sentence unit and a target text sentence is obtained, and a target subtitle is generated according to each of the target text sentences. By correcting the text sentence recognized by speech based on prior text information, the situation of typos in the subtitle is avoided, more accurate subtitle generation is realized, and the user experience is improved. Description of the Drawings
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required to be used in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 The flowchart shows a subtitle generation method according to an embodiment of the present invention;
[0014] Figure 2 The flowchart shows a subtitle generation method according to another embodiment of the present invention;
[0015] Figure 3 The structural diagram shows a subtitle generation device according to an embodiment of the present invention;
[0016] Figure 4 The structural diagram shows an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0018] It should be noted that those skilled in the art will easily think of other implementation manners of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.
[0019] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0020] It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0021] The embodiments of the present application provide a subtitle generation method, as Figure 1 shown, the method includes the following steps:
[0022] Step S101, obtain the speech to be recognized and the first text, perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm, and obtain at least one text sentence.
[0023] In this embodiment, the speech to be recognized is generated based on the first text, that is, the first text is the accurate text corresponding to the speech to be recognized. The speech to be recognized can be obtained by performing speech synthesis on the first text, or by recording the speech of the user reading the first text in real time. The speech to be recognized and the first text can be uploaded by the user, or can also be downloaded by the user from the cloud, or can also be sent from other terminals or servers.
[0024] Perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence. The accuracy of speech recognition is easily affected by factors such as speech volume and timbre. There may be typos in this text sentence. For example, if the first text is "I love the working people", and the text sentence is "I love the working name", the "name" in the text sentence is a typo. Therefore, subsequent correction of the text sentence needs to be based on the first text. Each text sentence carries timestamp information, and different text sentences are distinguished based on the timestamp information.
[0025] Optionally, the format of the speech to be recognized is any one of the formats including MP3, WAV, WMA, Flac, MIDI, RA, APE, AAC, CDA, MOV, etc. The formats of the first text and the text sentence are any one of the formats including txt, doc, etc. The preset speech recognition algorithm is any one of the algorithms including DTW (Dynamic Time Warping) algorithm, VQ (Vector Quantization) algorithm, HMM (Hidden Markov Model) algorithm, ANN (Artificial Neural Networks) algorithm, etc.
[0026] In some embodiments of the present application, before performing speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, the method further includes:
[0027] Perform noise reduction on the speech to be recognized based on a preset speech noise reduction algorithm.
[0028] In this embodiment, by performing noise reduction on the speech to be recognized before speech recognition, the accuracy of the recognized text sentence can be improved. Optionally, the preset speech noise reduction algorithm is any one of the noise reduction algorithms including LMS adaptive filter and notch filter noise reduction, subtraction method, Wiener filter noise reduction, double microphone noise reduction, AI noise reduction, etc.
[0029] Optionally, noise reduction can also be performed on the speech to be recognized based on a preset audio processing software, such as Adobe Audition CS6, VinylStudio, etc.
[0030] In some embodiments of the present application, the obtaining of the speech to be recognized and the first text includes:
[0031] Obtaining the first text and performing speech synthesis on the first text based on a preset speech synthesis algorithm to obtain the speech to be recognized; or,
[0032] Obtaining the speech content uploaded by the user and the first text, and using the speech content as the speech to be recognized.
[0033] In this embodiment, the speech to be recognized and the first text can be obtained in two ways. The first way is to first obtain the first text, and then perform speech synthesis on the first text based on a preset speech synthesis algorithm to obtain the speech to be recognized. The second way is that the user directly uploads the speech to be recognized and the first text. Based on different obtaining methods, a more flexible and reliable way to obtain the speech to be recognized and the first text is achieved.
[0034] Step S102, if the first text contains punctuation marks, dividing the first text into multiple sentence units based on the punctuation marks.
[0035] In this embodiment, the subtitles in the video should conform to the semantic features of the text, and the starting text in the subtitles should conform to the pause rules of the punctuation marks. In order to avoid including the fragmented text before and after the punctuation marks in the text in the complete subtitles and to correct the text sentences more accurately with the first text, when the first text contains punctuation marks (such as commas, periods, etc.), the first text needs to be divided into multiple sentence units based on the punctuation marks.
[0036] In some embodiments of the present application, after obtaining the speech to be recognized and the first text, the method further includes:
[0037] If the first text does not contain the punctuation marks, using the first text as the sentence unit.
[0038] In this embodiment, if the first text does not contain punctuation marks, it is determined that the first text is a single sentence unit and does not need to be further divided. The first text can be used as the sentence unit to ensure accurate obtaining of the sentence unit.
[0039] Step S103, performing word segmentation processing on each of the sentence units based on a preset word segmentation algorithm to obtain a plurality of word segments.
[0040] In this embodiment, it is necessary to use each word segment in the first text to match the text sentence to correct the text sentence. Word segmentation processing is performed on each sentence unit based on a preset word segmentation algorithm to obtain a plurality of word segments. For example, if the sentence unit is "I love the working people", each word segment can be ["I", "love", "working", "people"].
[0041] The preset word segmentation algorithms can be existing word segmentation algorithms, such as the maximum matching word segmentation algorithm, the shortest path word segmentation algorithm, the generative model word segmentation algorithm, the discriminative model word segmentation algorithm, the neural network word segmentation algorithm, etc. Those skilled in the art can use different preset word segmentation algorithms and different word segmentation strategies for word segmentation according to actual needs.
[0042] Step S104: Match the text statement with each of the word segments respectively, and determine the target statement unit corresponding to the text statement in each statement unit according to the matching result.
[0043] In this embodiment, the text statement is matched with each word segment respectively, and finally the target statement unit corresponding to the text statement can be determined from each statement unit. Compared with other statement units, the target statement unit is the statement unit with the highest similarity to the text statement. Therefore, if there are typos in the text statement, the target statement unit is the accurate text of the text statement.
[0044] In some embodiments of the present application, the step of matching the text statement with each of the word segments respectively and determining the target statement unit corresponding to the text statement in each statement unit includes:
[0045] Increment each of the word segments sequentially starting from a single word segment to form a plurality of consecutive word segment sequences;
[0046] Compare the text statement with each of the consecutive word segment sequences respectively, and determine the number of identical characters in the consecutive word segment sequence and the text statement;
[0047] Determine the matching score of each consecutive word segment sequence according to the ratio of the number of identical characters to the total number of characters in the text statement;
[0048] Determine the target statement unit according to the longest consecutive word segment sequence with the highest matching score.
[0049] In this embodiment, first, each word segment is incremented sequentially starting from a single word segment to form a plurality of consecutive word segment sequences. The consecutive word segment sequence is composed of a single word segment or several consecutive word segments. To ensure reliable matching, each word segment in a single consecutive word segment sequence belongs to one statement unit. Then, the text statement is compared with each consecutive word segment sequence respectively to determine the number of identical characters between the two. Next, the matching score of each consecutive word segment sequence is determined according to the ratio of the number of identical characters to the total number of characters in the text statement. The longest consecutive word segment sequence with the highest matching score has the highest similarity to the text statement. Therefore, this longest consecutive word segment sequence is determined as the target statement unit.
[0050] For example, if the first text is "I love the working people" and the sentence text is "I love the working names", and the individual word segments are ["I", "love", "working", "people"], then the corresponding continuous word segment sequences are: "I", "I love", "I love working", "I love the working people". Iterative matching is performed, that is, the following are carried out respectively:
[0051] Compare "I love the working names" with "I", and the matching score is: 1 / 6;
[0052] Compare "I love the working names" with "I love", and the matching score is: 1 / 3;
[0053] Compare "I love the working names" with "I love working", and the matching score is: 2 / 3;
[0054] Compare "I love the working names" with "I love the working people", and the matching score is: 5 / 6;
[0055] As can be seen from the above, the continuous word segment sequence "I love the working people" is the longest continuous word segment sequence with the highest matching score, and the target sentence unit is "I love the working people".
[0056] It should be noted that the solutions in the above embodiments are only a specific implementation solution proposed by this application, and other ways of matching the text sentence with each word segment belong to the protection scope of this application.
[0057] Step S105, correct the text sentence based on the target sentence unit to obtain a target text sentence, and generate target subtitles according to each target text sentence.
[0058] In this embodiment, the target sentence unit is the correct text corresponding to the text sentence. The misspelled words in the text sentence are corrected based on the target sentence unit to obtain a target text sentence without misspelled words. The target text sentences are combined together according to the time stamps to generate accurate target subtitles.
[0059] By applying the above technical solution, the speech to be recognized and the first text are obtained. Based on a preset speech recognition algorithm, the speech to be recognized is recognized to obtain at least one text sentence. Among them, the speech to be recognized is generated based on the first text. If the first text contains punctuation marks, the first text is divided into multiple sentence units based on the punctuation marks. Each sentence unit is segmented based on a preset word segmentation algorithm to obtain multiple segments. The text sentences are respectively matched with each segment, and according to the matching result, the target sentence unit corresponding to the text sentence in each sentence unit is determined. The text sentence is corrected based on the target sentence unit to obtain the target text sentence, and the target subtitle is generated according to each target text sentence. By correcting the text sentence recognized by speech based on the prior text information, the situation of misspelled words in the subtitle is avoided, and more accurate subtitle generation is realized, improving the user experience.
[0060] The embodiment of the present application also proposes a subtitle generation method, as Figure 2 shown, the method includes the following steps:
[0061] Step S201, obtain the speech to be recognized and the first text, and perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence.
[0062] In this embodiment, the speech to be recognized is generated based on the first text, that is, the first text is the accurate text corresponding to the speech to be recognized. The speech to be recognized can be obtained by performing speech synthesis on the first text. Or by real-time recording of the speech of the user reading the first text, the speech to be recognized is obtained. The speech to be recognized and the first text can be uploaded by the user, or can also be downloaded by the user from the cloud, or can also be sent from other terminals or servers.
[0063] Performing speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, the accuracy of speech recognition is easily affected by factors such as speech volume and tone, and there may be misspelled words in the text sentence. For example, if the first text is "I love the working people", and the text sentence is "I love the working person", the "person" in the text sentence is a misspelled word. Therefore, the text sentence needs to be corrected based on the first text later. Each text sentence carries timestamp information, and different text sentences are distinguished based on the timestamp information.
[0064] Step S202, if the first text contains punctuation marks, divide the first text into multiple sentence units based on the punctuation marks.
[0065] In this embodiment, the subtitles in the video should conform to the semantic features of the text, and the starting text in the subtitles should conform to the pause rules of punctuation marks. In order to avoid including fragmented words before and after punctuation marks in the text in the complete subtitles, so that the first text can more accurately correct the text statement, when there are punctuation marks (such as commas, periods, etc.) in the first text, the first text needs to be divided into multiple statement units based on the punctuation marks.
[0066] Step S203: Perform word segmentation processing on each of the statement units based on a preset word segmentation algorithm to obtain a plurality of segmented words.
[0067] In this embodiment, it is necessary to use each segmented word in the first text to match the text statement to correct the text statement. Perform word segmentation processing on each statement unit based on a preset word segmentation algorithm to obtain a plurality of segmented words. For example, if the statement unit is "I love the working people", each segmented word can be ["I", "love", "working", "people"].
[0068] Step S204: Match the text statement with each of the segmented words respectively, and determine the target statement unit corresponding to the text statement in each of the statement units according to the matching result.
[0069] In this embodiment, match the text statement with each segmented word respectively. Finally, the target statement unit corresponding to the text statement can be determined from each statement unit. Compared with other statement units, this target statement unit is the statement unit with the highest similarity to the text statement. Therefore, if there are typos in the text statement, this target statement unit is equivalent to the accurate text of the text statement.
[0070] Step S205: Correct the text statement based on the target statement unit to obtain a target text statement.
[0071] In this embodiment, the target statement unit is the correct text corresponding to the text statement. Correct the typos in the text statement based on the target statement unit to obtain a target text statement without typos.
[0072] Step S206: Determine whether the length of the target text statement is the same as the length of the target statement unit. If so, execute step S208; otherwise, execute step S207.
[0073] In this embodiment, since the text statements obtained by speech recognition have the characteristic of non-semantic consistency in some results, abnormal sentence breaks occur. For example, the first word of the current text statement is connected after the last word of the previous text statement. At this time, determine whether the length of the target text statement is the same as the length of the target statement unit. If they are not the same, it is necessary to further correct to make the target text statement more reasonable in terms of sentence pause.
[0074] Step S207, modify the length of the target text statement based on the target statement unit.
[0075] In this embodiment, if the length of the target text statement is inconsistent with the target statement unit, it indicates that there is an abnormal sentence break in the target statement unit. Modify the length of the target text statement based on the target statement unit, so as to obtain a more accurate target text statement and make the sentence break of the subtitle more accurate.
[0076] In some embodiments of the present application, the modifying the length of the target text statement based on the target statement unit includes:
[0077] If the length of the target text statement is greater than the length of the target statement unit, and the first target character segment that the target text statement has more than the target statement unit is consistent with the second target character segment, move the first target character segment before the first word in the next target text statement, where the second target character segment is the character segment at the starting part in the next statement unit after the target statement unit;
[0078] If the length of the target text statement is less than the length of the target statement unit, and the third target character segment that the target statement unit has more than the target text statement is consistent with the fourth target character segment, move the fourth target character segment after the last word in the target text statement, where the fourth target character segment is the character segment at the starting part in the next target text statement.
[0079] In this embodiment, when the length of the target text statement is inconsistent with the length of the target statement unit, there are two cases. The first case is that the length of the target text statement is greater than the length of the target statement unit. At this time, compare the first target character segment that the target text statement has more than the target statement unit with the second target character segment. The second target character segment is the character segment at the starting part in the next statement unit after the target statement unit. If the first target character segment is consistent with the second target character segment, it means that the first character segment should be the character segment at the starting part in the next target text statement. Therefore, move the first target character segment from the target text statement to before the first word in the next target text statement, so as to restore the length of the target text statement to the length at normal sentence break.
[0080] The second case is as follows: The length of the target text statement is less than the length of the target statement unit. At this time, the third target character segment that the target statement unit has more than the target text statement is compared with the fourth target character segment. The fourth target character segment is the character segment at the starting part of the next target text statement. If the third target character segment is consistent with the fourth target character segment, it indicates that the fourth target character segment should be the character segment at the end part of the target text statement. Therefore, the fourth target character segment is moved from the next target text statement to after the last word of the target text statement, so as to restore the length of the target text statement to the length at normal sentence segmentation.
[0081] It can be understood that the first target character segment and the third target character segment can be one word or multiple words.
[0082] For example, when the first target character segment and the third target character segment are one word, the following adjustment is made to the length of the target statement unit:
[0083] If the length of the target text statement is greater than the length of the target statement unit, and the last word in the target text statement is the same as the first word segment in the next statement unit after the target statement unit, it indicates that the last word in the target text statement should be the first word in the next target text statement. Move the last word in the target text statement to before the first word in the next target text statement;
[0084] If the length of the target text statement is less than the length of the target statement unit, and the last word segment of the target statement unit is the same as the first word in the next target text statement, it indicates that the first word in the next target text statement should be the last word of the target text statement. Move the first word in the next target text statement to after the last word in the target text statement.
[0085] By comparing the extra character segment in the target text statement with the next statement unit, or comparing the missing character segment in the target text statement with the next target text statement, the length of the target text statement is accurately corrected.
[0086] Step S208, generating target subtitles according to each of the target text statements.
[0087] In this embodiment, each target text statement is combined according to the time stamp to generate accurate target subtitles.
[0088] By applying the above technical solution, the speech to be recognized and the first text are obtained, and speech recognition is performed on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, where the speech to be recognized is generated based on the first text; if the first text contains punctuation marks, the first text is divided into multiple sentence units based on the punctuation marks; word segmentation processing is performed on each sentence unit based on a preset word segmentation algorithm to obtain multiple word segments; the text sentences are respectively matched with each word segment, and the target sentence unit corresponding to the text sentence in each sentence unit is determined according to the matching result; the text sentence is corrected based on the target sentence unit to obtain a target text sentence, and it is determined whether the length of the target text sentence is consistent with the length of the target sentence unit; if not, the length of the target text sentence is corrected based on the target sentence unit; target subtitles are generated according to each target text sentence. By correcting the text sentence recognized by speech based on prior text information, the situation of typos in the subtitles is avoided, and the length of the corrected text sentence is corrected to make its sentence pause more reasonable, realizing more accurate subtitle generation and improving the user experience.
[0089] An embodiment of the present application also proposes a subtitle generation device, as Figure 3 shown, the device includes:
[0090] A speech recognition module 301, configured to obtain the speech to be recognized and the first text, and perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, where the speech to be recognized is generated based on the first text;
[0091] A division module 302, configured to, if the first text contains punctuation marks, divide the first text into multiple sentence units based on the punctuation marks;
[0092] A word segmentation module 303, configured to perform word segmentation processing on each sentence unit based on a preset word segmentation algorithm to obtain multiple word segments;
[0093] A matching module 304, configured to respectively match each text sentence with each word segment, and determine the target sentence unit corresponding to the text sentence in each sentence unit according to the matching result;
[0094] A generation module 305, configured to correct the text sentence based on the target sentence unit to obtain a target text sentence, and generate target subtitles according to each target text sentence.
[0095] In a specific application scenario, the matching module 304 is specifically configured to:
[0096] Increment each word segment starting from a single word segment in sequence to form multiple consecutive word segment sequences;
[0097] Compare the text statement with each of the continuous word segmentation sequences respectively, and determine the number of identical characters in the continuous word segmentation sequence and the text statement;
[0098] Determine the matching score of each continuous word segmentation sequence according to the ratio of the number of identical characters to the total number of characters in the text statement;
[0099] Determine the target statement unit according to the longest continuous word segmentation sequence with the highest matching score.
[0100] In a specific application scenario, the device further includes a correction module, which is used for:
[0101] Judge whether the length of the target text statement is consistent with the length of the target statement unit;
[0102] If they are inconsistent, correct the length of the target text statement based on the target statement unit.
[0103] In a specific application scenario, the correction module is specifically used for:
[0104] If the length of the target text statement is greater than the length of the target statement unit, and the first target character segment that the target text statement has more than the target statement unit is consistent with the second target character segment, move the first target character segment to before the first word in the next target text statement, where the second target character segment is the character segment at the beginning of the next statement unit after the target statement unit;
[0105] If the length of the target text statement is less than the length of the target statement unit, and the third target character segment that the target statement unit has more than the target text statement is consistent with the fourth target character segment, move the fourth target character segment to after the last word of the target text statement, where the fourth target character segment is the character segment at the beginning of the next target text statement.
[0106] In a specific application scenario, the partitioning module 302 is further used for:
[0107] If the first text does not carry the punctuation mark, use the first text as the statement unit.
[0108] In a specific application scenario, the device further includes a noise reduction module, which is used for:
[0109] Perform noise reduction on the speech to be recognized based on a preset speech noise reduction algorithm.
[0110] In a specific application scenario, the speech recognition module 301 is specifically used for:
[0111] Obtain the first text and perform speech synthesis on the first text based on a preset speech synthesis algorithm to obtain the speech to be recognized; or,
[0112] Obtain the speech content uploaded by the user and the first text, and use the speech content as the speech to be recognized.
[0113] By applying the above technical solution, the subtitle generation device includes: a speech recognition module, configured to obtain the speech to be recognized and the first text, and perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, where the speech to be recognized is generated based on the first text; a division module, configured to divide the first text into multiple sentence units based on punctuation marks if the first text has punctuation marks; a word segmentation module, configured to perform word segmentation processing on each sentence unit based on a preset word segmentation algorithm to obtain multiple word segments; a matching module, configured to match each text sentence with each word segment respectively, and determine the target sentence unit corresponding to the text sentence in each sentence unit according to the matching result; a generation module, configured to correct the text sentence based on the target sentence unit to obtain a target text sentence, generate a target subtitle according to each target text sentence, and correct the text sentence recognized by speech based on prior text information, avoiding the situation of typos in the subtitle, achieving more accurate subtitle generation, and improving the user experience.
[0114] An embodiment of the present invention further provides an electronic device, as Figure 4 shown, including a processor 401, a communication interface 402, a memory 403, and a communication bus 404, where the processor 401, the communication interface 402, and the memory 403 complete mutual communication through the communication bus 404,
[0115] The memory 403 is used to store executable instructions of the processor;
[0116] The processor 401 is configured to execute via the execution of the executable instructions:
[0117] Obtain the speech to be recognized and the first text, and perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, where the speech to be recognized is generated based on the first text;
[0118] If the first text has punctuation marks, divide the first text into multiple sentence units based on the punctuation marks;
[0119] Perform word segmentation processing on each of the sentence units based on a preset word segmentation algorithm to obtain multiple word segments;
[0120] Match the text sentences with each of the word segments respectively, and determine the target sentence unit corresponding to the text sentence in each of the sentence units according to the matching result;
[0121] Correct the text statement based on the target statement unit to obtain a target text statement, and generate a target subtitle according to each of the target text statements.
[0122] The communication bus described above may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0123] The communication interface is used for communication between the terminal and other devices.
[0124] The memory may include a RAM (Random Access Memory), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0125] The processor described above may be a general-purpose processor, including a CPU (Central Processing Unit), an NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processing), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0126] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the subtitle generation method described above is implemented.
[0127] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, the computer is caused to execute the subtitle generation method described above.
[0128] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0129] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0130] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other. The focus of each embodiment is on the differences from other embodiments.
[0131] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included within the protection scope of the present invention.
Claims
1. A subtitle generation method, characterized in that, The method includes: Obtain the speech to be recognized and the first text, perform speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, where the speech to be recognized is generated based on the first text; If there are punctuation marks in the first text, divide the first text into multiple sentence units based on the punctuation marks; Perform word segmentation processing on each of the sentence units based on a preset word segmentation algorithm to obtain multiple word segments; Match the text sentences with each of the word segments respectively, and determine the target sentence unit corresponding to the text sentence in each of the sentence units according to the matching result; Correct the text sentence based on the target sentence unit and obtain a target text sentence, and generate a target subtitle according to each of the target text sentences; After correcting the text sentence based on the target sentence unit and obtaining the target text sentence, the method further includes: Judge whether the length of the target text sentence is the same as the length of the target sentence unit; If they are not the same, if the length of the target text sentence is greater than the length of the target sentence unit, and the first target character segment that the target text sentence has more than the target sentence unit is the same as the second target character segment, move the first target character segment before the first word in the next target text sentence, where the second target character segment is the character segment at the start of the next sentence unit after the target sentence unit; If the length of the target text sentence is less than the length of the target sentence unit, and the third target character segment that the target sentence unit has more than the target text sentence is the same as the fourth target character segment, move the fourth target character segment after the last word of the target text sentence, where the fourth target character segment is the character segment at the start of the next target text sentence.
2. The method according to claim 1, wherein The step of matching the text sentences with each of the word segments respectively and determining the target sentence unit corresponding to the text sentence in each of the sentence units according to the matching result includes: Increment each of the word segments starting from a single word segment in sequence to form multiple consecutive word segment sequences; Compare the text sentences with each of the consecutive word segment sequences respectively, and determine the number of identical characters in the consecutive word segment sequence and the text sentence; Determine the matching score of each of the consecutive word segment sequences according to the ratio of the number of identical characters to the total number of characters in the text sentence; Determine the target sentence unit according to the longest consecutive word segment sequence with the highest matching score.
3. The method according to claim 1, characterized in that After obtaining the speech to be recognized and the first text, the method further includes: If there are no such punctuation marks in the first text, use the first text as the sentence unit.
4. The method according to claim 1, wherein Before performing speech recognition on the speech to be recognized based on a preset speech recognition algorithm to obtain at least one text sentence, the method further includes: Perform noise reduction on the speech to be recognized based on a preset speech noise reduction algorithm.
5. The method according to claim 1, wherein The step of obtaining the speech to be recognized and the first text includes: Obtain the first text and perform speech synthesis on the first text based on a preset speech synthesis algorithm to obtain the speech to be recognized; or, Obtain the voice content uploaded by the user and the first text, and use the voice content as the voice to be recognized.
6. A subtitle generation device, characterized in that, The device includes: A voice recognition module, configured to obtain the voice to be recognized and the first text, and perform voice recognition on the voice to be recognized based on a preset voice recognition algorithm to obtain at least one text sentence, where the voice to be recognized is generated based on the first text; A division module, configured to, if the first text has punctuation marks, divide the first text into multiple sentence units based on the punctuation marks; A word segmentation module, configured to perform word segmentation processing on each of the sentence units based on a preset word segmentation algorithm to obtain multiple segmented words; A matching module, configured to match each of the text sentences with each of the segmented words respectively, and determine the target sentence unit corresponding to the text sentence in each of the sentence units according to the matching result; A generation module, configured to correct the text sentence based on the target sentence unit to obtain a target text sentence, and generate a target subtitle according to each of the target text sentences; The device further includes a correction module, configured to: Judge whether the length of the target text sentence is consistent with the length of the target sentence unit; If not, if the length of the target text sentence is greater than the length of the target sentence unit, and the first target character segment that the target text sentence has more than the target sentence unit is consistent with the second target character segment, move the first target character segment to before the first word in the next target text sentence, where the second target character segment is the character segment at the beginning of the next sentence unit after the target sentence unit; If the length of the target text sentence is less than the length of the target sentence unit, and the third target character segment that the target sentence unit has more than the target text sentence is consistent with the fourth target character segment, move the fourth target character segment to after the last word of the target text sentence, where the fourth target character segment is the character segment at the beginning of the next target text sentence.
7. An electronic device, characterized in that, Includes: A processor; And A memory, configured to store the executable instructions of the processor; Wherein, the processor is configured to execute the subtitle generation method according to any one of claims 1 to 5 by executing the executable instructions.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the subtitle generation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method, device, computer equipment and storage medium of voiceprint sample acquisition
CN109473106A
Recognition error correction device and program of the same, and subtitle generation system
JP2017021246A