Subtitle production method and device, computer-readable storage medium

By performing text correction and historical correction information update on audio files of the voice transfer system, the problem of low accuracy in transcribing people, place names, organization names and professional terms in the prior art is solved, and a more efficient subtitle production process is achieved.

CN114357979BActive Publication Date: 2025-06-17IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111672294.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-17
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

When the existing voice transcription system processes personal names, place names, organization names and professional terms, the transcription accuracy is low, resulting in large workloads and low efficiency.

Method used

By obtaining the initial transcribed text of the audio file, text correction is performed to obtain historical correction information, and using this information to update the error part in the transcribed text, thereby generating more accurate subtitle text.

Benefits of technology

It improves the accuracy of speech translation, reduces the workload of subtitle correction, and improves the efficiency of subtitle production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114357979B_ABST
    Figure CN114357979B_ABST
Patent Text Reader

Abstract

The present application discloses a subtitle production method and apparatus, and a computer-readable storage medium, belonging to the technical field of natural language processing. The subtitle production method first obtains a first transcription text corresponding to an audio file, then corrects the part of the first transcription text corresponding to the time before the current moment to obtain a first corrected text, then obtains historical correction information by using the first corrected text, and then updates the part of the first transcription text corresponding to the time after the current moment by using the historical correction information to obtain a subtitle corrected text. The present application enables the part after the current moment to be modified based on the correction historical information, thereby reducing the probability of recurrence of related errors. And with the accumulation of historical correction information, the text error rate in the updated first transcription text will gradually decrease, thereby improving the accuracy of speech transcription and reducing the workload of subtitle correction. The present application can improve the efficiency of subtitle production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and particularly to a subtitle production method, an apparatus, and a computer-readable storage medium. Background Art

[0002] Most existing subtitle production systems first perform audio transcription on speech sentences in an audio file to obtain text, and then combine post-processing to perform operations such as sentence segmentation and filtering of filler words on the transcribed text to obtain a transcribed text, and then perform subtitle correction on the transcribed text. The accuracy of audio transcription largely determines the workload of subtitle correction.

[0003] With the development of natural language processing technology, end-to-end speech transcription technology based on deep learning has been applied to subtitle production. However, the defects of current speech transcription systems are also very prominent. For example, there is still a large gap between the transcription accuracy of proper nouns such as personal names, place names, and organization names, and professional vocabulary in specific fields and expectations. This is mainly because the training data has poor coverage of proper nouns and domain terms, resulting in insufficient learning ability of the model for these words, resulting in low transcription accuracy, thereby increasing the workload in the subtitle correction process and reducing the efficiency of subtitle production. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide a subtitle production method, an apparatus, and a computer-readable storage medium, which can improve the accuracy of speech transcription and reduce the workload of subtitle correction.

[0005] To solve the above technical problem, a technical solution adopted by the present application is: to provide a subtitle production method, including:

[0006] Obtaining a first transcribed text corresponding to an audio file;

[0007] Performing text correction on the part of the first transcribed text corresponding to before the current moment to obtain a first corrected text;

[0008] Obtaining historical correction information by using the first corrected text;

[0009] Updating the part of the first transcribed text corresponding to after the current moment by using the historical correction information to obtain a subtitle corrected text.

[0010] Optionally, the step of obtaining historical correction information by using the first corrected text includes:

[0011] Preprocess the first corrected text to obtain a first segmented text, and preprocess the part of the first transcription text corresponding to the first corrected text to obtain a second segmented text; the first segmented text includes a plurality of discrete first words, and the second segmented text includes a plurality of discrete second words;

[0012] Align the first segmented text and the second segmented text so that the corresponding first words and second words are aligned to form word pairs;

[0013] Traverse all the word pairs. In response to the first word and the second word in the current word pair being different, and in response to the current word pair meeting a preset condition, generate a triple according to the current word pair; the triple includes a word pair and the number of times it appears;

[0014] Obtain the historical correction information by using the triple list composed of all the triples.

[0015] Optionally, the first corrected text includes at least one correction candidate sentence. The step of responding to the current word pair meeting the preset condition includes:

[0016] In response to the first word in the current word pair being a multi-word term; or,

[0017] In response to the first word in the current word pair being at the beginning of the corresponding correction candidate sentence and the first word after it not being a preset stop word; or,

[0018] In response to the first word in the current word pair being at the end of the corresponding correction candidate sentence and the first word before it not being a preset stop word; or,

[0019] In response to the first word in the current word pair being in the middle of the corresponding correction candidate sentence, determine that the current word pair meets the preset condition.

[0020] Optionally, the step of generating a triple according to the current word pair includes:

[0021] In response to the first word in the current pair being a multi-word term, form the triple by combining the current word pair and the corresponding cumulative number of occurrences;

[0022] In response to the first word in the current pair not being a multi-word term, obtain a current extended word pair according to the first word and its adjacent words, and form the triple by combining the current extended word pair and the corresponding cumulative number of occurrences.

[0023] Optionally, the first segmented text further includes first part-of-speech tags corresponding one-to-one to the first words, the second segmented text further includes second part-of-speech tags corresponding one-to-one to the second words, and the step of obtaining the historical correction information by using the triple list composed of all the triples includes:

[0024] Filtering out the triples in the triple list where the first part-of-speech tag or the second part-of-speech tag belongs to a preset part-of-speech list;

[0025] Taking the first words in the filtered triples as hot words, and taking the hot word list composed of all the hot words as the historical correction information.

[0026] Optionally, the first transcription text includes multiple first candidate sentence groups. Before the step of obtaining the first transcription text corresponding to the audio file, it further includes:

[0027] Obtaining a second transcription text; the second transcription text includes multiple second candidate sentence groups, one first candidate sentence group, one second candidate sentence group corresponds to a speech sentence in the audio file, and the first transcription text is obtained by reordering multiple second candidate sentences in each second candidate sentence group in the second transcription text;

[0028] The step of using the historical correction information to update the part of the first transcription text corresponding to after the current moment to obtain a subtitle correction text includes:

[0029] Using the hot word list and the second candidate sentence group in the second transcription text corresponding to after the current moment to create new candidate sentences, and adding the new candidate sentences to the corresponding second candidate sentence group to obtain a third candidate sentence group;

[0030] Reordering the third candidate sentences included in the third candidate sentence group to obtain a second correction text; the third candidate sentence is the same as the corresponding second candidate sentence or the new candidate sentence;

[0031] Replacing the first candidate sentence group in the first transcription text corresponding to after the current moment with the second correction text to obtain the subtitle correction text.

[0032] Optionally, the step of using the hot word list and the second candidate sentence group in the second transcription text corresponding to after the current moment to create new candidate sentences includes:

[0033] Construct at least one first mapping network using the hot word list, and construct at least one second mapping network using the second candidate sentence group corresponding to after the current moment in the second transcription text; wherein, the first mapping network corresponds one-to-one with the category of the hot words, the first mapping network includes at least one first mapping path, the first mapping path represents the mapping relationship between the hot word and its first pinyin sequence, and the input of the first mapping path is the first pinyin sequence and the output is the hot word; the second mapping network corresponds one-to-one with the second candidate sentence, the second mapping network includes multiple second mapping paths, the second mapping path represents the mapping relationship between the second candidate sentence and its second pinyin sequence, and the input of the second mapping path is the second candidate sentence and the output is the second pinyin sequence;

[0034] Determine whether there is a matching segment in the second pinyin sequence that satisfies the matching condition with the first pinyin sequence;

[0035] If so, add a category label between the start node and the end node of the second mapping path corresponding to the matching segment; the category label represents the type of the corresponding first mapping network;

[0036] Add a new sub-path between the start node and the end node corresponding to the category label; the sub-path is the same as the first mapping path corresponding to the matching segment;

[0037] Form a new mapping path from the part of the second mapping path that does not correspond to the matching segment and the sub-path, and use the text on the new mapping path as the new candidate sentence.

[0038] Optionally, the step of reordering the third candidate sentences included in the third candidate sentence group to obtain the second corrected text includes:

[0039] Obtain the language score of each third candidate sentence using the hot word list and the first corrected text;

[0040] Take the sum of the language score and the corresponding acoustic score as the final score; the acoustic score is obtained during the process of obtaining the first transcription text, and one second candidate sentence corresponds to one acoustic score;

[0041] Reorder the third candidate sentences included in the third candidate sentence group in the order of the magnitude of the final score to obtain the second corrected text.

[0042] Optionally, the step of obtaining the language score of each third candidate sentence using the hot word list and the first corrected text includes:

[0043] Obtain the patch language score of the third candidate sentence using the patch language model and the general language model, and obtain the main language score of the third candidate sentence using the class language model and the general language model; the inputs of the patch language model, the class language model, and the general language model are all a sentence, and the outputs are all the probabilities of the input sentence appearing. The patch language model is trained using the training text obtained from the hot word list and the first correction text, the class language model is trained using the training text including the category label, and the general language model is trained using the training text composed of general sentences without the category label;

[0044] Take the value obtained by weighted summation of the patch language score and the main language score as the language score.

[0045] Optionally, the first correction text includes at least one correction candidate sentence. The step of obtaining the patch language score of the third candidate sentence using the patch language model and the general language model includes:

[0046] Screen out the correction candidate sentences containing the hot words from the first correction text, copy them several times and add them to the first correction text to obtain the patch language training text;

[0047] Train the patch language model using the patch language training text;

[0048] Input the third candidate sentence into the patch language model and the general language model respectively, and take the value obtained by weighted summation of the output of the patch language model and the output of the general language model as the patch language score.

[0049] Optionally, the step of obtaining the main language score of the third candidate sentence using the class language model and the general language model includes:

[0050] In response to the third candidate sentence being the same as the corresponding second candidate sentence, input the second candidate sentence into the class language model and the general language model respectively; in response to the third candidate sentence being the same as the corresponding new candidate sentence, replace the hot word in the new candidate sentence with the corresponding category label, then input it into the class language model, and input the new candidate sentence into the general language model; and take the outputs of the class language model and the general language model as the class language score and the general language score of the third candidate sentence respectively;

[0051] Take the value obtained by weighted summation of the class language score and the general language score as the main language score.

[0052] Optionally, after the step of updating the part of the first transcription text corresponding to after the current moment by using the historical correction information to obtain a subtitle correction text, the method further includes:

[0053] Displaying the third candidate sentence with the highest score in each of the third candidate sentence groups of the second correction text as the subtitle of the corresponding speech sentence according to the time axis of the audio file.

[0054] To solve the above technical problems, another technical solution adopted by this application is: to provide a subtitle production device, including:

[0055] A transcription module, configured to obtain a first transcription text corresponding to an audio file;

[0056] A correction module, configured to perform text correction on the part of the first transcription text corresponding to before the current moment to obtain a first correction text;

[0057] An extraction module, configured to obtain historical correction information by using the first correction text;

[0058] An update module, configured to update the part of the first transcription text corresponding to after the current moment by using the historical correction information to obtain a subtitle correction text.

[0059] To solve the above technical problems, another technical solution adopted by this application is: to provide a subtitle production device, including a memory and a processor, where the memory stores program instructions, and the processor can execute the program instructions to implement the subtitle production method described in the above technical solution.

[0060] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, where program instructions are stored on the storage medium, and the program instructions can be executed by a processor to implement the subtitle production method described in the above technical solution.

[0061] The beneficial effects of the present application are as follows: The subtitle production method provided by the present application first obtains the first transcription text corresponding to the audio file, then corrects the part of the first transcription text corresponding to the time before the current moment to obtain the first corrected text, then obtains historical correction information by using the first corrected text, and then updates the part of the first transcription text corresponding to the time after the current moment by using the historical correction information to obtain the subtitle correction text. The present application accurately locates the correction object by using the historical correction information before the current moment, and then updates the part of the first transcription text corresponding to the time after the current moment by using the historical correction information, so that the part after the current moment is modified based on the correction history information, thereby reducing the probability of recurrence of related errors. And with the accumulation of historical correction information, the text error rate in the updated first transcription text will gradually decrease, thereby improving the accuracy of speech transcription and reducing the workload of subtitle correction. Therefore, the present application can improve the efficiency of subtitle production. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0063] Figure 1 is a schematic flowchart of an embodiment of the subtitle production method of the present application;

[0064] Figure 2 is Figure 1 a schematic flowchart of an embodiment of step S13 in ;

[0065] Figure 3 is Figure 2 a schematic flowchart of an embodiment of step S24;

[0066] Figure 4 is Figure 1 a schematic flowchart of an embodiment of step S14 in ;

[0067] Figure 5 is Figure 4 a schematic flowchart of an embodiment of step S41 in ;

[0068] Figure 6a is a schematic structural diagram of an embodiment of the first mapping network;

[0069] Figure 6b is a schematic structural diagram of an embodiment of the second mapping network;

[0070] Figure 6c is a schematic diagram of the principle of an embodiment of obtaining a new candidate sentence;

[0071] Figure 7 is Figure 4 A schematic flowchart of one embodiment of step S42 in

[0072] Figure 8 is Figure 7 A schematic flowchart of one embodiment of step S61 in

[0073] Figure 9 is Figure 8 A schematic flowchart of one embodiment of step S71 in

[0074] Figure 10 is Figure 8 A schematic flowchart of another embodiment of step S71 in

[0075] Figure 11 A schematic structural diagram of one embodiment of the subtitle production device of the present application;

[0076] Figure 12 is Figure 11 A schematic structural diagram of one embodiment of the hot word module in

[0077] Figure 13 is Figure 11 A schematic structural diagram of one embodiment of the update module in

[0078] Figure 14 A schematic structural diagram of one embodiment of the subtitle production device of the present application;

[0079] Figure 15 A schematic structural diagram of one embodiment of the computer-readable storage medium of the present application. Specific Embodiments

[0080] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0081] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of one embodiment of the subtitle production method of the present application. The subtitle production method includes the following steps.

[0082] Step S11: Obtain a first transcription text corresponding to the audio file.

[0083] The audio file for which subtitles are to be created has multiple speech sentences distributed according to its timeline, and the created subtitles correspond to these speech sentences. In this embodiment, first, a first transcription text is obtained for the audio file. The first transcription text includes multiple first candidate sentence groups, where one first candidate sentence group corresponds to one speech sentence in the audio file, and one first candidate sentence group includes multiple first candidate sentences.

[0084] Before the step of obtaining the first transcription text, this embodiment also obtains a second transcription text. The second transcription text includes multiple second candidate sentence groups. One second candidate sentence group is obtained by decoding one speech sentence, that is, one first candidate sentence group, one second candidate sentence group, and one speech sentence correspond. One second candidate sentence group includes multiple second candidate sentences, and the first transcription text is obtained by reordering the multiple second candidate sentences in each second candidate sentence group in the second transcription text.

[0085] That is to say, when the transcription system transcribes the current audio file, it first uses the speech transcription system to decode each speech sentence to obtain a second candidate sentence group, which includes multiple second candidate sentences. These multiple second candidate sentences have an order according to the scores, simply referred to as topN. Then, reordering is performed on topN, that is, two-pass scoring, to obtain the corresponding first candidate sentence group, and the multiple first candidate sentences in it also have an order according to the scores, simply referred to as N-best. Specifically, two-pass scoring can be combined with a 4th-order N-gram model and an acoustic model to obtain N-best.

[0086] However, each audio file has its own unique word usage habits and ways of organizing language. Even using a domain language model (i.e., the transcription system) that matches the audio content cannot guarantee full coverage of the vocabulary in the current audio. Therefore, the probability of mis-transcribed words appearing in either topN or N-best is relatively high.

[0087] Step S12: Perform text correction on the part of the first transcription text corresponding to before the current moment to obtain a first corrected text.

[0088] After obtaining the N-best for each speech sentence, the first candidate sentence with the highest score among them is given as the subtitle for automatic or manual correction to obtain a corrected candidate sentence. Then, before the current moment, at least one first candidate sentence with the highest score has been corrected. That is, the corrected candidate sentence is obtained by performing text correction on the first candidate sentence with the highest score in at least one first candidate sentence group corresponding to before the current moment in the first transcription text.

[0089] Step S13: Obtain historical correction information using the first corrected text.

[0090] During the above text correction process, the first candidate sentence with the highest score is corrected to a corrected candidate sentence, where the errors in the first candidate sentence usually manifest as lexical errors. For example, in an academic report introducing chronic diseases, due to the influence of many factors such as the speaker's accent, speech rate, and tone, "Plendil" appeared dozens of times, but was transcribed as words such as "not necessarily, Bo Yiding, Bo Yiding". That is, in the first candidate sentence group corresponding to those speech sentences containing "Plendil", the first candidate sentence with the highest score contains words such as "not necessarily, Bo Yiding, Bo Yiding". During text correction, these words are all corrected to "Plendil", making the frequently mis-transcribed "Plendil" a hot word, which can be focused on and incentivized during subsequent improvements to the transcription system. It can be understood that, in order to prevent false triggering, in this embodiment, hot words are mostly entity words or professional terms. The above historical correction information can be characterized by these hot words.

[0091] Step S14, update the part of the first transcription text corresponding to after the current moment using the historical correction information to obtain a subtitle correction text.

[0092] The historical correction information reflects the correction information of text correction before the current moment, that is, it reflects the weakest link when the transcription system transcribes the current audio file. It can be used to optimize the transcription system, thereby updating the part of the first transcription text corresponding to after the current moment, that is, updating the first candidate sentence group of the first transcription text corresponding to after the current moment, to obtain a subtitle correction text.

[0093] As mentioned above, before the update, the probability of errors in the first candidate sentence group is relatively high. After optimizing the transcription system using the historical correction information and improving the learning ability of the transcription system for mis-transcribed words, the probability of the same error reappearing in the subtitle correction text obtained after the update will be greatly reduced. Moreover, this update process can be carried out multiple times. For example, it is updated every N minutes, or triggered once when the number of modifications reaches the threshold M, that is, every time a certain amount of historical correction information accumulates, the transcription system is optimized once, gradually reducing the probability of mis-transcription. Thus, over time, the accuracy of speech transcription is gradually improved, the workload of subtitle correction is reduced, and the efficiency of subtitle production is increased.

[0094] In one embodiment, please combine Figure 1 Refer to Figure 2 , Figure 2 For Figure 1 a schematic flowchart of an embodiment of step S13 in

[0095] Step S21: Preprocess the first corrected text to obtain a first segmented text, and preprocess the part of the first transcription text corresponding to the first corrected text to obtain a second segmented text.

[0096] Among them, the first segmented text includes multiple discrete first words, and the second segmented text includes multiple discrete second words.

[0097] As mentioned above, the corrected candidate sentence and the corresponding first candidate sentence with the highest score are the transcription results before and after the correction of the same speech sentence, and both contain corresponding words. In this embodiment, the first corrected text is preprocessed by word segmentation and part-of-speech prediction to obtain a first segmented text, that is, the first segmented text includes multiple discrete first words and first part-of-speech tags corresponding to the first words one by one; at the same time, the part of the first transcription text corresponding to the first candidate sentence with the highest score before the current moment is preprocessed by word segmentation and part-of-speech prediction to obtain a second segmented text, that is, the second segmented text includes multiple discrete second words and second part-of-speech tags corresponding to the second words one by one.

[0098] For example, the first segmented text is: hypertension / n medical history / n for three / m years / m long-term / d taking / v Boyiding / n; the second segmented text is: hypertension / n died / v for three / m years / m long-term / d taking / v not necessarily / d.

[0099] Step S22: Align the first segmented text and the second segmented text so that the corresponding first words and second words are aligned to form word pairs.

[0100] After obtaining the first segmented text and the second segmented text, align them according to the minimum edit distance criterion so that the corresponding first words and second words are aligned to form word pairs. Inevitably, if the first word and the second word in the word pair are different, it means that a transcription error has occurred and text correction has been performed. For example, the word pair formed by "medical history" in the first segmented text example and "died" in the second segmented text example, and the word pair formed by "Boyiding" in the first segmented text example and "not necessarily" in the second segmented text example.

[0101] Step S23: Traverse all word pairs. In response to the first word and the second word in the current word pair being different, and in response to the current word pair meeting a preset condition, generate a triple according to the current word pair; the triple includes a word pair and the number of times it appears.

[0102] After aligning the first segmented text and the second segmented text to obtain multiple word pairs, traverse all the word pairs. When two conditions are met simultaneously, one is that the first word and the second word in the current word pair are different, and the other is that the current word pair meets the preset conditions. Generate a triple according to the current word pair and add it to the triple list, where the triple includes a word pair and the number of times it appears. For example, the triple list corresponding to the above examples of the first segmented text and the second segmented text is \(\{(medical history / n -> die of illness / v1), (Plendil / n -> not necessarily / d1)\}\), where the number "1" indicates that the word pair appears 1 time.

[0103] Specifically, the above preset conditions include the following situations. That is to say, when any one of the following conditions is met, it is determined that the current word pair meets the above preset conditions:

[0104] One is that the first word in the current word pair is a multi-word term;

[0105] Two is that the first word in the current word pair is at the beginning of the corresponding corrected candidate sentence, and the first word after it is not a preset stop word;

[0106] Three is that the first word in the current word pair is at the end of the corresponding corrected candidate sentence, and the first word before it is not a preset stop word;

[0107] Four is that the first word in the current word pair is in the middle of the corresponding corrected candidate sentence.

[0108] Among them, the preset stop words are a set of non-entity words such as preset modal particles, conjunctions, pronouns, prepositions, etc.

[0109] For example, both "medical history" and "Plendil" in the above examples are multi-word terms, and it is directly determined that the corresponding word pairs meet the preset conditions and generate the corresponding triples.

[0110] Specifically, when the first word in the current word pair is a multi-word term, directly form a triple with the current word pair and the corresponding cumulative number of occurrences. The triple can be generated according to the following steps:

[0111] In response to the first occurrence of the current word pair, form a triple with the current word pair and the number 1 and add it to the triple list; in response to the non-first occurrence of the current word pair, add 1 to the number in the triple in the triple list that includes the current word pair.

[0112] For example, when the word pair "medical history -> die of illness" appears for the second time, no new triple needs to be generated, but the number "1" in the existing triple \((medical history / n -> die of illness / v1)\) in the triple list is updated to "2".

[0113] For another example, the first segmented text is: Medicine / n, right / u, is / v, Byetta / v, sugar / n, apple / n, not / d, is / v, called / v, what / r, Byetta / n; The second segmented text is: Want / c, right / u, is / v, Byetta / v, pear / n, apple / n, not / d, is / v, called / v, what / r, heart / n. After aligning the two, the word pairs with different first and second words that can be extracted are "Medicine -> Want", "sugar -> pear", and "Byetta -> heart". Among them, "Medicine" is at the beginning of the sentence and the first word after it is the stop word "right", and "Byetta" is at the end of the sentence and the first word before it is the stop word "what". These two pairs do not meet the above preset conditions and no triples need to be generated based on them. However, "sugar" is in the middle of the corresponding corrected candidate sentence, which meets the preset conditions and triples need to be generated based on it.

[0114] Specifically, when the first word in the current word pair is not a multi-word term, that is, a single-word term, the current extended word pair is obtained based on the first word and its adjacent words, and the current extended word pair and the corresponding cumulative occurrence count are formed into a triple. Specifically, the triple can be generated according to the following steps:

[0115] In response to the current word pair appearing for the first time, the current extended word pair is obtained based on the current word pair, and the current extended word pair and the number 1 are formed into a triple and added to the triple list;

[0116] In response to the current word pair not appearing for the first time, the number in the triple including the current word pair in the triple list is incremented by 1.

[0117] Specifically, the current extended word pair can be obtained from the current word pair according to the following steps:

[0118] In response to the first word in the current word pair being at the beginning or end of the corresponding corrected candidate sentence, the first word and the first word after or before it are combined to form a first extended word, the second word in the corresponding first candidate sentence and the first word after or before it are combined to form a second extended word, and the first extended word and the second extended word are combined to form the current extended word pair;

[0119] In response to the first word in the current word pair being in the middle of the corresponding corrected candidate sentence, the first word and the first words before and after it are combined to form a first extended word, the second word in the corresponding first candidate sentence and the first words before and after it are combined to form a second extended word, and the first extended word and the second extended word are combined to form the current extended word pair.

[0120] For example, the current word pair in the above example is "Tang->Tang", the first expanded word is "Bai Tang Ping", the second expanded word is "Bai Tang Ping", and the corresponding triple is (Bai / v Tang / n Ping / n->Bai / v Tang / n Ping / n1). When the expanded word pair "Bai Tang Ping->Bai Tang Ping" appears again, the number "1" in the existing triple (Bai / v Tang / n Ping / n->Bai / v Tang / n Ping / n1) in the triple list is updated to "2".

[0121] It can be seen that as time goes by, the number of triples in the triple list will gradually increase, which can also reflect the probability of the same first word being incorrectly transcribed.

[0122] Step S24, obtaining historical correction information using the triple list consisting of all triples.

[0123] After obtaining the triple list, further use it to obtain historical correction information. Figure 2 See also Figure 3 , Figure 3 for Figure 2 This is a flowchart of an implementation of step S24. The hot word list can be obtained by using the triple list through the following steps.

[0124] Step S31, filtering out triples whose first part-of-speech tags or second part-of-speech tags belong to a preset part-of-speech list from the triple list.

[0125] After obtaining the triple list, filter out the triples that include meaningless words. Specifically, a part-of-speech list can be preset, including entity parts of speech such as nouns, personal nouns, place nouns, institutional nouns, proper nouns, and verbs. When the first part-of-speech tag or the second part-of-speech tag belongs to the list, filter out the corresponding triple. In this way, meaningless words such as pronouns, prepositions, and modal particles are filtered out, and the negative impact caused by over-excitation caused by adding non-entity words to the hot word list is avoided.

[0126] Step S32: taking the first word in the screened triple as a hot word, and taking a hot word list consisting of all hot words as historical correction information.

[0127] After selecting the triples, the first word is directly used as the hot word, and all the hot words are combined into a hot word list. It can be seen that the hot words are those meaningful words that have transcription errors.

[0128] In this embodiment, hot words are extracted from the calibration information before the current moment to accurately locate the calibration object, and then the first candidate sentence group after the current moment is updated, so that the first candidate sentence group after the current moment is modified based on the calibration history, thereby reducing the probability of recurrence of related errors. Moreover, with the accumulation of historical calibration information, the text error rate in the updated first candidate sentence group will gradually decrease, thereby improving the accuracy of speech transcription and reducing the workload of subtitle calibration.

[0129] In one embodiment, please refer to Figure 1 refer to Figure 4 , Figure 4 For Figure 1 a schematic flowchart of an embodiment of step S14 in

[0130] Step S41: Create new candidate sentences by using the hot word list and the second candidate sentence group corresponding to the current moment and after in the second transcription text, and add the new candidate sentences to the corresponding second candidate sentence group to obtain a third candidate sentence group.

[0131] As described above, the probability of transcription errors in the first candidate sentences is relatively high and needs to be updated. The hot word list obtained through the foregoing embodiment reflects the historical calibration information before the current moment, especially including the words with transcription errors. Based on this, new candidate sentences can be created by using the hot word list and the second candidate sentence group corresponding to the current moment and after in the second transcription text. The new candidate sentences are strongly related to the hot words, so as to reduce the probability of recurrence of the same transcription errors. The specific process of creating new candidate sentences will be described below.

[0132] Step S42: Reorder the third candidate sentences included in the third candidate sentence group to obtain a second corrected text.

[0133] After creating new candidate sentences and obtaining the third candidate sentence group, reordering the third candidate sentences included therein is equivalent to performing two-pass scoring again to obtain the N-best after the current moment. It can be understood that the third candidate sentences are the same as the corresponding second candidate sentences or new candidate sentences.

[0134] Step S43: Replace the first candidate sentence group corresponding to the current moment and after in the first transcription text with the second corrected text to obtain a subtitle corrected text.

[0135] After obtaining the N-best after the current moment again, the first candidate sentence group after the current moment is updated, that is, the first candidate sentence group corresponding to the current moment and after in the first transcription text is replaced with the third candidate sentence group corresponding to the same speech sentence. Since the speech transcription error rate of the third candidate sentence group is lower, the speech transcription error rate in the subtitle corrected text is also lower.

[0136] In this embodiment, hot words are extracted from the calibration information before the current moment, and then the first candidate sentence group after the current moment is updated, so that the first candidate sentence group after the current moment is updated based on the calibration history, thereby reducing the probability of recurrence of related errors.

[0137] In one embodiment, please refer to Figure 4 for reference Figure 5 , Figure 5 which is Figure 4 a schematic flowchart of an embodiment of step S41 in

[0138] Step S51: Construct at least one first mapping network by using the hot word list, and construct at least one second mapping network by using the second candidate sentence group corresponding to the second transcription text after the current moment.

[0139] As described above, the hot word list includes at least one hot word, and these hot words can be classified into different categories, such as drug name category, person name category, and so on. In this embodiment, the hot words in the hot word list are further segmented first, and at least one first mapping network is constructed according to the mapping relationship between the segmentation and the dictionary. Among them, the first mapping network corresponds to the category of the hot word one by one. The first mapping network includes at least one first mapping path, and the first mapping path represents the mapping relationship between the hot word and its first pinyin sequence, and the input of the first mapping path is the first pinyin sequence, and the output is the hot word.

[0140] Please refer to Figure 6a , Figure 6a which is a schematic structural diagram of an embodiment of the first mapping network. Taking the hot word list of the drug name category as an example, the hot words included are "Plendil", "Gan Kang", "Fenbid", etc. Each hot word corresponds to a first mapping path. For example, the first mapping path formed by nodes "0", "2", "3", "1" corresponds to the hot word "Plendil".

[0141] At the same time, this embodiment also constructs at least one second mapping network by using the second candidate sentence group corresponding to the second transcription text after the current moment. Among them, the second mapping network corresponds to the second candidate sentence one by one, that is, each second candidate sentence in each second candidate sentence group corresponds to a second mapping network. The second mapping network includes multiple second mapping paths, and the second mapping path represents the mapping relationship between the second candidate sentence and its second pinyin sequence, and the input of the second mapping path is the second candidate sentence, and the output is the second pinyin sequence. There is no need to re-segment here, and the output of the speech transcription system is in units of words.

[0142] Please refer to Figure 6b ,Figure 6b It is a schematic structural diagram of an embodiment of the second mapping network, only drawing the part corresponding to "not necessarily" in the corresponding second candidate sentence.

[0143] Step S52: Determine whether there is a matching segment in the second pinyin sequence that meets the matching condition with the first pinyin sequence.

[0144] After constructing the first mapping network and the second mapping network, compare each second mapping network with each first mapping network one by one to determine whether there is a matching segment in the second pinyin sequence in a certain second mapping network that meets the matching condition with the first pinyin sequence in a certain first mapping network. The matching condition can specifically be that a certain segment in the second pinyin sequence is the same as the first pinyin sequence, or that the pinyin string distance between a certain segment in the second pinyin sequence and the first pinyin sequence meets a preset threshold τ.

[0145] If it is determined that there is a matching segment in the second pinyin sequence that meets the matching condition with the first pinyin sequence, then execute the following step S53. Otherwise, continue the judgment in step S52.

[0146] Step S53: Add a category label between the start node and the end node corresponding to the matching segment of the second mapping path, and the category label represents the type of the corresponding first mapping network.

[0147] Each first mapping network includes hot words belonging to the same category and has a corresponding category label. For example Figure 6a the category label corresponding to the first mapping network in is "drug name" or "YaoMing". When it is determined that there is the above-mentioned matching segment in the second pinyin sequence, add this category label between the start node and the end node corresponding to the matching segment of the second mapping path.

[0148] Please combine Figure 6a and Figure 6b Refer to Figure 6c , Figure 6c is a schematic diagram of the principle of an embodiment for obtaining a new candidate sentence. Through comparison, it is found that Figure 6b there is a matching segment "not: bu necessarily: yi ding" in the second mapping path in that is very similar in pronunciation to a complete first mapping path "wave: bo yi ding: ding" in the first mapping network of the drug name category, and the distance between the two pinyin strings is less than the preset threshold τ. At this time, stick a class arc between the start node N and the end node M corresponding to the matching segment in the second mapping path, and add the corresponding category label (YaoMing) to this class arc, as shown in (a) of Figure 6c .

[0149] Step S54: Add a new sub-path between the start node and the end node corresponding to the category label, and the sub-path is the same as the first mapping path corresponding to the matching segment.

[0150] In the second mapping network after pasting the class arcs, a sub-path is added between the start node and the end node corresponding to the class label, and this sub-path is the same as the first mapping path of the corresponding matching segment. For example, between the start node N and the end node M of the second mapping network shown in (a) of Figure 6c , a sub-path "wave: bo depend: yi ding" is added, as shown in (c) of Figure 6c .

[0151] Specifically, the second mapping network can be first extended by N-gram and language scores are added to the original second mapping path and the newly added class arcs. The class arcs are pasted with class language scores. For example, the score of the YaoMing class N-gram is pasted between the above start node N and end node M, as shown in (b) of Figure 6c . Then the extended YaoMing class arc is restored to the first mapping path that is matched. For example, the class arc YaoMing between the start node N and the end node M in the second mapping network shown in (b) of Figure 6c corresponds to "bo yi ding" in the first mapping network. At this time, it needs to be restored to the first mapping path, as shown in (c) of Figure 6c .

[0152] Step S55: Combine the part of the second mapping path that does not correspond to the matching segment and the sub-path to form a new mapping path, and use the text on the new mapping path as the new candidate sentence.

[0153] Furthermore, combine the part of the second mapping path that does not correspond to the matching segment and the sub-path to form a new mapping path, and use the text on the new mapping path as the new candidate sentence. For example, in the second mapping network shown in (c) of Figure 6c , combine the part before the start node N, the newly added sub-path, and the part after the end node M to form a new mapping path.

[0154] In this embodiment, the second candidate sentence corresponding to the current moment generated by one-pass decoding is pasted with arcs using a hot word list to obtain a new candidate sentence, so as to obtain a third candidate sentence group, which is convenient for updating the first candidate sentence group after the current moment based on the correction history, thereby reducing the probability of the recurrence of hot word-related errors.

[0155] In one embodiment, please refer to Figure 4 and Figure 7 , Figure 7 which is Figure 4 a schematic flowchart of an embodiment of step S42 in

[0156] Step S61: Obtain the language score of each third candidate sentence by using the hot word list and the first corrected text.

[0157] Specifically, please refer to Figure 7 for Figure 8 , Figure 8 which Figure 7 is a schematic flowchart of an implementation manner of step S61 in

[0158] Step S71: Obtain the patch language score of the third candidate sentence by using the patch language model and the general language model, and obtain the main language score of the third candidate sentence by using the class language model and the general language model.

[0159] The patch language model, the class language model, and the general language model are all trained based on the N-gram model. The N-gram model is a language model (LM). A language model is a probability-based discriminative model. Its input is a sentence (a sequence of words in order), and the output is the probability of this sentence, that is, the joint probability of these words. That is, in this implementation manner, the inputs of the patch language model, the class language model, and the general language model are all a sentence, and the outputs are all the probabilities of the input sentences appearing.

[0160] N-gram itself also refers to a set composed of N words. The words have a sequence, and it is not required that the words are different from each other. Commonly seen are N = 2, N = 3, N = 4, etc. Suppose there is a sentence S = (w1, w2,..., w n ), and suppose the appearance of a word is only related to the (N - 1) words before it. For example, when N = 4, the appearance of a word w n is only related to the 3 words before it. That is, the probability P(S) of sentence S can be calculated by the following formula (1):

[0161] P(S) = p(w1)p(w2|w1)...p(w n |w n-3 w n-2 w n-1 )......(1).

[0162] In this implementation manner, the patch language model, the class language model, and the general language model are all obtained based on the 4th-order N-gram model. For simplicity of illustration, hereinafter, P(w n |w n-3 w n-2 w n-1 ) is used to represent the probability of a sentence.

[0163] Among them, the patch language model is trained using the training text obtained from the hot word list and the first corrected text. The class language model is trained using the training text including class labels (such as "YaoMing"). The general language model is trained using the training text composed of general statements without class labels. The class language model and the general language model are trained well using big data before subtitle production. The training text of the patch language model is strongly related to the above-mentioned hot words and is synchronously trained during the text correction process. The training text has a small volume and high training efficiency, so it can be trained multiple times as the hot words accumulate, for example, in accordance with the aforementioned update frequency, so as to continuously stimulate the hot words and improve the accuracy of the transcribing system.

[0164] Specifically, use P1(w n |w n-3 w n-2 w n-1 ) to represent the main language score of the third candidate sentence, and use P2(w n |w n-3 w n-2 w n-1 ) to represent the patch language score of the third candidate sentence.

[0165] For example, one of the second candidate sentences is "I have been taking it for a long time, not necessarily", and the new candidate sentence derived after personalized arc pasting is "I have been taking Amlodipine for a long time", and both before and after "Amlodipine" are with the class label "YaoMing", indicating that "Amlodipine" is a drug name. Both of these two sentences are third candidate sentences, and the patch language scores and main language scores of these two sentences are obtained respectively. The specific calculation process will be described below.

[0166] Step S72, take the value obtained by weighted summation of the patch language score and the main language score as the language score.

[0167] That is, the language score P(w n |w n-3 w n-2 w n-1 ) of the third candidate sentence can be calculated using the following formula (2): P1(w n |w n-3 w n- 2w n-1 ) = αP1(w n |w n-3 w n-2 w n-1 ) + (1 - α)P2(w n |w n-3 w n-2 w n-1 )...(2);

[0168] Among them, α is a weighting coefficient. Inevitably, α < 1. For example, α = 0.8.

[0169] Due to the excitation of hot words during the training process, the language score of the second candidate sentence "I have been taking it for a long time, not necessarily" will be significantly lower than that of the corresponding new candidate sentence "I have been taking Plendil for a long time", making the subsequent new candidate sentences rank higher.

[0170] Step S62: Use the sum of the language score and the corresponding acoustic score as the final score.

[0171] Among them, the acoustic score Acm is obtained during the process of obtaining the first transcription text. One second candidate sentence corresponds to one acoustic score, that is, the acoustic score of each second candidate sentence is obtained during the reordering process of the second candidate sentences in the second candidate sentence group. The acoustic score of the corresponding new candidate sentence can be adjusted and obtained based on the acoustic score of the second candidate sentence, which is the same as in the prior art.

[0172] For example, in the above example, if the acoustic score of the second candidate sentence "I have been taking it for a long time, not necessarily" is Acm(topN1), then the acoustic score Acm(topN2) of the corresponding new candidate sentence "I have been taking Plendil for a long time" can be obtained using the following formula (3):

[0173] Acm(topN2) = Acm(topN1) + plenty……(3);

[0174] Among them, plenty is the distance between "not necessarily" and "Plendil" calculated according to the pre-trained pronunciation confusion matrix.

[0175] Furthermore, the final score P of the third candidate sentence final can be expressed by the following formula (4):

[0176] P final = P(w n |w n-3 w n-2 w n-1 ) + Acm……(4).

[0177] It can be seen that the second candidate sentence and the corresponding new candidate sentence will each obtain a final score P final , that is, each third candidate sentence has a final score P final .

[0178] Step S63: Reorder the third candidate sentences included in the third candidate sentence group in the order of the final score size to obtain the second corrected text.

[0179] After obtaining a final score P for each third candidate sentence finalAfter that, reorder the third candidate sentences in each third candidate sentence group according to their sizes, that is, obtain the second corrected text.

[0180] This embodiment reorders the third candidate sentences included in the third candidate sentence group based on language scores and acoustic scores, which can give higher score incentives to the new candidate sentences related to hot words, improve the scores of the corresponding third candidate sentences, make them occupy a forward position in the reordering, thereby reducing the probability of the recurrence of hot word-related errors, improving the transcription accuracy of the transcription system, and reducing the workload of subtitle correction.

[0181] In one embodiment, please refer to Figure 8 refer to Figure 9 , Figure 9 is Figure 8 a schematic flowchart of an embodiment of step S71 in, and the patch language scores of the third candidate sentences can be obtained through the following steps using the patch language model and the general language model.

[0182] Step S81: Screen out the corrected candidate sentences containing hot words from the first corrected text, and copy them several times and then add them to the first corrected text to obtain the patch language training text.

[0183] At the current moment, after obtaining the first corrected text and the hot word list, use the hot word list to retrieve the first corrected text, screen out the corrected candidate sentences containing hot words, and these corrected candidate sentences need to be enhanced when training the patch language model. They can be copied several times and then added to the first corrected text to obtain the patch language training text.

[0184] Step S82: Train the patch language model using the patch language training text.

[0185] Further train the patch language model using the patch language training text. Specifically, first perform the above-mentioned word segmentation operation on the patch language training text, and then obtain the intersection with the word set (the word set obtained based on big data) when training the above general language model as the word set for training the patch language model, so that the patch language model can stimulate sentences containing hot words and avoid the problem that the trained model is very sparse due to less training text and too large a word set.

[0186] Step S83: Input the third candidate sentences into the patch language model and the general language model respectively, and take the value obtained by weighted summation of the output of the patch language model and the output of the general language model as the patch language score.

[0187] After training, further input the third candidate sentences into the patch language model and the general language model respectively, and the output P of the patch language model N (w n |w n-3 wn-2 w n-1 ) and the output P of the general language model B (w n |w n-3 w n-2 w n-1 ) The value obtained by weighted summation is used as the patch language score, and the patch language score P2(w n |w n-3 w n-2 w n-1 ) of the third candidate sentence can be obtained using the following formula (5):

[0188] P2(w n |w n-3 w n-2 w n-1 ) = βP N (w n |w n-3 w n-2 w n-1 ) + (1 - β)P B (w n |w n-3 w n-2 w n-1 )... (5);

[0189] Among them, β is the weighting coefficient, and necessarily, β < 1. For example, β = 0.6.

[0190] This embodiment trains the patch language model using the hot word list and the first corrected text to improve the score incentive for sentences containing hot words, and appropriately adjusts the output of the patch language model using the general language model, which can avoid the problem that the output probability of the trained patch language model is too high due to less training text.

[0191] In one embodiment, please refer to Figure 8 for Figure 10 , Figure 10 which is Figure 8 a schematic flowchart of another embodiment of step S71 in

[0192] Step S91: In response to the third candidate sentence being the same as the corresponding second candidate sentence, input the second candidate sentence into the class language model and the general language model respectively; in response to the third candidate sentence being the same as the corresponding new candidate sentence, replace the hot words in the new candidate sentence with the corresponding category labels, then input it into the class language model, and input the new candidate sentence into the general language model; and use the outputs of the class language model and the general language model as the class language score and the general language score of the third candidate sentence respectively.

[0193] As described above, the third candidate sentence is the same as the corresponding second candidate sentence or the new candidate sentence. When the third candidate sentence is the same as the second candidate sentence, the second candidate sentence is input into the class language model and the general language model respectively, and the outputs of the class language model and the general language model are used as the class language score P of the third candidate sentence A (w n |w n-3 w n-2 w n-1 ) and the general language score P B (w n |w n-3 w n-2 w n-1 ). It can be understood that at this time, the class language score and the general language score are the same.

[0194] When the third candidate sentence is the same as the new candidate sentence, the new candidate sentence is input into the general language model, and the output of the general language model is used as the general language score P of the third candidate sentence B (w n |w n-3 w n-2 w n-1 ), meanwhile, the hot words in the new candidate sentence are replaced with the corresponding category labels, and then input into the class language model, and the output of the class language model is used as the class language score P of the third candidate sentence A (w n |w n-3 w n-2 w n-1 ). At this time, the class language score and the general language score are not the same, and the training process of the class language model gives score incentives to sentences with category labels, making the class language score higher, thus improving the score incentive for the new candidate sentence.

[0195] For example, for the above example, the second candidate sentence "I have been taking it for a long time, not necessarily" is input into the class language model and the general language model respectively, and the obtained class language score and general language score are P A(topN1) (w n |w n-3 w n-2 w n-1 ) and P B(topN1) (w n |w n-3 w n-2 w n-1 ), and these two are essentially equal. For the corresponding new candidate sentence "I have been taking Plendil for a long time", when it is input into the general language model, the obtained general language score is P B(topN2) (w n |w n-3 w n-2 w n-1)。Meanwhile, replace "Boyiding" in it with the category label "YaoMing" to get "I have been taking YaoMing for a long time", and then input it into the class language model. The obtained class language is divided into P A(topN2) (w n |w n-3 w n-2 w n-1 )。

[0196] Step S92, take the value obtained by weighted summation of the class language score and the ordinary language score as the main language score.

[0197] Furthermore, the class language score P of the third candidate sentence A (w n |w n-3 w n-2 w n-1 ) and the ordinary language score P B (w n |w n-3 w n- 2w n-1 ) are weighted and summed to obtain the above-mentioned main language score P1(w n |w n-3 w n-2 w n-1 ) of the third candidate sentence. That is, the main language score of the third candidate sentence can be obtained through the following formula (6):

[0198] P1(w n |w n-3 w n-2 w n-1 ) = γP A (w n |w n-3 w n-2 w n-1 ) + (1 - γ)P B (w n |w n-3 w n-2 w n-1 )...(6);

[0199] Among them, γ is the weighting coefficient. Necessarily, γ < 1. For example, γ = 0.7.

[0200] This embodiment uses the class language model to improve the score incentive for sentences containing hot words, can give higher score incentives to new candidate sentences related to hot words, improve the score of the corresponding third candidate sentence, and make it occupy a forward position in the re-ranking.

[0201] Furthermore, after using the hot word list to update the first candidate sentence group corresponding to after the current moment in the first transcription text to obtain the subtitle correction text, for each third candidate sentence group in the second correction text, the final score Pfinal The highest third candidate sentence is displayed as the subtitle of the corresponding speech sentence according to the time axis of the audio file for automatic or manual correction. At this time, the transcribing system has been optimized according to the above embodiments, and the transcribing accuracy of the subtitle correction text has been improved, and the incorrect transcriptions that need to be corrected have been greatly reduced.

[0202] This application captures the modification behavior in the subtitle correction process, accurately locates the modification object, and applies it to the topN candidate personalized arc obtained by the first-pass decoding, updating the first transcription text after the current moment. At the same time, using the historical correction information before the current moment, rapid patch language model training is performed. The combined use of hot words and the patch language model significantly improves the transcribing accuracy. More importantly, the entire process does not require the first-pass decoding and is completed during the second-pass rescoring, making full use of the log information recorded by the non-real-time transcribing system with low cost. The user can control the frequency of updating the first transcription text after the current moment by themselves. The patch language model can learn the language style and domain vocabulary of the current audio file more fully. As time goes by, the content that needs to be modified in the subtitle correction process will become less and less, achieving the effect of intelligent error correction. And the continuous update of the first transcription text of the uncorrected part after the current moment has no impact on the efficiency of the subtitle correction process, presenting the technical effect of being more accurate with each update, thereby greatly improving the subtitle production efficiency.

[0203] Based on the same inventive concept, this application also provides a subtitle production device. Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an embodiment of the subtitle production device of this application. The subtitle production device includes a transcribing module 11, a correction module 12, an extraction module 13, and an update module 14.

[0204] Among them, the transcribing module 11 is used to obtain the first transcription text corresponding to the audio file, where the first transcription text includes multiple first candidate sentence groups, one first candidate sentence group corresponds to one speech sentence in the audio file, and one first candidate sentence group includes multiple first candidate sentences. The correction module 12 is used to perform text correction on the part of the first transcription text corresponding to before the current moment to obtain the first corrected text; the first corrected text includes at least one correction candidate sentence, and the correction candidate sentence is obtained by performing text correction on the highest-scoring first candidate sentence in at least one first candidate sentence group corresponding to before the current moment in the first transcription text. The extraction module 13 is used to obtain historical correction information by using the first corrected text. The update module 14 is used to update the part of the first transcription text corresponding to after the current moment by using the historical correction information to obtain the subtitle correction text.

[0205] This embodiment can improve the accuracy of speech transcribing and reduce the workload of subtitle correction, thereby improving the efficiency of subtitle production.

[0206] In one embodiment, refer to Figure 12 , Figure 12 which is Figure 11 a schematic structural diagram of an embodiment of the extraction module in

[0207] The extraction module 13 includes a preprocessing module 131, a triple module 132, and an obtaining module 133. Among them, the preprocessing module 131 is used to preprocess the first corrected text to obtain a first segmented text, and preprocess the part of the first transcription text corresponding to the first corrected text (that is, the first candidate sentence with the highest score before the current moment in the first transcription text) to obtain a second segmented text; the first segmented text includes a plurality of discrete first words, and the second segmented text includes a plurality of discrete second words. The processing module 131 is further used to align the first segmented text and the second segmented text so that the corresponding first words and second words are aligned to form word pairs.

[0208] Among them, the triple module 132 is used to traverse all word pairs. In response to the first word and the second word in the current word pair being different, and in response to the current word pair satisfying a preset condition, generate a triple according to the current word pair; the triple includes a word pair and the number of times it appears.

[0209] Specifically, the triple module 132 determines whether the current word pair satisfies any one of the following multiple conditions. If so, it determines that the current word pair satisfies the above preset condition.

[0210] One is that the first word in the current word pair is a multi-word term;

[0211] Two is that the first word in the current word pair is at the beginning of the corresponding corrected candidate sentence, and the first word after it is not a preset stop word;

[0212] Three is that the first word in the current word pair is at the end of the corresponding corrected candidate sentence, and the first word before it is not a preset stop word;

[0213] Four is that the first word in the current word pair is in the middle of the corresponding corrected candidate sentence.

[0214] Among them, when the first word in the current word pair is a multi-word term, the triple module 132 is used to form a triple with the current word pair and the corresponding cumulative number of occurrences. Specifically, the triple module 132, in response to the current word pair appearing for the first time, forms a triple with the current word pair and the number 1, and adds it to the triple list; the triple module 132, in response to the current word pair not appearing for the first time, adds 1 to the number in the triple in the triple list that includes the current word pair.

[0215] Wherein, when the first word in the current word pair is not a multi-word term, the triple module 132 is used to obtain the current extended word pair based on the first word and its adjacent words, and form a triple by combining the current extended word pair and the corresponding cumulative occurrence count. Specifically, in response to the current word pair appearing for the first time, the triple module 132 obtains the current extended word pair according to the current word pair, forms a triple by combining the current extended word pair and the number 1, and adds it to the triple list. In response to the current word pair not appearing for the first time, the triple module 132 increments the number in the triple including the current word pair in the triple list.

[0216] Specifically, in response to the first word in the current word pair being at the beginning or end of the corresponding corrected candidate sentence, the triple module 132 forms a first extended word by combining the first word and the first word after or before it, forms a second extended word by combining the second word in the corresponding first candidate sentence and the first word after or before it, and forms the current extended word pair by combining the first extended word and the second extended word; in response to the first word in the current word pair being in the middle of the corresponding corrected candidate sentence, the triple module 132 forms a first extended word by combining the first word and the first word after and before it, forms a second extended word by combining the second word in the corresponding first candidate sentence and the first word after and before it, and forms the current extended word pair by combining the first extended word and the second extended word.

[0217] Wherein, the obtaining module 133 is used to obtain historical correction information by using the triple list composed of all triples. The above first segmented text further includes a first part-of-speech tag corresponding to the first word one by one, and the second segmented text further includes a second part-of-speech tag corresponding to the second word one by one. The obtaining module 133 is specifically used to filter out the triples in the triple list whose first part-of-speech tag or second part-of-speech tag belongs to a preset part-of-speech list. The obtaining module 133 is also specifically used to use the first word in the filtered triples as hot words, and use the hot word list composed of all hot words as the above historical correction information.

[0218] This embodiment extracts hot words from the correction information before the current moment to accurately locate the correction object, facilitating the subsequent update of the first candidate sentence group after the current moment, so that the first candidate sentence group after the current moment is modified based on the correction history, thereby reducing the probability of related errors occurring again.

[0219] In one embodiment, please refer to Figure 13 , Figure 13 For Figure 11 a structural schematic diagram of an embodiment of the update module in

[0220] In this embodiment, before obtaining the first transcription text corresponding to the audio file, the transcription module 11 is further configured to obtain a second transcription text; the second transcription text includes a plurality of second candidate sentence groups, one second candidate sentence group is obtained by decoding one speech sentence, one second candidate sentence group includes a plurality of second candidate sentences, and the first transcription text is obtained by reordering the plurality of second candidate sentences in each second candidate sentence group in the second transcription text.

[0221] Among them, the creation module 141 is configured to create new candidate sentences by using the hot word list and the second candidate sentence group corresponding to the time after the current time in the second transcription text, and add the new candidate sentences to the corresponding second candidate sentence group to obtain a third candidate sentence group.

[0222] Specifically, the creation module 141 includes a mapping module 1411, a judgment module 1412, and an execution module 1413.

[0223] Among them, the mapping module 1411 is configured to construct at least one first mapping network by using the hot word list, and construct at least one second mapping network by using the second candidate sentence group corresponding to the time after the current time in the second transcription text; wherein, the first mapping network corresponds to the category of hot words one by one, the first mapping network includes at least one first mapping path, the first mapping path represents the mapping relationship between the hot word and its first pinyin sequence, and the input of the first mapping path is the first pinyin sequence and the output is the hot word; the second mapping network corresponds to the second candidate sentence one by one, the second mapping network includes a plurality of second mapping paths, the second mapping path represents the mapping relationship between the second candidate sentence and its second pinyin sequence, and the input of the second mapping path is the second candidate sentence and the output is the second pinyin sequence.

[0224] The judgment module 1412 is configured to judge whether there is a matching segment in the second pinyin sequence that satisfies the matching condition with the first pinyin sequence. The execution module 1413 is configured to, when the judgment result of the judgment module 1412 is yes, add the category label of the type of the corresponding first mapping network between the start node and the end node of the matching segment corresponding to the second mapping path; add a sub-path identical to the first mapping path of the corresponding matching segment between the start node and the end node of the corresponding category label; form a new mapping path with the part of the second mapping path that does not correspond to the matching segment and the sub-path, and use the text on the new mapping path as the new candidate sentence.

[0225] Among them, the sorting module 142 is configured to reorder the third candidate sentences included in the third candidate sentence group to obtain a second corrected text; the third candidate sentence is the same as the corresponding second candidate sentence or new candidate sentence.

[0226] Specifically, the sorting module 142 includes a scoring module 1421 and a calculation module 1422. Among them, the scoring module 1421 is used to obtain the language score of each third candidate sentence by using the hot word list and the first corrected text; the calculation module 1422 is used to take the sum value of the language score and the corresponding acoustic score as the final score; the acoustic score is obtained during the process of obtaining the first transcript, and one second candidate sentence corresponds to one acoustic score; and it is used to re-sort the third candidate sentences in the order of the size of the final score to obtain the second corrected text.

[0227] Among them, the scoring module 1421 includes a language module 14211 and a weighting module 14212.

[0228] The language module 14211 is used to obtain the patch language score of the third candidate sentence by using the patch language model and the general language model, and to obtain the main language score of the third candidate sentence by using the class language model and the general language model. Among them, the inputs of the patch language model, the class language model, and the general language model are all a sentence, and the outputs are all the probabilities of the input sentence appearing. The patch language model is trained by using the training text obtained from the hot word list and the first corrected text, the class language model is trained by using the training text including category labels, and the general language model is trained by using the training text composed of general sentences without category labels. The weighting module 14212 is used to take the value obtained by weighted summation of the patch language score and the main language score as the language score.

[0229] Please continue to refer to Figure 13 that the language module 14211 further includes a patch language module 14211A and a main language module 14211B.

[0230] The patch language module 14211A is specifically used to screen out the corrected candidate sentences containing hot words from the first corrected text, copy them several times and then add them to the first corrected text to obtain the patch language training text; use the patch language training text to train the patch language model; input the third candidate sentences into the patch language model and the general language model respectively, and take the value obtained by weighted summation of the output of the patch language model and the output of the general language model as the patch language score.

[0231] The main language module 14211B is specifically used to, in response to the third candidate sentence being the same as the corresponding second candidate sentence, input the second candidate sentence into the class language model and the general language model respectively; in response to the third candidate sentence being the same as the corresponding new candidate sentence, replace the hot words in the new candidate sentence with the corresponding category labels, then input it into the class language model, and input the new candidate sentence into the general language model; and take the outputs of the class language model and the general language model as the class language score and the general language score of the third candidate sentence respectively; take the value obtained by weighted summation of the class language score and the general language score as the main language score.

[0232] Among them, the replacement module 143 is used to replace the first candidate sentence group corresponding to after the current moment in the first transcription text with the second corrected text to obtain a subtitle corrected text.

[0233] In some embodiments, the replacement module 143 is further used to display the third candidate sentence with the highest score in each third candidate sentence group of the second corrected text as the subtitle of the corresponding speech sentence according to the time axis of the audio file.

[0234] This embodiment reorders the third candidate sentences included in the third candidate sentence group based on language scores and acoustic scores, and can give a higher score incentive to the new candidate sentences related to hot words, improve the scores of the corresponding third candidate sentences, and make them occupy a front position in the reordering, thereby reducing the probability of the recurrence of hot word-related errors, improving the transcription accuracy of the transcription system, and reducing the workload of subtitle correction.

[0235] In addition, the present application also provides a subtitle production device. Please refer to Figure 14 , Figure 14 , which is a schematic structural diagram of an embodiment of the subtitle production device of the present application. The subtitle production device includes a memory 1401 and a processor 1402. The memory 1401 stores program instructions, and the processor 1402 can execute the program instructions to implement the subtitle production method described in any of the above embodiments. For details, please refer to any of the above embodiments, and details will not be described here again.

[0236] In addition, the present application also provides a computer-readable storage medium. Please refer to Figure 15 , Figure 15 , which is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The storage medium 150 stores program instructions 151, and the program instructions 151 can be executed by a processor to implement the subtitle production method described in any of the above embodiments. For details, please refer to any of the above embodiments, and details will not be described here again.

[0237] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A subtitle production method, characterized in that, Including: Obtaining a first transcript text corresponding to an audio file; Performing text correction on the part of the first transcript text before the current moment to obtain a first corrected text; Obtaining historical correction information by using the first corrected text; wherein, the historical correction information corresponds to a hot word list, screening triples whose first part-of-speech tag or second part-of-speech tag belongs to a preset part-of-speech list, and using the first word in the triples as the hot words in the hot word list, the triples include word pairs and the number of times they appear, the word pairs include a first word and a second word, the first segmented text includes multiple discrete first words and the first part-of-speech tags corresponding to the first words one by one, the second segmented text includes multiple discrete second words and the second part-of-speech tags corresponding to the second words one by one, the first segmented text is obtained by preprocessing the first corrected text, the second segmented text is obtained by preprocessing the part of the first transcript text corresponding to the first corrected text, and the first segmented text and the second segmented text are aligned so that the corresponding first words and second words are aligned to form word pairs; Updating the part of the first transcript text after the current moment by using the historical correction information to obtain a subtitle corrected text; Wherein, the first transcript text includes multiple first candidate sentence groups, and the first transcript text is obtained by reordering multiple second candidates in each second candidate sentence group in the second transcript text, the second transcript text includes the first candidate sentence groups and multiple second candidate sentence groups, and the second candidate sentence groups correspond to the speech sentences in the audio file; The step of updating the part of the first transcript text after the current moment by using the historical correction information to obtain a subtitle corrected text includes: creating new candidate sentences by using the hot word list and the second candidate sentence groups in the second transcript text after the current moment, and adding the new candidate sentences to the corresponding second candidate sentence groups to obtain third candidate sentence groups; reordering the third candidate sentences included in the third candidate sentence groups to obtain a second corrected text; and replacing the first candidate sentence groups in the first transcript text after the current moment with the second corrected text to obtain the subtitle corrected text.

2. The subtitle production method according to claim 1, characterized in that, The step of obtaining historical correction information by using the first corrected text includes: Preprocessing the first corrected text to obtain a first segmented text, and preprocessing the part of the first transcript text corresponding to the first corrected text to obtain a second segmented text; aligning the first segmented text and the second segmented text so that the corresponding first words and second words are aligned to form word pairs; Traversing all the word pairs, in response to the first word and the second word in the current word pair being different, and in response to the current word pair meeting a preset condition, generating a triple according to the current word pair; the triple includes a word pair and the number of times it appears; Obtaining the historical correction information by using a triple list composed of all the triples.

3. The subtitle production method according to claim 2, characterized in that, The first corrected text includes at least one correction candidate sentence. The step of responding to the current word pair satisfying a preset condition includes: responding to the first word in the current word pair being a multi-word term; or, responding to the first word in the current word pair being located at the beginning of the corresponding correction candidate sentence and the first word after it not being a preset stop word; or, responding to the first word in the current word pair being located at the end of the corresponding correction candidate sentence and the first word before it not being a preset stop word; or, responding to the first word in the current word pair being located in the middle of the corresponding correction candidate sentence and determining that the current word pair satisfies the preset condition.

4. The subtitle production method according to claim 3, characterized in that, The step of generating a triple according to the current word pair includes: responding to the first word in the current pair being a multi-word term, and forming the triple by the current word pair and the corresponding cumulative occurrence times; responding to the first word in the current pair not being a multi-word term, obtaining a current extended word pair according to the first word and its adjacent words, and forming the triple by the current extended word pair and the corresponding cumulative occurrence times.

5. The subtitle production method according to claim 2, characterized in that, The step of obtaining the historical correction information by using the triple list composed of all the triples includes: screening out the triples in the triple list whose first part-of-speech tag or second part-of-speech tag belongs to a preset part-of-speech list; taking the first word in the screened triples as hot words, and taking the hot word list composed of all the hot words as the historical correction information.

6. The subtitle production method according to claim 5, characterized in that, The step of creating a new candidate sentence by using the hot word list and the second candidate sentence group corresponding to the second transcript text after the current moment includes: constructing at least one first mapping network by using the hot word list, and constructing at least one second mapping network by using the second candidate sentence group corresponding to the second transcript text after the current moment; wherein, the first mapping network corresponds to the category of the hot words one by one, the first mapping network includes at least one first mapping path, the first mapping path represents the mapping relationship between the hot word and its first pinyin sequence, and the input of the first mapping path is the first pinyin sequence and the output is the hot word; the second mapping network corresponds to the second candidate sentence one by one, the second mapping network includes multiple second mapping paths, the second mapping path represents the mapping relationship between the second candidate sentence and its second pinyin sequence, and the input of the second mapping path is the second candidate sentence and the output is the second pinyin sequence; judging whether there is a matching segment in the second pinyin sequence that satisfies the matching condition with the first pinyin sequence; if so, adding a category label between the start node and the end node of the second mapping path corresponding to the matching segment; the category label represents the type of the corresponding first mapping network; adding a new sub-path between the start node and the end node corresponding to the category label; the sub-path is the same as the first mapping path corresponding to the matching segment; Combine the part of the second mapping path that does not correspond to the matching segment and the sub-path to form a new mapping path, and use the text on the new mapping path as the new candidate sentence.

7. The subtitle production method according to claim 6, characterized in that, The step of reordering the third candidate sentences included in the third candidate sentence group to obtain the second corrected text includes: Obtain the language score of each third candidate sentence by using the hot word list and the first corrected text; Use the sum of the language score and the corresponding acoustic score as the final score; the acoustic score is obtained during the process of obtaining the first transcription text, and one second candidate sentence corresponds to one acoustic score; Reorder the third candidate sentences included in the third candidate sentence group in the order of the magnitude of the final score to obtain the second corrected text.

8. The subtitle production method according to claim 7, characterized in that The step of obtaining the language score of each third candidate sentence by using the hot word list and the first corrected text includes: Obtain the patch language score of the third candidate sentence by using the patch language model and the general language model, and obtain the main language score of the third candidate sentence by using the class language model and the general language model; the inputs of the patch language model, the class language model and the general language model are all a sentence, and the outputs are all the probabilities of the input sentence appearing. The patch language model is trained by using the training text obtained from the hot word list and the first corrected text, the class language model is trained by using the training text including the category label, and the general language model is trained by using the training text composed of ordinary sentences without the category label; Use the weighted sum of the patch language score and the main language score as the language score.

9. The subtitle production method according to claim 8, characterized in that The first corrected text includes at least one corrected candidate sentence. The step of obtaining the patch language score of the third candidate sentence by using the patch language model and the general language model includes: Select the corrected candidate sentences containing the hot words from the first corrected text, copy them several times and add them to the first corrected text to obtain the patch language training text; Train the patch language model by using the patch language training text; Input the third candidate sentences into the patch language model and the general language model respectively, and use the weighted sum of the outputs of the patch language model and the general language model as the patch language score.

10. The subtitle production method according to claim 8, characterized in that The step of obtaining the main language score of the third candidate sentence by using the class language model and the general language model includes: In response to the third candidate sentence being the same as the corresponding second candidate sentence, input the second candidate sentence into the class language model and the general language model respectively; in response to the third candidate sentence being the same as the corresponding new candidate sentence, replace the hot word in the new candidate sentence with the corresponding category label, then input it into the class language model, and input the new candidate sentence into the general language model; and use the outputs of the class language model and the general language model as the class language score and the general language score of the third candidate sentence respectively; The value obtained by weighted summation of the pseudo-language part and the ordinary language part is used as the main language part.

11. The subtitle production method according to claim 5, characterized in that After the step of using the historical correction information to update the part of the first transcription text corresponding to after the current moment to obtain a subtitle correction text, the method further includes: For each third candidate sentence group of the second correction text, the third candidate sentence with the highest score is displayed as the subtitle of the corresponding speech sentence according to the time axis of the audio file.

12. A subtitle production device, characterized in that It includes: A transcription module, configured to obtain a first transcription text corresponding to an audio file; A correction module, configured to perform text correction on the part of the first transcription text corresponding to before the current moment to obtain a first correction text An extraction module, configured to obtain historical correction information by using the first correction text; wherein, the historical correction information corresponds to a hot word list, screening triples whose first part-of-speech tag or second part-of-speech tag belongs to a preset part-of-speech list, and using the first word in the triples as the hot words in the hot word list, the triples include word pairs and the number of times they appear, the word pairs include a first word and a second word, the first segmented text includes multiple discrete first words and the first part-of-speech tags corresponding to the first words one by one, the second segmented text includes multiple discrete second words and the second part-of-speech tags corresponding to the second words one by one, the first segmented text is obtained by preprocessing the first correction text, the second segmented text is obtained by preprocessing the part of the first transcription text corresponding to the first correction text, and the first segmented text and the second segmented text are aligned so that the corresponding first words and second words are aligned to form word pairs; An update module, configured to update the part of the first transcription text corresponding to after the current moment by using the historical correction information to obtain a subtitle correction text; Wherein, the first transcription text includes multiple first candidate sentence groups, the first transcription text is obtained by reordering multiple second candidate sentences in each second candidate sentence group of a second transcription text, the second transcription text includes the first candidate sentence groups and multiple second candidate sentence groups, and the second candidate sentence groups correspond to speech sentences in the audio file; The step of using the historical correction information to update the part of the first transcription text corresponding to after the current moment to obtain a subtitle correction text includes: creating new candidate sentences by using the hot word list and the second candidate sentence groups corresponding to after the current moment in the second transcription text, and adding the new candidate sentences to the corresponding second candidate sentence groups to obtain third candidate sentence groups; reordering the third candidate sentences included in the third candidate sentence groups to obtain a second correction text; replacing the first candidate sentence groups corresponding to after the current moment in the first transcription text with the second correction text to obtain the subtitle correction text.

13. A subtitle production device, characterized in that It includes a memory and a processor, the memory stores program instructions, and the processor can execute the program instructions to implement the subtitle making method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that Program instructions are stored on the storage medium, and the program instructions can be executed by a processor to implement the subtitle production method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Automatic error correction method for real-time court hearing speech recognition, storage medium and computing device

    CN108984529A