Subtitle correction method and related device

CN122693641APending Publication Date: 2026-09-04TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610920784.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

但该强对齐纠错算法适合音频内容和提供的文本完全一致的情况,在音频内容和提供的文本不一致的情况下进行强制对齐,会出现音频文本不对应的情况

Benefits of technology

[0080] This application embodiment generates a character matching score matrix between the sentence to be corrected and the target reference sentence based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence. Based on the character matching score matrix, the sentence to be corrected and the target sentence are aligned to obtain a character-level alignment path between the sentence to be corrected and the reference sentence. Using the character-level alignment path, the target character in the sentence to be corrected and the reference character corresponding to the target character are determined. Based on the reference character, the target character in the sentence to be corrected is corrected to obtain the corrected subtitle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122693641A_ABST
    Figure CN122693641A_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a subtitle correction method, which is used for realizing character-level alignment and correction between a to-be-corrected sentence in a to-be-corrected subtitle and a target reference sentence. The method comprises the following steps: screening a target reference sentence corresponding to the to-be-corrected sentence of the to-be-corrected subtitle from a reference text; generating a character matching score matrix between the to-be-corrected sentence and the target reference sentence based on a matching relationship between each subtitle character in the to-be-corrected sentence and each reference character in the target reference sentence; performing sequence alignment processing on the to-be-corrected sentence and the target reference sentence based on the character matching score matrix, so as to obtain a character-level alignment path between the to-be-corrected sentence and the target reference sentence; determining a target character in the to-be-corrected sentence and a reference character corresponding to the target character based on the character-level alignment path; and performing correction processing on the target character in the to-be-corrected sentence based on the reference character, so as to obtain a corrected subtitle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for correcting subtitle errors. Background Technology

[0002] With the development of AI technology, in audio scenarios (such as music and audiobooks), AI can directly convert sound into text, thereby making it convenient to obtain the text corresponding to the sound.

[0003] Currently, in scenarios where AI is used to convert sound into text, there are often issues such as unclear sound, pauses, or homophones, which cause discrepancies between the text converted by AI and the original sound.

[0004] To address this issue, a strong alignment error correction algorithm exists for subtitle correction. This algorithm precisely matches the timestamps of the audio signal with the corresponding characters / words in the text. Its core principle is to predict the correspondence between audio frames and text units (phonemes or characters) using an acoustic model. However, this strong alignment error correction algorithm is suitable when the audio content and the provided text are completely identical. Forcing alignment when the audio content and the provided text are inconsistent will result in mismatches between the audio and text. Summary of the Invention

[0005] This invention provides a subtitle correction method and related apparatus, which are used to achieve character-level alignment and correction between the subtitle to be corrected and the target reference statement when the subtitle to be corrected is not strongly aligned with the target reference statement.

[0006] The first aspect of this application provides a subtitle error correction method, including:

[0007] Select the target reference sentences from the reference text that correspond to the sentences to be corrected in the subtitles;

[0008] Based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence, a character matching score matrix is ​​generated between the sentence to be corrected and the target reference sentence; the character matching score matrix is ​​used to characterize the degree of matching between each subtitle character and each reference character.

[0009] Based on the character matching score matrix, the statement to be corrected and the target reference statement are subjected to sequence alignment processing to obtain the character-level alignment path between the statement to be corrected and the target reference statement;

[0010] Based on the character-level alignment path, determine the target character in the statement to be corrected and the reference character corresponding to the target character;

[0011] Based on the reference characters, the target characters in the sentence to be corrected are corrected to obtain the corrected subtitles.

[0012] As an optional embodiment, before generating the character matching score matrix, the method further includes:

[0013] Obtain at least two of the word embedding, position embedding, and pinyin embedding of each subtitle character in the sentence to be corrected, and fuse at least two of the word embedding, position embedding, and pinyin embedding of each subtitle character to obtain the first fusion feature of each subtitle character;

[0014] Obtain at least two of the word embedding, position embedding, and pinyin embedding of each reference character in the target reference statement, and fuse at least two of the word embedding, position embedding, and pinyin embedding of each reference character to obtain a second fused feature for each reference character;

[0015] Multiple first fusion features of the statement to be corrected and multiple second fusion features of the reference statement are concatenated into a feature sequence matrix, and the multiple first fusion features and multiple second fusion features are respectively identified in the feature sequence matrix;

[0016] The first fusion feature and the second fusion feature in the feature sequence matrix are encoded using a multi-head attention mechanism to obtain a deep feature vector matrix that represents the context information of the first fusion feature and the context information of the second fusion feature.

[0017] As an optional embodiment, generating a character matching score matrix between the sentence to be corrected and the target reference sentence based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence includes:

[0018] Based on the first fusion feature identifier and the second fusion feature identifier, the deep feature vector matrix is ​​split into a first deep feature vector matrix for representing the context information of each first fusion feature and a second deep feature vector matrix for representing the context information of each second fusion feature.

[0019] The first deep feature vector matrix and the second deep feature vector matrix are processed to obtain a character matching score matrix. The element in the i-th row and j-th column of the character matching score matrix is ​​used to characterize the similarity between the i-th subtitle character in the sentence to be corrected and the j-th reference character in the reference sentence.

[0020] As an optional embodiment, the step of performing sequence alignment processing on the statement to be corrected and the target reference statement based on the character matching score matrix to obtain the character-level alignment path between the statement to be corrected and the target reference statement includes:

[0021] Construct an initial dynamic programming matrix and a preset empty space penalty;

[0022] Based on the vacancy penalty, fill the boundary vacancy penalty sequence into the initialized dynamic programming matrix;

[0023] Based on the first formula for calculating the matching path score, the second formula for calculating the deletion path score, the third formula for calculating the insertion path score, the character matching score matrix, and the empty space penalty, the maximum score corresponding to the target element in the initialized dynamic programming matrix on the matching path, deletion path, and insertion path is calculated respectively, and the maximum score is determined as the element value of the target element, thus obtaining the corresponding dynamic programming matrix. The element value in the i-th row and j-th column of the dynamic programming matrix is ​​used to represent the maximum score of the first i subtitle characters in the sentence to be corrected and the first j reference characters in the target reference sentence on the matching path, deletion path, and insertion path.

[0024] Based on the backtracking path formula and the character matching score matrix, the corresponding elements are traversed sequentially from the lower right element of the dynamic programming matrix along a preset direction until the boundary element of the dynamic programming matrix is ​​reached, thereby obtaining the character-level alignment path between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence.

[0025] As an optional embodiment, the backtracking path formula includes a character matching judgment formula, which is used to characterize that the subtitle characters in the sentence to be corrected are completely aligned with the corresponding reference characters in the reference sentence;

[0026] The step of traversing the corresponding elements sequentially from the bottom right corner of the dynamic programming matrix along a preset direction based on the backtracking path formula and the character matching score matrix includes:

[0027] Based on the character matching judgment formula and the character matching score matrix, if it is determined that the subtitle character in the sentence to be corrected is completely aligned with the corresponding reference character in the reference sentence, then backtracking continues from the lower right element of the dynamic programming matrix along the diagonal direction, and the subtitle character in the sentence to be corrected is aligned with the corresponding reference character in the reference sentence in the alignment result.

[0028] As an optional embodiment, the backtracking path formula includes a reference text character missing judgment formula, which is used to characterize the missing corresponding subtitle character in the target reference statement;

[0029] The step of traversing the corresponding elements sequentially from the bottom right corner of the dynamic programming matrix along a preset direction based on the backtracking path formula and the character matching score matrix includes:

[0030] Based on the missing reference text character judgment formula and the character matching score matrix, if it is determined that the target reference statement is missing a corresponding subtitle character, then backtracking continues from the lower right corner element of the dynamic programming matrix upwards, and the subtitle character is added to the alignment result, and an empty character marker is added at the reference character position corresponding to the subtitle character.

[0031] As an optional embodiment, the backtracking path formula includes a subtitle character missing judgment formula, which is used to characterize the absence of a corresponding reference character in the statement to be corrected;

[0032] The step of traversing the corresponding elements sequentially from the bottom right corner of the dynamic programming matrix along a preset direction based on the backtracking path formula and the character matching score matrix includes:

[0033] Based on the missing subtitle character judgment formula and the character matching score matrix, if it is determined that the corresponding reference character is missing in the sentence to be corrected, then backtracking continues from the lower right element of the dynamic programming matrix to the left, and the reference character is added to the alignment result, and an empty character mark is added at the position of the subtitle character corresponding to the reference character.

[0034] As an optional embodiment, the method further includes:

[0035] Based on the alignment results, after generating the character-level alignment path for each character in the statement to be corrected and each reference character in the reference statement, the matching score of each character pair in the alignment path is recorded.

[0036] As an optional embodiment, the step of filtering out the target reference statement corresponding to the statement to be corrected in the subtitle from the reference text includes:

[0037] Based on at least one of text similarity, pinyin similarity, and contextual similarity, target reference paragraphs with a similarity greater than a first preset threshold to the subtitle to be corrected are selected from the reference text.

[0038] Based on each sentence to be corrected in the subtitles to be corrected, a multi-level matching strategy is used to select target reference sentences corresponding to each sentence to be corrected from the target reference paragraphs. The multi-level matching strategy includes pinyin similarity matching and semantic similarity matching.

[0039] As an optional embodiment, the step of filtering target reference paragraphs from the reference text based on at least one of text similarity, pinyin similarity, and contextual similarity, with a similarity greater than a first preset threshold to the subtitle to be corrected, includes:

[0040] The reference text and the subtitle to be corrected are respectively divided into multiple text segments to obtain multiple segmented reference text segments and multiple segmented subtitle segments to be corrected;

[0041] Generate a monotonically changing first position sequence identifier for the plurality of reference text segments, and generate a monotonically changing second position sequence identifier for the plurality of subtitle segments to be corrected;

[0042] Based on at least one of text similarity, pinyin similarity, and contextual similarity, target reference text segments and target subtitle segments to be corrected are selected from the plurality of reference text segments and the plurality of subtitle segments to be corrected, with a similarity greater than the second preset threshold.

[0043] Obtain the first position number i1 corresponding to the start position of the target subtitle segment to be corrected in the subtitle to be corrected, and the second position number i2 corresponding to the end position of the target subtitle segment to be corrected in the subtitle to be corrected;

[0044] Obtain the third position number j1 corresponding to the start position of the target reference text fragment in the reference text and the fourth position number j2 corresponding to the end position of the target reference text fragment in the reference text;

[0045] The reference text fragments at positions j1-i1 are determined as the starting positions of the target subtitle fragment to be corrected in the reference text, or the reference text fragments at positions j2-i2 are determined as the starting positions of the target subtitle fragment to be corrected in the reference text.

[0046] As an optional embodiment, after obtaining multiple target reference statements corresponding to multiple statements to be corrected, the method further includes:

[0047] Based on multiple consecutive statements to be corrected, determine whether the position identifiers of multiple target reference statements corresponding to the multiple consecutive statements to be corrected in the target reference paragraph are monotonically increasing.

[0048] If not, then obtain the abnormal reference statements whose position identifiers are not monotonically increasing among the multiple target reference statements, and delete the abnormal reference statements.

[0049] As an optional embodiment, the step of using a multi-level matching strategy to filter out the target reference statements corresponding to each statement to be corrected from the target reference paragraph includes:

[0050] The sentence to be corrected and the target reference paragraph are converted into pinyin sequences respectively, to obtain the pinyin sequence of the sentence to be corrected and the pinyin sequence of the target reference paragraph;

[0051] If the pinyin sequence of the target reference paragraph contains the pinyin sequence of the statement to be corrected, then a reference statement containing the pinyin sequence of the statement to be corrected is determined in the target reference paragraph, and the reference statement is determined as the target reference statement corresponding to the statement to be corrected.

[0052] As an optional embodiment, the method further includes:

[0053] If the pinyin sequence of the target reference paragraph does not contain the pinyin sequence of the sentence to be corrected, then the sentence to be corrected is converted into a first vector, and each reference sentence in the target reference paragraph is converted into a second vector, resulting in multiple second vectors;

[0054] Based on multiple similarity values ​​between the first vector and multiple second vectors, a target reference sentence corresponding to the sentence to be corrected is determined from the target reference paragraph.

[0055] As an optional embodiment, the step of performing error correction processing on the target character in the statement to be corrected based on the reference character includes:

[0056] A multidimensional confidence scoring mechanism is used to correct the target characters in the statement to be corrected using a hierarchical decision strategy. The multidimensional confidence scoring mechanism includes at least two of the following: character-level matching score, sentence-level matching score, empty space ratio, neighborhood reliability score, and harmful domain detection score. The hierarchical decision strategy includes a first confidence correction strategy, a second confidence correction strategy, and a third confidence correction strategy. The error correction threshold of the first confidence correction strategy is less than the error correction threshold of the second confidence correction strategy, and the error correction threshold of the second confidence correction strategy is less than the error correction threshold of the third confidence correction strategy.

[0057] As an optional embodiment, before using a multi-dimensional confidence scoring mechanism to correct the target characters in the statement to be corrected using a hierarchical decision-making strategy, the method further includes:

[0058] Based on the matching score of each character pair in the alignment path, obtain the character-level matching score between the statement to be corrected and the target reference statement; and / or,

[0059] Calculate the semantic similarity between the statement to be corrected and the target reference statement to obtain a sentence-level matching score between them; and / or,

[0060] The percentage of empty characters in the alignment path is calculated, and the empty ratio is obtained based on the percentage of empty characters; and / or,

[0061] Calculate the character pair matching scores of each surrounding character within a preset distance of the target character, and determine the neighborhood reliability score of the target character based on the character pair matching scores of each surrounding character; and / or,

[0062] Identify the punctuation marks in the target reference statement, and determine whether there is a sentence break at the corresponding position of the statement to be corrected based on the punctuation marks.

[0063] If there is no punctuation break at the corresponding position of the statement to be corrected, then it is determined that there is a harmful neighborhood in the statement to be corrected; or, if there is a punctuation break in the statement to be corrected, but the character pair matching score corresponding to at least one character at a preset distance from the punctuation break is lower than a preset threshold, then it is determined that there is a harmful neighborhood in the statement to be corrected.

[0064] As an optional embodiment, after using a multi-dimensional confidence scoring mechanism to correct the target characters in the statement to be corrected using a hierarchical decision-making strategy, the method further includes:

[0065] Identify the first and last characters in the statement to be corrected;

[0066] If the first and last characters of the statement to be corrected are not aligned with the corresponding target reference statement, then calculate the first perplexity of the statement to be corrected before correction and the second perplexity of the statement to be corrected after correction.

[0067] If the first perplexity is less than the second perplexity, then the correction of the target character in the statement to be corrected is abandoned;

[0068] If the first perplexity is greater than the second perplexity, then error correction is received for the target character in the statement to be corrected.

[0069] As an optional embodiment, after determining the sentence-level matching score between the statement to be corrected and the target reference statement, the method further includes:

[0070] Identify the polyphonic characters in the sentence to be corrected, and the first pronunciation of the polyphonic characters in the sentence to be corrected;

[0071] Identify the polyphonic characters in the statement to be corrected, the corresponding polyphonic characters in the target reference statement, and the second pronunciation of the polyphonic characters in the target reference statement;

[0072] Calculate the similarity between the first pronunciation and the second pronunciation;

[0073] If the similarity is less than a preset threshold, the sentence-level matching score is reduced.

[0074] As an optional embodiment, after using a multi-dimensional confidence scoring mechanism to correct the target characters in the statement to be corrected using a hierarchical decision-making strategy, the method further includes:

[0075] The first timestamp of the corrected statement is mapped to the second timestamp of the statement to be corrected before the correction, so that the timestamp of the corrected statement is consistent with the timestamp of the statement to be corrected before the correction.

[0076] A second aspect of this application provides a computer device including a processor, which, when executing a computer program stored in a memory, implements the subtitle correction method provided in the first aspect of this application.

[0077] A third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the subtitle correction method provided in the first aspect of this application.

[0078] The fourth aspect of this application provides a computer program product having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the subtitle correction method provided in the first aspect of this application.

[0079] As can be seen from the above technical solutions, the embodiments of the present invention have the following advantages:

[0080] This application embodiment generates a character matching score matrix between the sentence to be corrected and the target reference sentence based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence. Based on the character matching score matrix, the sentence to be corrected and the target sentence are aligned to obtain a character-level alignment path between the sentence to be corrected and the reference sentence. Using the character-level alignment path, the target character in the sentence to be corrected and the reference character corresponding to the target character are determined. Based on the reference character, the target character in the sentence to be corrected is corrected to obtain the corrected subtitle.

[0081] Because the embodiments of this application can utilize the character matching score matrix to obtain the character-level alignment path between the statement to be corrected and the reference statement, even when the statement to be corrected and the reference statement are not forcibly aligned, the target character in the statement to be corrected and the reference character corresponding to the target character can be determined based on the character-level alignment path, and the error correction processing can be performed on the target character in the statement to be corrected based on the reference character, thereby improving the accuracy of error correction of the statement to be corrected. Attached Figure Description

[0082] Figure 1 This is a schematic diagram of the subtitle correction system in the embodiments of this application;

[0083] Figure 2 This is a schematic diagram of one embodiment of the subtitle correction method in this application;

[0084] Figure 3 for Figure 2 Detailed steps of step 201 in the embodiment;

[0085] Figure 4 for Figure 3 Detailed steps of step 301 in the embodiment;

[0086] Figure 5 for Figure 3 Detailed steps of step 302 in the embodiment;

[0087] Figure 6 for Figure 2 Detailed steps of step 202 in the embodiment;

[0088] Figure 7 for Figure 2 Another detailed step of step 202 in the embodiment;

[0089] Figure 8 for Figure 2 Detailed steps of step 203 in the embodiment;

[0090] Figure 9 for Figure 2 Detailed steps of step 204 in the embodiment;

[0091] Figure 10 This is a schematic diagram of another embodiment of the subtitle correction method in this application. Detailed Implementation

[0092] This invention provides a subtitle correction method for achieving character-level alignment and correction between the subtitle to be corrected and the target reference statement when the subtitle to be corrected is not strongly aligned with the reference statement.

[0093] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0094] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0095] This application provides a subtitle correction method. The general principle of this method is as follows: First, a target reference statement corresponding to the subtitle to be corrected is selected from the reference text. Second, a character matching score matrix is ​​generated between the subtitle characters in the subtitle to be corrected and the target reference statement based on the matching relationship between each reference character in the target reference statement and each subtitle character in the target reference statement. Third, the character matching score matrix characterizes the degree of matching between each subtitle character and each reference character. Fourth, based on the character matching score matrix, sequence alignment is performed between the subtitle to be corrected and the target reference statement to obtain a character-level alignment path between them. Fifth, based on the character-level alignment path, the target character in the subtitle to be corrected and its corresponding reference character are determined. Finally, the target character in the subtitle to be corrected is corrected based on the reference character to obtain the corrected subtitle. The subtitle correction method provided in this application can obtain the character-level alignment path between the sentence to be corrected and the reference sentence in the subtitle to be corrected by using the character matching score matrix between the sentence to be corrected and the reference sentence. Then, using the character-level alignment path, the target character in the sentence to be corrected and the reference character corresponding to the target character are determined. The target character is then corrected using the reference character. This achieves character-level correction and alignment between the sentence to be corrected and the target reference sentence in the subtitle to be corrected when they are not strongly aligned.

[0096] To better implement the above-mentioned subtitle correction method, this application provides a subtitle correction system. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of the architecture of a subtitle correction system provided in an embodiment of this application. The subtitle correction system may include at least one terminal device 101 and a server 102. Different types of applications may be installed on the terminal device 101, such as instant messaging applications, live streaming applications, music applications, conferencing applications, etc. The terminal device 101 may be a smartphone, tablet, laptop, desktop computer, smart vehicle, etc. The server 102 may be used to store subtitle data generated by different types of applications on the terminal device 101. The server 102 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc.

[0097] The aforementioned subtitle correction method is executed by either terminal device 101 or server 102. When the subtitle correction method is executed by terminal device 101, the subtitles to be corrected generated by terminal device 101 in different types of applications can be included in the server. When terminal device 101 needs to correct the subtitles to be corrected, it can obtain the subtitles to be corrected from server 102. After obtaining the subtitles to be corrected from server 102, terminal device 101 can obtain the target reference statement corresponding to the subtitle to be corrected from the reference text. Based on the matching relationship between each subtitle character in the subtitle to be corrected and each reference character in the target reference statement, a character matching score matrix is ​​generated between the subtitle to be corrected and the target reference statement. Based on the character matching score matrix, the character-level alignment path between the subtitle to be corrected and the target reference statement is obtained. Using the character-level alignment path, the target character in the subtitle to be corrected and the reference character corresponding to the target character are determined. The target character is then corrected using the reference character.

[0098] For ease of understanding, the subtitle correction method in this application is described below. Please refer to [link / reference]. Figure 2 One embodiment of the subtitle error correction method in this application includes:

[0099] 201. Select the target reference sentences from the reference text that correspond to the sentences to be corrected in the subtitles;

[0100] When outputting audio to various audio output terminals (such as music platforms, audiobook platforms, etc.), an AI model can be used to convert the audio output to AI subtitles, thereby obtaining subtitles corresponding to the audio, which are also the subtitles to be corrected in this application. Simultaneously with the audio output, each audio also has a corresponding reference text. For example, when an audio platform plays a song, the reference text is the lyrics of that song; when an audiobook platform plays an audiobook, the reference text is the original text of that audiobook.

[0101] To avoid discrepancies between AI-generated subtitles and the reference text, this application compares the AI-generated subtitles (i.e., the subtitles to be corrected in this application) with the reference text to correct misplaced characters. Specifically, the content of the subtitles to be corrected is often quite extensive (e.g., containing multiple paragraphs). Therefore, when comparing the subtitles to be corrected with the reference text, this application needs to compare the sentences to be corrected in the subtitles with the corresponding target reference sentences sentence by sentence to correct erroneous characters in the sentences to be corrected.

[0102] In order to achieve sentence-by-sentence comparison between the statement to be corrected and the corresponding target reference statement, this application needs to first filter out the target reference statement corresponding to the statement to be corrected in the subtitle from the reference text. In the process of filtering the target reference statement, the reference statement with a similarity greater than a preset threshold to the statement to be corrected can be filtered out from the reference text first. Then, step 202 is performed for the statement to be corrected and the target reference statement respectively.

[0103] Specifically, the process of selecting target reference sentences from the reference text can be based on sentence-by-sentence comparison, or by comparing paragraphs first and then sentence-by-sentence. There are no specific restrictions on the process of selecting reference sentences from the reference text.

[0104] 202. Based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence, generate a character matching score matrix between the sentence to be corrected and the target reference sentence; the character matching score matrix is ​​used to characterize the degree of matching between each subtitle character and each reference character;

[0105] After obtaining the statement to be corrected and the reference statement, this application can generate a character matching score matrix between the statement to be corrected and the target reference statement based on the matching relationship between the subtitle characters in the statement to be corrected and each reference character in the target reference statement. The character matching score matrix is ​​used to characterize the degree of matching between each subtitle character and each reference character.

[0106] Specifically, when obtaining the matching relationship between the subtitle characters in the sentence to be corrected and each reference character in the target reference sentence, the sentence to be corrected and the reference sentence can be encoded separately to turn the two sentences into a computable encoded representation, and the pairwise matching scores at the character / token level can be calculated to form a matching score matrix.

[0107] Specifically, when encoding the statement to be corrected and the reference statement, the statement to be corrected and the reference statement can be split into characters or encoded and mapped to obtain character sequences. Then, each character of the statement to be corrected and each character of the reference statement are traversed, and the matching score between the two is calculated. The specific matching rule can be that when the two characters are the same, a high score (such as 1.0) is given, and when the two characters are different, a low score (such as 0.0) is given, thereby generating an M*N matching score matrix, where M is the length of the statement to be corrected and N is the length of the reference statement.

[0108] 203. Based on the character matching score matrix, perform sequence alignment processing on the statement to be corrected and the target reference statement to obtain the character-level alignment path between the statement to be corrected and the reference statement;

[0109] After obtaining the matching score matrix, the optimal path can be solved based on the matching score using a dynamic programming algorithm. This involves performing sequence alignment processing on the statement to be corrected and the target reference statement. The optimal path satisfies the following conditions: the character matching degree is the highest, the character order in the statement to be corrected and the reference statement remains unchanged, and insertion, deletion, and replacement operations are allowed in this path. The output character-level alignment path needs to clearly indicate the position of each character in the statement to be corrected in the reference statement, as well as which characters in the statement to be corrected are redundant characters and which are missing characters.

[0110] Specifically, the dynamic programming algorithm in this application embodiment may be such as Needleman-Wunsch global alignment algorithm, Smith-Waterman local alignment algorithm, or DTW dynamic time warping algorithm, etc. There are no specific restrictions on the implementation method of the dynamic programming algorithm here.

[0111] 204. Based on the character-level alignment path, determine the target character in the statement to be corrected and the reference character corresponding to the target character;

[0112] After obtaining the character-level alignment path between the statement to be corrected and the target reference statement, the target character in the statement to be corrected and the reference character corresponding to the target character in the target reference statement can be determined based on the path. The target character in the statement to be corrected can be any character in the statement to be corrected. Since the character-level alignment path is calculated under non-strong alignment between the statement to be corrected and the target reference statement in this embodiment, the reference character corresponding to the target character in the character-level alignment path may be a character or it may be empty. In the case of "empty", it means that the target character is an extra character.

[0113] 205. Based on the reference characters, perform error correction processing on the target characters in the statement to be corrected to obtain the corrected subtitles.

[0114] After obtaining the target character and the corresponding reference character, the error correction process is performed on the target character in the statement to be corrected based on the reference character. If the target character in the statement to be corrected is different from the reference character in the target reference statement, the character is identified as an error character and replaced with the correct character at the corresponding position in the reference statement. If there are extra characters in the statement to be corrected, the character is deleted. If there are missing characters in the statement to be corrected, the correct character is added to the statement to be corrected to complete the correction of the statement to be corrected.

[0115] This application embodiment generates a character matching score matrix between the sentence to be corrected and the target reference sentence based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence. Based on the character matching score matrix, the sentence to be corrected and the target sentence are aligned to obtain a character-level alignment path between the sentence to be corrected and the reference sentence. Using the character-level alignment path, the target character in the sentence to be corrected and the reference character corresponding to the target character are determined. Based on the reference character, the target character in the sentence to be corrected is corrected to obtain the corrected subtitle.

[0116] Because the embodiments of this application can utilize the character matching score matrix to obtain the character-level alignment path between the statement to be corrected and the reference statement, even when the statement to be corrected and the reference statement are not forcibly aligned, the target character in the statement to be corrected and the reference character corresponding to the target character can be determined based on the character-level alignment path, and the error correction processing can be performed on the target character in the statement to be corrected based on the reference character, thereby improving the accuracy of error correction of the statement to be corrected.

[0117] The following is a detailed description of step 201, based on step 201. Please refer to [link / reference]. Figure 3 , Figure 3 Detailed steps for step 201:

[0118] 301. Based on at least one of text similarity, pinyin similarity, and contextual similarity, select target reference paragraphs from the reference text whose similarity to the subtitle to be corrected is greater than a first preset threshold.

[0119] In order to filter out the target reference statement corresponding to the statement to be corrected from the reference text, this application embodiment can adopt the method of first filtering out the target reference paragraph from the reference text to narrow down the filtering range, and then filtering out the target reference statement from the target reference paragraph.

[0120] During the screening process, at least one of the following can be used: text similarity, pinyin similarity, and contextual similarity. Target reference segments with a similarity greater than a first preset threshold to the subtitle to be corrected can be selected from the reference text. When calculating text similarity, methods such as edit distance (Levenshtein), Jaccard similarity, or character overlap rate can be used to capture the degree of text matching between segments. When calculating pinyin similarity, methods such as edit distance and cosine similarity can be used. When calculating contextual similarity, contextual segment matching or sequence similarity methods can be used.

[0121] 302. Based on each sentence to be corrected in the subtitles, a multi-level matching strategy is used to select the target reference sentences corresponding to each sentence to be corrected from the target reference paragraphs. The multi-level matching strategy includes pinyin similarity matching and semantic similarity matching.

[0122] After obtaining the reference paragraphs that match the subtitles to be corrected, a multi-level matching strategy can be used to select the target reference sentences corresponding to each sentence in the subtitles to be corrected. The multi-level matching strategy includes pinyin similarity matching and semantic similarity matching.

[0123] In order to improve the efficiency and accuracy of selecting target reference sentences that match the sentence to be corrected from the reference text, this embodiment of the application adopts a method of first selecting target reference paragraphs from the reference text, and then selecting reference sentences from the target reference paragraphs, thereby improving the efficiency of selecting target reference sentences. Furthermore, in the process of selecting target reference paragraphs and target reference sentences, at least one of textual similarity, pinyin similarity and contextual similarity is used for calculation, thereby improving the accuracy of similarity calculation.

[0124] based on Figure 3 Step 301 in the embodiment is described in detail below. Please refer to [link / reference]. Figure 4 , Figure 4 Detailed steps for step 301:

[0125] 401. Split a reference text and a subtitle to be corrected into a plurality of text segments respectively, to obtain a plurality of split reference text segments and a plurality of split subtitle segments to be corrected;

[0126] For convenience of understanding, the following description is given by way of example:

[0127] Assume that the AI-generated subtitles to be corrected are as shown in Table 1:

[0128]

[0129] The corresponding reference text is: "... Dafeng, the prison of Jingzhao Mansion. Xu Qi'an woke up faintly, and smelled the damp putrid smell in the air, which caused slight discomfort and stomach acid churning. What on earth is this overwhelming stench? Did my family's Erha run to the bed and defecate again? Judging by the pungent degree, I am afraid it was defecated right above my head. Xu Qi'an keeps a Siberian husky at home, commonly known as Erha...".

[0130] It should be noted that there are multiple differences between the reference text and the subtitle to be corrected, which is exactly the manifestation of non-strong consistency: for example, word number difference: "the prison of Jingzhao Mansion" in the reference text has one more word "the" than the subtitle to be corrected; word modification difference: the reference text uses "smelled" while the subtitle to be corrected uses "scented" (synonym replacement; homophone error: the reference text has "Dafeng" while the subtitle to be corrected has "Big Phoenix", etc.

[0131] In the present application, the subtitle to be corrected and the reference text are first split into text segments of fixed size respectively, for example, the subtitle to be corrected and the reference text are respectively split into a plurality of text segments according to a fixed length of 6 characters, and then step 402 is performed for the subtitle segments in the subtitle to be corrected and the reference text segments in the reference text.

[0132] 402. Generate monotonically changing first position sequence identifiers for the plurality of reference text segments, and generate monotonically changing second position sequence identifiers for the plurality of subtitle segments to be corrected;

[0133] After splitting the reference text into a plurality of reference text segments, generate monotonically changing first position sequence identifiers for the plurality of reference text segments, such as i0, i1, i2....in; after splitting the subtitle to be corrected into a plurality of text segments to be corrected, generate monotonically changing second position sequence identifiers for the plurality of text segments to be corrected, such as j0, j1, j2....jn.

[0134] 403. Based on at least one of character similarity, pinyin similarity and context similarity, screen out target reference text segments and target subtitle segments to be corrected with similarity greater than a second preset threshold from the plurality of reference text segments and the plurality of subtitle segments to be corrected;

[0135] After obtaining multiple reference text fragments and multiple subtitle fragments to be corrected, target reference text fragments and target subtitle fragments to be corrected can be selected from the multiple reference text fragments and multiple subtitle fragments to be corrected based on at least one of text similarity, pinyin similarity and contextual similarity.

[0136] 404. Obtain the first position number i1 corresponding to the start position of the target subtitle segment to be corrected in the subtitle segment to be corrected and the second position number i2 corresponding to the end position of the target subtitle segment to be corrected in the subtitle segment to be corrected;

[0137] Assume that the starting position of the target segment to be corrected is the first position number i1=4 and the second position number i2=7 in the subtitle to be corrected.

[0138] 405. Obtain the third position number j1 corresponding to the start position of the target reference text fragment in the reference text and the fourth position number j2 corresponding to the end position of the target reference text fragment in the reference text;

[0139] Correspondingly, assume that the starting position of the target reference text fragment is the third position number j1=5 and the fourth position number j2=8 in the reference text.

[0140] 406. Determine the reference text fragments at j1-i1 as the starting position of the target subtitle fragment to be corrected in the reference text, or determine the reference text fragments at j2-i2 as the starting position of the target subtitle fragment to be corrected in the reference text.

[0141] After obtaining i1=4, i2=7, j1=5, j2=8, this application determines the reference text at j1 as the starting position of the subtitle to be corrected in the reference text. Then, based on the offset of the ending position of the subtitle to be corrected relative to the starting position of the subtitle to be corrected, the ending position of the subtitle to be corrected in the reference text is calculated.

[0142] If the offset of the end position of the subtitle to be corrected relative to the start position of the subtitle to be corrected is determined to be 7 text segments, then starting from position j1 ​​in the reference text, offset 7 text segments backward to obtain the end position of the subtitle to be corrected in the reference text. That is, j8 is determined to be the end position of the subtitle to be corrected in the reference text.

[0143] In this embodiment, the process of identifying a target reference segment similar to the subtitle to be corrected from the reference text is described in detail. Furthermore, this embodiment calculates the offset of the subtitle to be corrected relative to the reference text by segmenting the text and calculating similarity, thereby improving the convenience of identifying the target reference segment from the reference text.

[0144] Figure 4 In this embodiment, after selecting reference paragraphs from the reference text whose similarity to the subtitle to be corrected is greater than a first preset threshold, the target reference statement corresponding to each statement to be corrected in the subtitle can be selected from the reference paragraphs. This process is described in detail below; please refer to [link to relevant documentation]. Figure 5 , Figure 5 Detailed steps for step 302:

[0145] 501. Convert the sentence to be corrected and the target reference paragraph into pinyin sequences respectively to obtain the pinyin sequence of the sentence to be corrected and the pinyin sequence of the target reference paragraph;

[0146] The text preprocessing is performed on the error-correction statement and the target reference paragraph to remove punctuation, spaces, and special symbols; the Chinese character pinyin conversion tool is called to convert each character into a continuous pinyin sequence without tone marks or separators.

[0147] 502. If the pinyin sequence of the target reference paragraph contains the pinyin sequence of the statement to be corrected, then identify the reference statement in the target reference paragraph that contains the pinyin sequence of the statement to be corrected, and determine the reference statement as the target reference statement corresponding to the statement to be corrected.

[0148] After obtaining the pinyin sequences of the statement to be corrected and the reference paragraph, a substring matching algorithm is used to locate the start and end character positions of the pinyin sequence of the statement to be corrected in the pinyin sequence of the target reference paragraph, and to map the start and end character positions back to the reference paragraph in order to extract the target reference statement corresponding to the statement to be corrected from the reference paragraph. If the pinyin sequence of the reference paragraph does not contain the pinyin sequence of the statement to be corrected, then step 503 is executed.

[0149] 503. If the pinyin sequence of the target reference paragraph does not contain the pinyin sequence of the sentence to be corrected, then convert the sentence to be corrected into a first vector, and convert each reference sentence in the target reference paragraph into a second vector, resulting in multiple second vectors.

[0150] For sentences that fail to match pinyin, a bag-of-words model is used for feature extraction: first, the sentence to be corrected and the individual reference sentences split within the reference paragraph are segmented into words to construct a local feature dictionary; then, the sentence to be corrected is mapped to a first feature vector, and each reference sentence is mapped to an independent second feature vector to complete text vectorization.

[0151] 504. Based on multiple similarity values ​​between the first vector and multiple second vectors, determine the target reference sentence corresponding to the sentence to be corrected from the target reference paragraph.

[0152] The cosine similarity between the first vector and each of the second vectors is calculated one by one to obtain a set of similarity values. The reference sentence with the highest similarity value is selected as the target reference sentence corresponding to the current sentence to be corrected, so as to solve the matching problem that the subtitle to be corrected and the reference text are semantically consistent but use different words.

[0153] In this embodiment, the method of pinyin similarity matching is used to solve the matching problem of homophones and similar pronunciations between the reference text and the sentence to be corrected. For the sentence to be corrected that does not match in pinyin, the method of semantic similarity matching is used to determine the target reference sentence corresponding to the sentence to be corrected from the reference paragraph, so as to solve the matching problem of the sentence to be corrected and the reference sentence having the same semantics but different wording.

[0154] Further based on Figure 5 In the embodiments described above, after determining the reference statements corresponding to each statement to be corrected, in order to prevent erroneous exact matching, the embodiments of this application can also use a sliding window mechanism to verify the contextual consistency of the matching results. Specifically, after determining multiple statements to be corrected, based on multiple consecutive statements to be corrected, it can be determined whether the position identifiers of multiple target reference statements corresponding to multiple statements to be corrected in the target reference paragraph are monotonically increasing. If the position identifiers of multiple target reference statements in the target reference paragraph are monotonically increasing, it indicates that the matching is correct. If the position identifiers of multiple target reference statements in the target reference paragraph are not monotonically increasing, such as the position of the target reference statement in the previous sentence being greater than that in the next sentence, it indicates that the matching is abnormal. In this case, the abnormal reference statements whose position identifiers are not monotonically increasing are obtained and deleted to ensure the logical coherence of the overall matching sequence.

[0155] based on Figure 2 Before step 202 in the embodiment, the following steps may also be performed, please refer to [link to relevant documentation]. Figure 6 :

[0156] 601. Obtain at least two of the word embedding, position embedding, and pinyin embedding of each subtitle character in the sentence to be corrected, and fuse at least two of the word embedding, position embedding, and pinyin embedding of each subtitle character to obtain the first fused feature of each subtitle character;

[0157] To improve the accuracy of comparing each character in the sentence to be corrected with each reference character in the target reference sentence, this application can extract multimodal features of each subtitle character in the sentence to be corrected, such as extracting at least two of the word embedding, position embedding and pinyin embedding of each subtitle character, and fusing at least two of the word embedding, position embedding and pinyin embedding of each subtitle character to obtain the first fused feature of each subtitle character.

[0158] Specifically, before obtaining the word embedding, position embedding and pinyin embedding of each subtitle character in the to-be-corrected statement, it is necessary to split the to-be-corrected statement to obtain a subtitle character sequence. For example, when the to-be-corrected statement is "Xu Qi An You You Xing Lai", it needs to be split into 7 characters: "Xu", "Qi", "An", "You", "You", "Xing", "Lai".

[0159] Then, for the above 7 characters, the word embedding, position embedding and pinyin embedding of each subtitle character are extracted respectively. Specifically, when extracting the word embedding feature of each subtitle character, a pre-trained character embedding model (such as CharBERT, Word2Vec character-level model) can be used to map each subtitle character to a word embedding vector of a fixed dimension (such as 128 dimensions), so as to capture the semantic information of the character (such as the semantic difference between "其 (Qi)" and "七 (Qi)"); when extracting the position embedding feature of each subtitle character, based on the position of each character in the to-be-corrected statement "Xu Qi An You You Xing Lai", a sine-cosine position encoding formula (or learnable position encoding) is used to generate a position embedding vector, so as to distinguish the order information of characters (for example, "Xu" is at the beginning of the sentence and "Lai" is at the end of the sentence, so their position embeddings are different), and the dimension is consistent with that of word embedding; when extracting the pinyin embedding of each subtitle character, it is necessary to first obtain the pinyin (including tones) of each subtitle character, split the pinyin into three parts: initial consonant, final vowel and tone, map them to vectors respectively and then splice them, or use a pre-trained pinyin model (such as PinyinBERT) to generate the pinyin embedding vector, whose dimension is consistent with that of word embedding, so as to capture the pronunciation characteristics of characters (solve the problems of similar shape and similar pronunciation errors. For example, "其" and "七" have similar pronunciations, with pinyin "qí" and "qī" respectively; "悠" and "幽" have the same pronunciation, both "yōu", which can be further distinguished by combining with word embedding).

[0160] After obtaining the multimodal features of each subtitle character, weighted summation or concatenation can be used to fuse at least two selected embeddings into the first fused feature.

[0161] 602. Obtain at least two of the word embedding, position embedding and pinyin embedding of each reference character in the target reference statement, fuse at least two of the word embedding, position embedding and pinyin embedding of each reference character, and obtain the second fused feature of each reference character;

[0162] Specifically, if the reference statement is "Xu Qi An You You Xing Zhuan Guo Lai", the process of obtaining the second fused feature of each reference character in the reference statement is similar to the process of obtaining the first fused feature of each subtitle character in step 601, and will not be repeated here.

[0163] 603. Concatenate multiple first fusion features of the statement to be corrected and multiple second fusion features of the reference statement into a feature sequence matrix, and mark the multiple first fusion features and multiple second fusion features in the feature sequence matrix respectively;

[0164] After obtaining multiple first fusion features of the statement to be corrected and multiple second fusion features of the reference statement, the multiple first fusion features and multiple second fusion features are concatenated into a feature sequence matrix. For example, the 7 first fusion features in step 601 and the 9 second fusion features in step 302 are concatenated into a feature sequence matrix to obtain a (7+9)*128 feature sequence matrix, where 128 is the dimension of the first fusion feature and the second fusion feature.

[0165] To distinguish between the first and second fusion features in the feature sequence matrix, a feature type identifier can be added. For example, a identifier bit can be appended to the end of each fusion feature. For instance, the identifier bit of the subtitle character can be set to 0, and the identifier bit of the reference character can be set to 1. In this case, the dimension of the feature sequence matrix becomes 16×129. The identifier bit does not participate in the subsequent attention calculation and is only used to distinguish the feature type.

[0166] After obtaining the feature sequence matrix, the first 7 rows of the feature sequence matrix contain the first fusion feature of the sentence to be corrected, “Xu Qi’an woke up slowly”, and the last 9 rows contain the second fusion feature of the reference sentence, “Xu Qi’an woke up slowly”.

[0167] 604. Encode the first and second fusion features in the feature sequence matrix using a multi-head attention mechanism to obtain a deep feature vector matrix that represents the context information of the first and second fusion features.

[0168] To capture the contextual relationships and cross-character interaction information between subtitle characters and reference characters, shallow fusion features can be upgraded to deep features to improve the discriminative power and representational ability of the features, providing more accurate feature support for subsequent character alignment and error correction.

[0169] After obtaining the 16×129-dimensional feature sequence matrix in step 603, the feature sequence matrix is ​​used as Q (query vector), K (key vector), and V (value vector) of the multi-head attention mechanism, and then multi-head attention calculation is performed. The last identifier bit does not participate in the calculation and is only used for identification.

[0170] When performing multi-head attention calculation, Q, K and V can be respectively split into sub-Q, sub-K and sub-V of multiple heads through multiple parallel linear layers (the number of heads can be set to 4, 8, or 16, which can be adjusted according to actual requirements). For example, when the number of heads is set to 8, 128-dimensional Q, K and V are split into 8 16-dimensional sub-Q, sub-K and sub-V; then attention weights are calculated separately for each head. For example, the first fusion feature (Q) of the Chinese character "其" in the to-be-corrected sentence will be used to calculate similarity with the second fusion features (K) of all characters in the reference sentence, wherein the similarity with the Chinese character "七" is the highest (similar semantics and pronunciation), thus the attention weight of "其" corresponding to "七" is the largest. After obtaining the input features of each head, the output features of all heads are spliced, and a unified-dimensional attention output feature is obtained through a linear transformation layer, that is, a deep feature vector matrix used for characterizing the context information of the first fusion feature and the context information of the second fusion feature is obtained, wherein each feature vector in the deep feature vector matrix contains the context information of subtitle characters and reference characters, and cross-sentence interaction information (such as the association between "其" and "七", the association between "悠" and "幽", and the new features of "转" and "过"), which provides accurate support for subsequent character-level alignment (such as matching "其"→"七", matching "悠"→"幽", and supplementing "转" and "过").

[0171] When adopting multimodal features for each character, the embodiments of the present application can compensate for the defect that a single embedding cannot兼顾 semantics, order and pronunciation, can accurately distinguish similar glyph characters and homophones / near-homophones (such as "其" and "七" having similar pronunciation but different semantics, and "悠" and "幽" having the same pronunciation but different semantics), and simultaneously capture character order information, providing more comprehensive feature support for subsequently distinguishing "悠悠" and "幽幽", matching "其" and "七", and identifying new characters "转" and "过". In addition, the deep feature vector matrix calculated by the multi-head attention mechanism can capture cross-character associations between the to-be-corrected sentence and the reference sentence (such as the semantic and pronunciation association between the subtitle character "其" and the reference character "七", and the semantic association between "悠" and "幽"), and simultaneously mine the context information within a single sentence (such as the order of "悠悠" and the coherent semantics of "幽幽醒转过来"), upgrades shallow fusion features to deep features, significantly improves the discrimination and pertinence of features, and is particularly adapted to scenarios where the length of the to-be-corrected sentence and the reference sentence are different.

[0172] Based on Figure 6 the embodiments described above, the process of matching subtitle characters in the to-be-corrected sentence with each reference character in the reference sentence to obtain a matching score matrix between the to-be-corrected sentence and the reference sentence in step 202 will be described in detail below, please refer to Figure 7 , Figure 7 are the refined steps of step 202:

[0173] 701. Based on the first fusion feature identifier and the second fusion feature identifier, the deep feature vector matrix is ​​split into a first deep feature vector matrix for representing the context information of each first fusion feature and a second deep feature vector matrix for representing the context information of each second fusion feature;

[0174] exist Figure 6 In the embodiment, after concatenating multiple first fusion features of the statement to be corrected and multiple second fusion features of the reference statement into a feature sequence matrix in step 603, the multiple first fusion features and multiple second fusion features in the feature sequence matrix are identified respectively. Therefore, the first fusion features and the second fusion features in the feature sequence matrix are encoded using a multi-head attention mechanism to obtain a deep feature vector matrix used to represent the context information of the first fusion features and the context information of the second fusion features. Therefore, after obtaining the deep feature vector matrix, this embodiment of the application also needs to split the deep feature vector into a first deep feature vector matrix and a second deep feature vector matrix according to the pre-identification of multiple first fusion features and multiple second fusion features.

[0175] As another optional embodiment, during the process of encoding the first fused feature and the second fused feature in the feature sequence matrix using the multi-head attention mechanism, the order of the first fused feature and the second fused feature in the feature sequence is not changed. That is, the order of the first fused feature in the feature sequence matrix is ​​the same as the order of the first deep feature vector matrix in the deep feature vector matrix. Therefore, in this embodiment, when splitting the deep feature vector matrix into the first deep feature vector matrix and the second deep feature vector matrix, the first deep feature vector matrix and the second deep feature vector matrix in the deep feature vector matrix can also be split according to the order of the first fused feature and the second fused feature in the feature sequence.

[0176] 702. Perform operations on the first deep feature vector matrix and the second deep feature vector matrix to obtain the character matching score matrix. The element in the i-th row and j-th column of the character matching score matrix is ​​used to characterize the similarity between the i-th subtitle character in the sentence to be corrected and the j-th reference character in the reference sentence.

[0177] After obtaining the first deep feature vector matrix and the second deep feature vector matrix, the first deep feature vector matrix and the second deep feature vector matrix can be operated on to obtain the character matching score matrix. The element in the i-th row and j-th column of the character matching score matrix is ​​used to characterize the similarity between the i-th subtitle character in the sentence to be corrected and the j-th reference character in the reference sentence.

[0178] Specifically, when performing operations on the first deep eigenvector matrix and the second deep eigenvector matrix, in order to ensure the consistency of the feature distribution, L2 normalization can be performed on the elements in the first deep eigenvector matrix and the second deep eigenvector matrix first, so as to improve the stability of subsequent calculations.

[0179] After obtaining the normalized first deep feature vector matrix and the normalized second deep feature vector matrix, the character-level attention score matrix can be generated by calculating the inner product.

[0180] To make it easier to understand, the following example is provided:

[0181] Assume the first deep feature vector matrix is The second deep feature vector matrix is ;

[0182] Then attention score matrix ;

[0183] Where dim=-1 indicates that softmax is applied to each row of the matrix;

[0184] Then W[i][j] represents the matching score between the i-th character in the statement to be corrected and the j-th character in the reference statement. The higher the score, the more similar the two characters are in terms of both meaning and pronunciation.

[0185] In this embodiment, before performing operations on the first deep feature vector matrix and the second deep feature vector matrix, L2 normalization is performed on the elements in the first deep feature vector matrix and the second deep feature vector matrix respectively, so that the inner product of each element in the first deep feature vector matrix and the second deep feature vector matrix is ​​equal to the cosine similarity, and the probability distribution after Softmax can truly reflect the matching weight between characters and will not be disturbed by the magnitude of the feature vector.

[0186] based on Figure 2 In the aforementioned embodiment, after obtaining the matching score matrix between the statement to be corrected and the reference statement, the following describes the process of performing sequence alignment between the statement to be corrected and the target reference statement based on this matching score matrix to obtain the character-level alignment path between the statement to be corrected and the target reference statement. Please refer to [link to relevant documentation]. Figure 8 , Figure 8 The following are the detailed steps for step 203:

[0187] 801. Construct the initial dynamic programming matrix and the preset empty space penalty;

[0188] After obtaining the matching score matrix between the statement to be corrected and the reference statement, this embodiment of the application generates the character-level alignment path between the statement to be corrected and the reference statement based on the dynamic programming matrix method. Specifically, the dynamic programming matrix needs to first construct an initialized dynamic programming matrix and a preset space penalty.

[0189] To make it easier to understand, the following example is provided:

[0190] Assuming the number of characters to be corrected, N=4, and the number of reference characters, M=4, then the corresponding matching score matrix is ​​W.

[0191] in:

[0192] The penalty gap for open shots is -0.2.

[0193] The resulting initial dynamic programming matrix is ​​(4+1)×(4+1)=5×5.

[0194] 802. Based on the vacancy penalty, fill the boundary vacancy penalty sequence into the initialized dynamic programming matrix;

[0195] After obtaining the vacancy penalty, the boundary vacancy penalty sequence is filled into the initialized dynamic programming matrix DP, where:

[0196] ; ;

[0197] This yields the initial dynamic programming matrix filled with the penalty sequence for boundary vacancies:

[0198]

[0199] 803. Based on the first formula for calculating the score of the matching path, the second formula for calculating the score of the deletion path, the third formula for calculating the score of the insertion path, the matching score matrix, and the empty space penalty, calculate the maximum score of the target element in the initialized dynamic programming matrix on the matching path, deletion path, and insertion path, and determine the maximum score as the element value of the target element. The corresponding dynamic programming matrix is ​​obtained. The element value of the i-th row and j-th column in the dynamic programming matrix is ​​used to represent the maximum score of the first i characters in the statement to be corrected and the first j characters in the reference statement on the matching path, deletion path, and insertion path.

[0200] Among them, the first formula S1 is used to calculate the matching path score, the second formula S2 is used to calculate the deletion path score, which is the missing character in the statement to be corrected; the third formula S3 is used to calculate the insertion path score, which is the missing character in the reference statement.

[0201]

[0202]

[0203]

[0204]

[0205] Following the formula above, each element in the initialized dynamic programming matrix is ​​filled in sequentially, resulting in the following dynamic programming matrix DP:

[0206]

[0207] 804. Based on the backtracking path formula and the character matching score matrix, traverse the corresponding elements sequentially from the bottom right element of the dynamic programming matrix along the preset direction until the boundary element of the dynamic programming matrix is ​​reached, thereby obtaining the character-level alignment path of each subtitle character in the sentence to be corrected and each reference character in the reference sentence.

[0208] After obtaining the dynamic programming matrix, the corresponding elements are traversed sequentially from the bottom right element of the dynamic programming matrix along a preset direction according to the backtracking path formula and the character matching score matrix, until the boundary element of the dynamic programming matrix is ​​reached, thereby obtaining the character-level alignment path of each subtitle character in the sentence to be corrected and each reference character in the reference sentence.

[0209] Specifically, the backtracking path formula in this application embodiment includes: a character matching judgment formula, a reference text character missing judgment formula, and a subtitle character missing judgment formula. The character matching judgment formula is used to indicate that the subtitle character in the statement to be corrected is completely aligned with the corresponding reference character in the target reference statement. The reference text character missing judgment formula is used to indicate that the target reference statement is missing a corresponding subtitle character. The subtitle character missing judgment formula is used to indicate that the statement to be corrected is missing a corresponding reference character.

[0210] For ease of description, let's assume the character matching judgment formula is C1, the missing reference text character judgment formula is C2, and the missing subtitle character judgment formula is C3. Then, the corresponding C1, C2, and C3 are as follows:

[0211] C1: ;

[0212] C2: ;

[0213] C3: ;

[0214] As mentioned above, if the bottom right element of the dynamic programming matrix satisfies the character matching judgment formula C1, it means that the character to be corrected and the reference character are completely matched. Then, we continue to backtrack from the bottom right element of the dynamic programming matrix along the diagonal direction and perform alignment operation on the character to be corrected and the reference character in the alignment result.

[0215] If the bottom right element of the dynamic programming matrix satisfies the missing reference text character judgment formula C2, it means that the corresponding character to be corrected is missing in the reference statement. Then, we continue to backtrack from the bottom right element of the dynamic programming matrix, add the character to be corrected to the alignment result, and add an empty character mark at the reference character position corresponding to the character to be corrected.

[0216] If the bottom right element of the dynamic programming matrix satisfies the formula C3 for judging missing subtitle characters, it means that the corresponding reference character is missing in the sentence to be corrected. Then, we continue to backtrack from the bottom right element of the dynamic programming matrix to the left, add the reference character to the alignment result, and add an empty character mark at the position of the character to be corrected corresponding to the reference character.

[0217] As described above, by tracing back the target elements in the dynamic programming matrix sequentially until the boundary elements in the dynamic programming matrix are reached, the sequence alignment between the error-correcting statement and the target reference statement is completed, and the character-level alignment path between the subtitle characters and the reference characters is obtained.

[0218] To facilitate understanding, the following description uses the dynamic programming matrix DP and the matching score matrix W as examples to illustrate the process of generating character-level matching paths:

[0219] For example, regarding the DP(4,4) element:

[0220] (DP[3][3]+W[3][3]=1.2+0.8=2.0) Tracing back to the top left → character to be corrected 4 Matching reference 4, jump to (3,3)

[0221] For the DP(3,3) element:

[0222] (DP[2][2]+W[2][2]=0.5+0.9=1.2) Tracing back to the top left → matching reference 3 for the character to be corrected, jump to (2,2)

[0223] For the DP (2,2) element:

[0224] (DP[2][1]-0.2=0.7-0.2=0.5) Trace back to the left → reference character 2 Insert an empty character, jump to (2,1)

[0225] For the DP (2,1) element:

[0226] (DP[1][0]+W[1][0]=-0.2+0.9=0.7) Tracing back to the top left → the character to be corrected 2 matches the reference character 1, jump to (1,0)

[0227] For the DP (1,0) element, the character to be corrected 1 has no corresponding reference character. Tracing back to the top →, jump to (0,0), and the backtracking ends.

[0228] Furthermore, after obtaining the character-level alignment path between the character to be corrected and the reference character, this embodiment of the application can also record the matching score between each pair of characters in the alignment path. For example, when the character to be corrected and the reference character are completely matched, the score in the matching score matrix is ​​recorded accordingly. For example, when the i-th character in the character to be corrected and the j-th character in the reference character are completely matched, the value of the W[i-1][j-1] element in the matching score matrix is ​​the matching score between the i-th character in the character to be corrected and the j-th character in the reference character.

[0229] This application uses a dynamic programming algorithm to trace the alignment path between the character to be corrected and the reference character. This dynamic programming algorithm is naturally adapted to scenarios such as normal matching, redundant character deletion, and missing character insertion, making it suitable for actual subtitle correction scenarios such as adding or omitting characters. After accurately locating the positions of missing and extra characters, the algorithm provides a clear positional basis for subsequent correction operations of the text to be corrected.

[0230] based on Figure 8 The following describes in detail the process of correcting the target character in the statement to be corrected based on the reference character, as described in the embodiment above. Please refer to [link to relevant documentation]. Figure 9 :

[0231] 901. Based on the matching score of each character pair in the alignment path, obtain the character-level matching score between the statement to be corrected and the target reference statement; and / or;

[0232] To adapt to different error correction scenarios, this application can use different confidence levels to correct the target character in the character to be corrected. For example, for high confidence scenarios, the character to be corrected is directly replaced with the reference character; for medium confidence scenarios, the character to be corrected is retained and a second verification is performed; for low confidence scenarios, the character to be corrected is not modified to avoid false corrections.

[0233] To improve the accuracy of confidence calculation, embodiments of this application can calculate confidence from multiple dimensions. The multi-dimensional confidence scoring mechanism includes at least two of the following: character-level matching score, sentence-level matching score, empty space ratio, neighborhood reliability score, and harmful domain detection score. When the confidence dimension includes character-level matching score, the character-level matching score between the statement to be corrected and the reference statement is obtained based on the matching score of each character pair in the alignment path.

[0234] Specifically, when calculating the character-level matching score between the statement to be corrected and the reference statement, the alignment results between each pair of characters can be traversed. If the character to be corrected matches the reference character, W[i-1][j-1] is taken as the score for that pair. If the character to be corrected is deleted or inserted, it is recorded as 0 points. Thus, the character-level matching score = the average score of all matching pairs / the sum.

[0235] 902. Calculate the semantic similarity between the sentence to be corrected and the reference sentence to obtain the sentence-level matching score between the sentence to be corrected and the target reference sentence; and / or;

[0236] When calculating the semantic similarity between the sentence to be corrected and the target reference sentence, a text semantic model (such as BERT, Sentence-BERT) can be used to encode the sentence to be corrected and the target reference sentence into vectors, and then the similarity between the two vectors can be calculated to obtain the sentence-level matching score between the sentence to be corrected and the target reference sentence.

[0237] As another optional embodiment, in order to avoid errors caused by polyphonic characters in the statement to be corrected and the target reference statement, this application embodiment may also perform the following operations:

[0238] 1. Identify the polyphonic characters in the sentence to be corrected, and their first pronunciation in the sentence;

[0239] To avoid errors caused by polyphonic characters, embodiments of this application may also introduce polyphonic character matching optimization, such as determining the polyphonic characters in the sentence to be corrected, and the first pronunciation of the polyphonic characters in the sentence to be corrected.

[0240] 2. Identify the polyphonic characters in the sentence to be corrected, their corresponding polyphonic characters in the target reference sentence, and the second pronunciation of the polyphonic characters in the target reference sentence;

[0241] Correspondingly, in this embodiment of the application, after obtaining the first pronunciation of the polyphonic character in the sentence to be corrected, the corresponding polyphonic character in the target reference sentence and the second pronunciation of the polyphonic character in the target reference sentence are further obtained.

[0242] 3. Calculate the similarity between the first and second pronunciations;

[0243] After obtaining the first and second pronunciations, the similarity between the first and second pronunciations is further calculated, and the following steps are performed based on the similarity value.

[0244] 4. If the similarity is less than the preset threshold, the sentence-level matching score will be reduced.

[0245] If the similarity between a polyphonic character in the sentence to be corrected and a polyphonic character in the corresponding reference sentence is lower than a preset threshold, it indicates that the two characters may be mismatched due to the polyphonic character. In this case, the sentence-level matching score will be reduced accordingly. The specific value for reducing the sentence-level matching score can be set according to actual needs, but no specific setting is made here.

[0246] 903. Calculate the percentage of empty characters in the alignment path, and obtain the empty ratio based on the percentage of empty characters; and / or;

[0247] When calculating the proportion of empty spaces, you can count the empty characters in the alignment path, then calculate the total number of aligned characters, and then use the number of empty characters / the total number of aligned characters to get the proportion of empty spaces.

[0248] 904. Calculate the character pair matching scores of each surrounding character at a preset distance from the target character, and determine the neighborhood reliability score of the target character based on the character pair matching scores of each surrounding character; and / or;

[0249] Specifically, when calculating the neighborhood reliability score of a target character, one can take the matching scores of the K characters before and after the target character (e.g., K=2), and then calculate the neighborhood reliability score of the target character according to the following formula, where:

[0250] Neighborhood reliability score = the average score of surrounding character matching.

[0251] 905. Identify the punctuation marks in the target reference statement and determine whether there is a punctuation break at the corresponding position in the statement to be corrected based on the punctuation marks.

[0252] Specifically, after identifying the punctuation marks in the reference sentence, the system can identify whether there are corresponding punctuation marks at the corresponding positions in the sentence to be corrected based on the alignment path. If there are corresponding punctuation marks, it means that there are punctuation breaks in the sentence to be corrected; otherwise, it means that there are no punctuation breaks in the sentence to be corrected.

[0253] 906. If there is no punctuation at the corresponding position of the statement to be corrected, then it is determined that there is a harmful neighborhood in the statement to be corrected; or, if there is punctuation in the statement to be corrected, but the character pair matching score corresponding to at least one character at a preset distance from the punctuation is lower than a preset threshold, then it is determined that there is a harmful neighborhood in the statement to be corrected.

[0254] Specifically, when determining whether a harmful neighborhood exists in a statement to be corrected, it can be determined that the statement to be corrected has a harmful neighborhood if any of the following conditions are met:

[0255] 1. The corresponding reference statement has punctuation, but the statement to be corrected has no punctuation, indicating that the statement to be corrected has sentence concatenation and contains harmful neighboring clauses.

[0256] 2. The sentence to be corrected has a punctuation mark, but the character matching score near the punctuation mark is less than the preset threshold, indicating that the part of the sentence to be corrected is unreliable and there is a harmful neighborhood.

[0257] 907. Using a multidimensional confidence scoring mechanism, a hierarchical decision-making strategy is adopted to correct the target characters in the sentence to be corrected. The multidimensional confidence scoring mechanism includes at least two of the following: character-level matching score, sentence-level matching score, empty space ratio, neighborhood reliability score, and harmful domain detection score. The hierarchical decision-making strategy includes a first confidence correction strategy, a second confidence correction strategy, and a third confidence correction strategy. The error correction threshold of the first confidence correction strategy is less than the error correction threshold of the second confidence correction strategy, and the error correction threshold of the second confidence correction strategy is less than the error correction threshold of the third confidence correction strategy.

[0258] After obtaining the character-level matching score, sentence-level matching score, empty space ratio, neighborhood reliability score, and harmful domain detection score, at least two of these can be used as confidence scoring dimensions, and a hierarchical decision-making strategy can be adopted to correct erroneous characters in the statement to be corrected.

[0259] The hierarchical decision-making strategy includes a first-confidence error correction strategy, a second-confidence error correction strategy, and a third-confidence error correction strategy. The error correction threshold of the first-confidence error correction strategy is lower than that of the second-confidence error correction strategy, and the error correction threshold of the second-confidence error correction strategy is lower than that of the third-confidence error correction strategy. That is, for high-confidence scenarios, the first-confidence threshold is used, and the character to be corrected is directly replaced with the reference character. For medium-confidence scenarios, the second-confidence threshold is used, the character to be corrected is retained, and a second verification is performed. For low-confidence scenarios, the third-confidence threshold is used, and the character to be corrected is not modified to avoid false corrections.

[0260] In this embodiment, multiple confidence dimensions are used to score the confidence level for erroneous characters in the alignment path, thereby improving the reliability of confidence level calculation. On the other hand, for high confidence level scenarios, a lower first confidence level threshold is used to directly correct errors, while for low confidence level scenarios, a higher third confidence level threshold is used to avoid miscorrection of characters to be corrected, thus satisfying different error correction scenarios and improving the reliability of the error correction process.

[0261] based on Figure 9 In the embodiments described above, this application can also determine the reliability of the corrected statement after performing error correction operations on the target characters using alignment paths. Please refer to [link to relevant documentation]. Figure 10 Another embodiment of the subtitle correction method in this application includes:

[0262] 1001. Identify the first and last characters in the statement to be corrected;

[0263] After correcting the misplaced characters in the statement to be corrected, the first character after the correction is performed can be used as the first character and the last character as the last character.

[0264] 1002. If the first and last characters of the statement to be corrected are not aligned with the corresponding target reference statement, calculate the first perplexity of the statement to be corrected before correction and the second perplexity of the statement to be corrected after correction.

[0265] Furthermore, obtain the first and last characters of the reference statement corresponding to the statement to be corrected, and determine whether the first and last characters in the statement to be corrected and the first and last characters in the reference statement are aligned, that is, whether they are the same. If they are the same, it is determined that the two are aligned; if they are not the same, it is determined that the two are not aligned.

[0266] 1003. If the first level of perplexity is less than the second level of perplexity, then abandon the correction of the target character in the correction statement;

[0267] If the first and last characters of the statement to be corrected are not aligned with the corresponding reference statement, then calculate the first perplexity of the statement before correction and the second perplexity of the statement after correction.

[0268] Specifically, when calculating perplexity, the sentence to be corrected before and after correction can be input into a language model (such as GPT2, LSTM, or BERT base) to obtain the sentence fluency score output by the language model, which is also the first and second perplexity output by the language model.

[0269] 1004. If the first perplexity is greater than the second perplexity, then accept the correction of the target character in the statement to be corrected.

[0270] The perplexity level is used to characterize the non-fluency of a language. That is, the higher the perplexity level, the less fluent the language is. Correspondingly, when the first perplexity level is greater than the second perplexity level, it means that the statement to be corrected becomes more fluent after the correction is performed, and the error characters in the statement to be corrected are corrected.

[0271] In this embodiment of the application, after the error correction is performed on the statement to be corrected, the fluency of the corrected statement is further verified by the perplexity degree, thereby improving the reliability of the error correction operation performed on the statement to be corrected.

[0272] based on Figure 9 In the embodiments described above, after performing error correction on the target character using the alignment path, in order to maintain time consistency, the embodiments of this application also need to map the first timestamp of the corrected statement to the second timestamp of the statement to be corrected before the error correction, so that the timestamp of the corrected statement remains consistent with the timestamp of the statement to be corrected before the error correction, thereby avoiding the problem of the corrected statement being out of sync with the original audio.

[0273] It is understood that, in various embodiments of the present invention, the order of the steps does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0274] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0275] The computer device in the embodiments of the present invention will now be described from the perspective of hardware processing:

[0276] One embodiment of the computer device in this invention includes:

[0277] Processor and memory;

[0278] The memory is used to store computer programs, and when the processor executes the computer programs stored in the memory, it can implement the various steps in the above method embodiments.

[0279] It is understood that when the processor in the computer device described above executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be repeated here. For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the gateway device / browser. For example, the computer program can be divided into units in the aforementioned gateway device, and each unit can implement the specific functions described in the corresponding gateway device above.

[0280] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the processor and memory are merely examples of a computer device and do not constitute a limitation on the computer device. It may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0281] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device, connecting various parts of the computer device via various interfaces and lines.

[0282] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0283] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor can be used to perform the various steps in the above method embodiments.

[0284] It is understood that if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0285] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0286] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0287] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0288] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for correcting subtitle errors, characterized in that, include: Select the target reference sentences from the reference text that correspond to the sentences to be corrected in the subtitles; Based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence, a character matching score matrix is ​​generated between the sentence to be corrected and the target reference sentence; the character matching score matrix is ​​used to characterize the degree of matching between each subtitle character and each reference character. Based on the character matching score matrix, the statement to be corrected and the target reference statement are subjected to sequence alignment processing to obtain the character-level alignment path between the statement to be corrected and the target reference statement; Based on the character-level alignment path, determine the target character in the statement to be corrected and the reference character corresponding to the target character; Based on the reference characters, the target characters in the sentence to be corrected are corrected to obtain the corrected subtitles.

2. The method according to claim 1, characterized in that, Before generating the character matching score matrix, the method further includes: Obtain at least two of the word embedding, position embedding, and pinyin embedding of each subtitle character in the sentence to be corrected, and fuse at least two of the word embedding, position embedding, and pinyin embedding of each subtitle character to obtain the first fusion feature of each subtitle character; Obtain at least two of the word embedding, position embedding, and pinyin embedding of each reference character in the target reference statement, and fuse at least two of the word embedding, position embedding, and pinyin embedding of each reference character to obtain a second fused feature for each reference character; Multiple first fusion features of the statement to be corrected and multiple second fusion features of the reference statement are concatenated into a feature sequence matrix, and the multiple first fusion features and multiple second fusion features are respectively identified in the feature sequence matrix; The first fusion feature and the second fusion feature in the feature sequence matrix are encoded using a multi-head attention mechanism to obtain a deep feature vector matrix that represents the context information of the first fusion feature and the context information of the second fusion feature.

3. The method according to claim 2, characterized in that, The step of generating a character matching score matrix between the sentence to be corrected and the target reference sentence based on the matching relationship between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence includes: Based on the first fusion feature identifier and the second fusion feature identifier, the deep feature vector matrix is ​​split into a first deep feature vector matrix for representing the context information of each first fusion feature and a second deep feature vector matrix for representing the context information of each second fusion feature. The first deep feature vector matrix and the second deep feature vector matrix are processed to obtain the character matching score matrix. The element in the i-th row and j-th column of the character matching score matrix is ​​used to characterize the similarity between the i-th subtitle character in the sentence to be corrected and the j-th reference character in the target reference sentence.

4. The method according to claim 1, characterized in that, The step of performing sequence alignment processing on the statement to be corrected and the target reference statement based on the character matching score matrix to obtain the character-level alignment path between the statement to be corrected and the target reference statement includes: Construct an initial dynamic programming matrix and a preset empty space penalty; Based on the vacancy penalty, fill the boundary vacancy penalty sequence into the initialized dynamic programming matrix; Based on the first formula for calculating the matching path score, the second formula for calculating the deletion path score, the third formula for calculating the insertion path score, the character matching score matrix, and the empty space penalty, the maximum score corresponding to the target element in the initialized dynamic programming matrix on the matching path, deletion path, and insertion path is calculated respectively, and the maximum score is determined as the element value of the target element, thus obtaining the dynamic programming matrix. The element value in the i-th row and j-th column of the dynamic programming matrix is ​​used to represent the maximum score of the first i subtitle characters in the sentence to be corrected and the first j reference characters in the target reference sentence on the matching path, deletion path, and insertion path. Based on the backtracking path formula and the character matching score matrix, the corresponding elements are traversed sequentially from the lower right element of the dynamic programming matrix along a preset direction until the boundary element of the dynamic programming matrix is ​​reached, thereby obtaining the character-level alignment path between each subtitle character in the sentence to be corrected and each reference character in the target reference sentence.

5. The method according to claim 4, characterized in that, The backtracking path formula includes a character matching judgment formula, which is used to characterize that the subtitle characters in the sentence to be corrected are completely aligned with the corresponding reference characters in the target reference sentence. The step of traversing the corresponding elements sequentially from the bottom right corner of the dynamic programming matrix along a preset direction based on the backtracking path formula and the character matching score matrix includes: Based on the character matching judgment formula and the character matching score matrix, if it is determined that the subtitle character in the sentence to be corrected is completely aligned with the corresponding reference character in the target reference sentence, then backtracking continues from the lower right element of the dynamic programming matrix along the diagonal direction, and the subtitle character in the sentence to be corrected is aligned with the corresponding reference character in the target reference sentence in the alignment result.

6. The method according to claim 4, characterized in that, The backtracking path formula includes a reference text character missing judgment formula, which is used to characterize the missing corresponding subtitle character in the target reference statement; The step of traversing the corresponding elements sequentially from the bottom right corner of the dynamic programming matrix along a preset direction based on the backtracking path formula and the character matching score matrix includes: Based on the missing reference text character judgment formula and the character matching score matrix, if it is determined that the target reference statement is missing a corresponding subtitle character, then backtracking continues from the lower right corner element of the dynamic programming matrix upwards, and the subtitle character is added to the alignment result, and an empty character marker is added at the reference character position corresponding to the subtitle character.

7. The method according to claim 4, characterized in that, The backtracking path formula includes a subtitle character missing judgment formula, which is used to characterize the missing corresponding reference character in the sentence to be corrected. The step of traversing the corresponding elements sequentially from the bottom right corner of the dynamic programming matrix along a preset direction based on the backtracking path formula and the character matching score matrix includes: Based on the missing subtitle character judgment formula and the character matching score matrix, if it is determined that the corresponding reference character is missing in the sentence to be corrected, then backtracking continues from the lower right element of the dynamic programming matrix to the left, and the reference character is added to the alignment result, and an empty character mark is added at the position of the subtitle character corresponding to the reference character.

8. The method according to any one of claims 5 to 7, characterized in that, The method further includes: Based on the alignment results, after generating the character-level alignment path for each subtitle character in the statement to be corrected and each reference character in the target reference statement, the matching score of each character pair in the alignment path is recorded.

9. The method according to claim 1, characterized in that, The step of selecting target reference sentences from the reference text that correspond to the sentences to be corrected in the subtitles includes: Based on at least one of text similarity, pinyin similarity, and contextual similarity, target reference paragraphs with a similarity greater than a first preset threshold to the subtitle to be corrected are selected from the reference text. Based on each sentence to be corrected in the subtitles to be corrected, a multi-level matching strategy is used to select target reference sentences corresponding to each sentence to be corrected from the target reference paragraphs. The multi-level matching strategy includes pinyin similarity matching and semantic similarity matching.

10. The method according to claim 9, characterized in that, The step of selecting target reference paragraphs from the reference text whose similarity to the subtitle to be corrected is greater than a first preset threshold, based on at least one of text similarity, pinyin similarity, and contextual similarity, includes: The reference text and the subtitle to be corrected are respectively divided into multiple text segments to obtain multiple segmented reference text segments and multiple segmented subtitle segments to be corrected; Generate a monotonically changing first position sequence identifier for the plurality of reference text segments, and generate a monotonically changing second position sequence identifier for the plurality of subtitle segments to be corrected; Based on at least one of text similarity, pinyin similarity, and contextual similarity, target reference text segments and target subtitle segments to be corrected are selected from the plurality of reference text segments and the plurality of subtitle segments to be corrected, with a similarity greater than the second preset threshold. Obtain the first position number i1 corresponding to the start position of the target subtitle segment to be corrected in the subtitle to be corrected, and the second position number i2 corresponding to the end position of the target subtitle segment to be corrected in the subtitle to be corrected; Obtain the third position number j1 corresponding to the start position of the target reference text fragment in the reference text and the fourth position number j2 corresponding to the end position of the target reference text fragment in the reference text; The reference text fragments at positions j1-i1 are determined as the starting positions of the target subtitle fragment to be corrected in the reference text, or the reference text fragments at positions j2-i2 are determined as the starting positions of the target subtitle fragment to be corrected in the reference text.

11. The method according to claim 9, characterized in that, After obtaining multiple target reference statements corresponding to multiple statements to be corrected, the method further includes: Based on multiple consecutive statements to be corrected, determine whether the position identifiers of multiple target reference statements corresponding to the multiple consecutive statements to be corrected in the target reference paragraph are monotonically increasing. If not, then obtain the abnormal reference statements whose position identifiers are not monotonically increasing among the multiple target reference statements, and delete the abnormal reference statements.

12. The method according to claim 9, characterized in that, The step of using a multi-level matching strategy to filter out the target reference statements corresponding to each statement to be corrected from the target reference paragraph includes: The sentence to be corrected and the target reference paragraph are converted into pinyin sequences respectively, to obtain the pinyin sequence of the sentence to be corrected and the pinyin sequence of the target reference paragraph; If the pinyin sequence of the target reference paragraph contains the pinyin sequence of the statement to be corrected, then a reference statement containing the pinyin sequence of the statement to be corrected is determined in the target reference paragraph, and the reference statement is determined as the target reference statement corresponding to the statement to be corrected.

13. The method according to claim 12, characterized in that, The method further includes: If the pinyin sequence of the target reference paragraph does not contain the pinyin sequence of the sentence to be corrected, then the sentence to be corrected is converted into a first vector, and each reference sentence in the target reference paragraph is converted into a second vector, resulting in multiple second vectors; Based on multiple similarity values ​​between the first vector and multiple second vectors, a target reference sentence corresponding to the sentence to be corrected is determined from the target reference paragraph.

14. The method according to claim 1, characterized in that, The step of correcting the target character in the statement to be corrected based on the reference character includes: A multidimensional confidence scoring mechanism is used to correct the target characters in the statement to be corrected using a hierarchical decision strategy. The multidimensional confidence scoring mechanism includes at least two of the following: character-level matching score, sentence-level matching score, empty space ratio, neighborhood reliability score, and harmful domain detection score. The hierarchical decision strategy includes a first confidence correction strategy, a second confidence correction strategy, and a third confidence correction strategy. The error correction threshold of the first confidence correction strategy is less than the error correction threshold of the second confidence correction strategy, and the error correction threshold of the second confidence correction strategy is less than the error correction threshold of the third confidence correction strategy.

15. The method according to claim 14, characterized in that, Before using a multidimensional confidence scoring mechanism to correct the target character in the statement to be corrected using a hierarchical decision-making strategy, the method further includes: Based on the matching score of each character pair in the alignment path, obtain the character-level matching score between the statement to be corrected and the target reference statement; Calculate the semantic similarity between the statement to be corrected and the target reference statement to obtain a sentence-level matching score between them; and / or, The percentage of empty characters in the alignment path is calculated, and the empty ratio is obtained based on the percentage of empty characters; and / or, Calculate the character pair matching scores of each surrounding character within a preset distance of the target character, and determine the neighborhood reliability score of the target character based on the character pair matching scores of each surrounding character; and / or, Identify the punctuation marks in the target reference statement, and determine whether there is a sentence break at the corresponding position of the statement to be corrected based on the punctuation marks. If there is no punctuation break at the corresponding position of the statement to be corrected, then it is determined that there is a harmful neighborhood in the statement to be corrected; or, if there is a punctuation break in the statement to be corrected, but the character pair matching score corresponding to at least one character at a preset distance from the punctuation break is lower than a preset threshold, then it is determined that there is a harmful neighborhood in the statement to be corrected.

16. The method according to claim 14, characterized in that, After using a multidimensional confidence scoring mechanism to correct the target characters in the statement to be corrected using a hierarchical decision-making strategy, the method further includes: Identify the first and last characters in the statement to be corrected; If the first and last characters of the statement to be corrected are not aligned with the corresponding target reference statement, then calculate the first perplexity of the statement to be corrected before correction and the second perplexity of the statement to be corrected after correction. If the first perplexity is less than the second perplexity, then the correction of the target character in the statement to be corrected is abandoned; If the first perplexity is greater than the second perplexity, then error correction is received for the target character in the statement to be corrected.

17. The method according to claim 14, characterized in that, After determining the sentence-level matching score between the statement to be corrected and the target reference statement, the method further includes: Identify the polyphonic characters in the sentence to be corrected, and the first pronunciation of the polyphonic characters in the sentence to be corrected; Identify the polyphonic characters in the statement to be corrected, the corresponding polyphonic characters in the target reference statement, and the second pronunciation of the polyphonic characters in the target reference statement; Calculate the similarity between the first pronunciation and the second pronunciation; If the similarity is less than a preset threshold, the sentence-level matching score is reduced.

18. The method according to claim 14, characterized in that, After using a multidimensional confidence scoring mechanism to correct the target characters in the statement to be corrected using a hierarchical decision-making strategy, the method further includes: The first timestamp of the corrected statement is mapped to the second timestamp of the statement to be corrected before the correction, so that the timestamp of the corrected statement is consistent with the timestamp of the statement to be corrected before the correction.

19. A computer device comprising a processor, characterized in that, When the processor executes a computer program stored in the memory, it is used to implement the subtitle correction method as described in any one of claims 1 to 18.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the subtitle correction method as described in any one of claims 1 to 18.

21. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the subtitle correction method as described in any one of claims 1 to 18.