Multi-text difference recognition method, device and electronic device

By identifying the difference between the standard text of the audio and multiple target texts, and using the context information of the different elements to match and annotate the target elements, the problem of difficulty in analyzing the error differences between multiple target texts in the prior art is solved, and an intuitive analysis of the functional differences of the speech to subtitle tool is realized.

CN113962211BActive Publication Date: 2025-05-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111258033.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-05-06
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

The prior art is difficult to intuitively reflect the error differences between multiple target texts relative to standard texts, and it is impossible to effectively analyze the functional differences between speech to subtitles tools.

Method used

By determining standard text for audio and multiple target texts, identify the difference between each target text relative to the standard text, obtain the difference element, and determine the matching target element from the standard text and other text based on the context information of the difference element, and if matched, pre-labeled in other texts.

Benefits of technology

It realizes the intuitive reflection of error differences between multiple target texts relative to standard texts, providing a basis for analyzing the functional differences of the speech to subtitle tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113962211B_ABST
    Figure CN113962211B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for identifying differences between multiple texts, and relates to the field of computer technology, in particular to the field of text recognition technology, and specifically to a method, device, and electronic device for identifying differences between multiple texts. The specific implementation scheme is: determine a standard text for audio, and multiple target texts whose differences are to be identified; for each target text, identify the difference between the target text and the standard text to obtain a difference element; from the target text, determine the context information of the first element included in the difference element; wherein the first element is an element of a change type; based on the context information of the first element, determine the target element that matches the position of the first element from the standard text and other texts respectively; if the determined target elements are the same, perform a first predetermined annotation on the determined target element in the other texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the field of text recognition technology, and specifically to a method, device and electronic device for identifying differences between multiple texts. Background Art

[0002] In scenarios such as subtitle recognition, it is usually necessary to analyze the functions of speech-to-subtitle tools.

[0003] In the related art, a diff algorithm is used to identify the difference between the text to be identified, that is, the subtitle text obtained by processing the audio through a speech-to-subtitle tool, and the standard text of the audio, so as to obtain the difference elements of the text to be identified relative to the standard text, and mark the difference elements with colors, thereby reflecting the difference between the text to be identified and the standard text. Summary of the invention

[0004] The present invention provides a method, device, equipment and storage medium for identifying differences between multiple texts.

[0005] According to one aspect of the present disclosure, a method for identifying differences between multiple texts is provided, comprising:

[0006] Determine a standard text for the audio, and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio;

[0007] For each target text, identifying the difference between the target text and the standard text to obtain a difference element;

[0008] Determining context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type;

[0009] Based on the context information of the first element, determine the target element matching the position of the first element from the standard text and other texts respectively; wherein the other text is a text other than the target text among the multiple target texts;

[0010] If the determined target elements are the same, a first predetermined marking is performed on the determined target elements in the other texts.

[0011] According to another aspect of the present disclosure, a multi-text difference identification device is provided, comprising:

[0012] A first determination module is used to determine a standard text for the audio and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio;

[0013] A first recognition module is used to identify the difference between each target text and the standard text to obtain a difference element;

[0014] A second determination module is used to determine context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type;

[0015] A first matching module is used to determine, based on the context information of the first element, a target element that matches the position of the first element from the standard text and other texts, respectively; wherein the other text is a text other than the target text among multiple target texts;

[0016] The first marking module is used to perform a first predetermined marking on the determined target element in the other text if the determined target element is the same.

[0017] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the above-mentioned multi-text difference identification method.

[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the steps of the above-mentioned multi-text difference identification method.

[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the steps of the above-mentioned multi-text difference identification method are implemented.

[0020] Through this solution, the error differences between multiple texts relative to the standard text can be intuitively reflected, thereby providing an analytical basis for analyzing the functional differences between speech-to-subtitle tools.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0023] Figure 1 is a flow chart of a multi-text difference identification method according to the present disclosure;

[0024] Figure 2 is another flow chart of a multi-text difference identification method according to the present disclosure;

[0025] Figure 3 is another flow chart of a multi-text difference identification method according to the present disclosure;

[0026] Figure 4 is an application flow chart of the multi-text difference identification method disclosed in the present invention;

[0027] Figure 5 is a flowchart of a specific example of the multi-text difference identification method according to the present disclosure;

[0028] Figure 6 is a structural schematic diagram of a multi-text difference recognition device according to the present disclosure;

[0029] Figure 7 The block diagram is a block diagram of an electronic device for implementing the multi-text difference recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] In scenarios such as subtitle recognition, it is usually necessary to analyze the functions of speech-to-subtitle tools. For example, text A and text B for the same audio are identified by speech-to-subtitle tool A and speech-to-subtitle tool B, and then the error differences between text A and text B relative to the standard text are compared, that is, which content is correctly identified in text A (i.e., the same as the standard text content) and incorrectly identified in text B (i.e., different from the standard text content), or which content is correctly identified in text B and incorrectly identified in text A. By comparing the error differences between text A and text B relative to the standard text, the functional gap between the two speech-to-subtitle tools can be analyzed. Among them, the standard text is usually obtained by manually listening to the audio content.

[0032] In the related art, a diff (different) algorithm is used to compare the differences between two texts. Since the diff algorithm can only compare the differences between two texts, when there are more than one speech-to-subtitle tool that recognizes multiple target texts for the same audio, the diff algorithm can only mark the differences between any target text and the standard text, or the differences between any target text and other target texts, but cannot mark the elements in any target text that are the same as the standard text but different from the other target texts. Therefore, the error differences between multiple target texts relative to the standard text cannot be intuitively reflected. Subsequently, the functional analysis between the various speech-to-subtitle tools cannot be achieved.

[0033] Based on the above content, in order to intuitively reflect the error differences between multiple target texts and a standard text, the embodiments of the present disclosure provide a multi-text difference identification method, apparatus, device and storage medium.

[0034] The following first introduces a multi-text difference recognition method provided by an embodiment of the present disclosure.

[0035] The multi-text difference recognition method provided by the embodiment of the present disclosure can be applied to an electronic device. In a specific application, the electronic device can be a server or a terminal device, which is reasonable. In practical applications, the terminal device can be: a smart phone, a tablet computer, a desktop computer, etc.

[0036] Specifically, the execution subject of the multi-text difference recognition method can be a multi-text difference recognition device. Exemplarily, when the multi-text difference recognition method is applied to a terminal device, the multi-text difference recognition device can be a functional software running in the terminal device, for example: an APP (Application) or a web client with a function of identifying differences between multiple texts; of course, the multi-text difference recognition device can also be a plug-in in the existing functional software. Exemplarily, when the multi-text difference recognition method is applied to a server, the multi-text difference recognition device can be a computer program running in the server, and the computer program can be used to realize the recognition of differences between multiple texts.

[0037] The method for identifying differences between multiple texts provided by the embodiment of the present disclosure may include the following steps:

[0038] Determine a standard text for the audio, and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio;

[0039] For each target text, identifying the difference between the target text and the standard text to obtain a difference element;

[0040] Determining context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type;

[0041] Based on the context information of the first element, determine the target element matching the position of the first element from the standard text and other texts respectively; wherein the other text is a text other than the target text among the multiple target texts;

[0042] If the determined target elements are the same, a first predetermined marking is performed on the determined target elements in the other texts.

[0043] In the scheme provided by the present disclosure, by identifying the difference between each target text and the standard text, a difference element is obtained; from each target text, the context information of the first element of the change type included in the difference element is determined; based on the context information of the first element, the target element matching the position of the first element is determined from the standard text and other texts respectively; if the determined target elements are the same, it means that the target element that matches the position of the first element in the standard text is equivalent to the target element that matches the position of the first element in other texts. At this time, the first predetermined annotation is performed on the determined target element in other texts. In this way, for each target text, the elements that are correctly identified in other texts but incorrectly identified in the target text can be intuitively displayed. It can be seen that through this scheme, the error differences between multiple target texts relative to the standard text can be intuitively reflected, thereby providing an analytical basis for the functional differences between speech-to-subtitle tools.

[0044] The following describes the multi-text difference recognition method provided by the embodiment of the present disclosure in conjunction with the accompanying drawings.

[0045] like Figure 1 As shown, a multi-text difference recognition method provided by an embodiment of the present disclosure may include the following steps:

[0046] S101, determining a standard text for audio and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio;

[0047] In this embodiment, the audio may be any file that is in an audio format and contains voice content, and the voice content included in the audio is not limited in this embodiment. In addition, there may be multiple ways to obtain the audio. For example, the audio may be a file selected from a file stored locally in an electronic device, or a file downloaded from a network terminal; of course, the audio may also be extracted from a target video by a predetermined audio extraction device, wherein the target video may be a video pre-downloaded by a video player software. For example, the predetermined audio extraction device may be a ffmpeg tool (Fast Forward Mpeg, Fast Forward Mpeg), which is a set of open source computer programs that can be used to record and convert digital audio and video, and can convert them into streams. The audio stream can be extracted from the video by the ffmpeg tool.

[0048] Among them, the multiple target texts whose differences are to be identified can be subtitle texts obtained by identifying the audio through multiple different speech-to-subtitle tools; and the standard text can be the manual verification result of any text among the multiple target texts, that is, the standard text can be obtained by manually listening to the audio and then proofreading any target text, thereby greatly improving the generation efficiency of the standard text.

[0049] S102, for each target text, identifying the difference between the target text and the standard text, and obtaining the difference element;

[0050] For each target text, when comparing it with the standard text, the text content can be segmented by characters, and then a text comparison algorithm, such as a diff algorithm, is used to identify the difference between each target text and the standard text, and obtain difference elements, that is, difference characters. The difference elements may include one or more elements, and the type of elements included in the difference elements may be either text character elements or punctuation elements.

[0051] Furthermore, after the difference element is identified, the index position and difference type of the element included in the difference element in the target text can be recorded, and the difference type can be addition, deletion, change, etc. Among them, when the difference type is addition, it means that the difference element is an element added in the target text relative to the standard text; when the difference type is deletion, it means that the difference element is an element deleted in the target text relative to the standard text; when the difference type is change, it means that the difference element is an element changed in the target text relative to the standard text.

[0052] S103, determining context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type;

[0053] After identifying the difference elements of the target text relative to the standard text through step S102, for the first element, that is, the element whose difference type in the target text is the change type, the context information of the first element is determined. In one implementation, the index position of the first element in the target text can be used to find the elements at multiple positions before and after the index position as the context information of the first element. Exemplarily, if the index position of the first element in the target text is 8, the elements at 3 positions before and after the index position can be found as context information, then the above information of the first element is the elements at the 5th to 7th positions in the target text, and the following information is the elements at the 9th to 11th positions in the target text. For example: if the target text is "The market is not ideal recently and people have been asking if there are any recommended funds", the first element is "one", the above information is "too ideal", and the following information is "straight people".

[0054] Optionally, in an implementation manner, before determining the context information of the first element included in the difference element from the target text, the method further includes:

[0055] The elements included in the difference element are merged; wherein the merging includes merging elements with consecutive positions.

[0056] After obtaining the difference elements in step S102, the difference elements with consecutive positions in the target text are merged, and the index positions of the merged continuous difference elements are an interval range. At this time, if the difference elements include first elements with consecutive index positions, the first elements with consecutive index positions are merged. When searching for context information later, it is not necessary to search for context for each first element, but only to search for the elements at the first and last positions of the index interval where the merged first element is located as the context information of the merged first element, thereby improving the search efficiency.

[0057] S104, based on the context information of the first element, determining a target element that matches the position of the first element from the standard text and other texts respectively; wherein the other text is a text other than the target text among multiple target texts;

[0058] After determining the context information of the first element through step S103, the context information is found from the standard text and other texts, so that the target element matching the position of the first element in the standard text and other texts is found through the index position information of the context information in the standard text and other texts. In addition, in order to better understand the so-called other texts, an example is taken in which the multiple target texts are 2 texts, and the 2 texts include text A and text B. For text A, text B is other text; similarly, an example is taken in which the multiple target texts are 3 texts: the 3 texts include text A, text B and text C. For text A, text B can be other text, or text C can also be other text.

[0059] It is understandable that due to the presence of added and deleted elements in the target text, the elements at the same index position in the target text and the standard text will be different, making it impossible to determine whether the aforementioned elements at the same index position are misidentified elements. Therefore, in this embodiment, for the first element of the changed type, context information is searched to determine the target element that matches the position of the first element in the standard text and other texts, that is, the element at the position between the index interval where the previous information is located and the index interval where the following information is located, so that it can be determined whether the target element that matches the position of the first element in other texts is correctly identified.

[0060] Optionally, in an implementation, the determining, based on the context information of the first element, target elements matching the position of the first element from the standard text and other texts respectively includes steps A1-A2:

[0061] A1, detecting, from the standard text, a first position of an element that matches the preceding information of the first element, and a second position of an element that matches the following information of the first element; and determining an element between the first position and the second position as a target element that matches the position of the first element in the standard text;

[0062] It is understandable that after determining the context information of the first element in the target text, the context information can be found in the standard text. The first position of the element in the standard text that matches the previous information of the first element is the index interval where the previous information in the standard text is located; the second position of the element in the standard text that matches the following information of the first element is the index interval where the following information in the standard text is located. Therefore, the element at the position between the first position and the second position is the target element in the standard text that matches the position of the first element.

[0063] For example, if the index interval of the previous information of the first element in the standard text is [5-7], and the index interval of the following information in the standard text is [9-11], then the target element that matches the position of the first element in the standard text is the element with index position 8.

[0064] A2, detecting, from the other text, a third position of an element that matches the preceding information of the first element, and a fourth position of an element that matches the following information of the first element; determining an element between the third position and the fourth position as a target element in the other text that matches the position of the first element.

[0065] It is understandable that after determining the context information of the first element in the target text, the context information can be found from other target texts (i.e., other texts) other than the target text. The third position of the element in other texts that matches the previous information of the first element is the index interval where the previous information in the other text is located; the fourth position of the element in other texts that matches the following information of the first element is the index interval where the following information in the other text is located. Therefore, the element between the third position and the fourth position is the target element in other texts that matches the position of the first element.

[0066] S105: If the determined target elements are the same, a first predetermined annotation is performed on the determined target elements in the other texts.

[0067] It can be understood that if the determined target elements are the same, then the target element corresponding to the first element position in the other text is the same as the target element at the position in the standard text, indicating that the content of the other text at the position is correctly recognized, while the content of the target text at the position is incorrectly recognized. At this time, in the other text, the determined target element is first predeterminedly marked, so that the speech-to-subtitle tool used by the target text can be compared with the speech-to-subtitle tool used by the other text by analyzing the elements that are correctly recognized in the other text and incorrectly recognized in the target text.

[0068] Exemplarily, the first predetermined marking may be an intuitive marking method such as color marking, bold font marking, highlight marking, etc., to facilitate user viewing and analysis.

[0069] It can be seen that through this scheme, the error differences between multiple target texts relative to the standard text can be intuitively reflected, thereby providing an analytical basis for the functional differences between speech-to-subtitle tools.

[0070] Optionally, in another embodiment of the present disclosure, Figure 2 As shown, the multi-text difference identification method may include steps S201-S206:

[0071] S201, determining a standard text for the audio and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio;

[0072] S202, for each target text, identifying the difference between the target text and the standard text, and obtaining the difference element;

[0073] S203, determining context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type;

[0074] S204, based on the context information of the first element, determine a target element that matches the position of the first element from the standard text and other texts respectively; wherein the other text is a text other than the target text among multiple target texts;

[0075] S205, if the determined target elements are the same, performing a first predetermined marking on the determined target elements in the other texts;

[0076] S206: In the target text, perform a second predetermined annotation on each element included in the difference element; wherein the second predetermined annotation is different from the first predetermined annotation.

[0077] The contents of steps S201-S205 are the same as those of steps S101-S105, and are not described again here.

[0078] In this embodiment, in the target text, a second predetermined annotation can also be performed for each element included in the difference element. Among them, each element included in the difference element can be an element of the added, deleted or changed type, and the second predetermined annotation can be a color annotation or a bold font annotation, etc. It can be understood that the second predetermined annotation for the difference element can intuitively reflect the situation of the speech-to-subtitle tool recognizing errors, that is, the difference of the target text relative to the standard text. It can be understood that if there are both elements of the second predetermined annotation and elements of the first predetermined annotation in a target text, then the difference between the standard text, that is, the content that is incorrectly recognized relative to the standard text, can be intuitively reflected through the elements of the second predetermined annotation, and the content that is incorrectly recognized in another text other than the target text can be intuitively reflected through the elements of the first predetermined annotation, and the target text recognizes the correct content. Thus, the recognition accuracy of multiple target texts can be compared.

[0079] It can be seen that by performing a second predetermined marking on the elements in the target text that are different from the standard text through this scheme, the recognition errors in the target text can be intuitively displayed. Moreover, by performing a first predetermined marking on the elements that are correctly recognized in other texts but incorrectly recognized in the target text, the error differences in different target text recognitions relative to the standard text can be intuitively reflected.

[0080] Optionally, in another embodiment of the present disclosure, before identifying the difference between the target text and the standard text and obtaining the difference element for each target text, the method further includes:

[0081] Detecting a current marking mode; wherein the marking mode is a character marking mode or a punctuation marking mode;

[0082] In this embodiment, before identifying the difference between the target text and the standard text, the current annotation mode can also be detected. If the annotation mode is a character annotation mode, when identifying the difference, only the difference between the text characters in the target text and the standard text is identified, but the difference between the punctuation marks is not identified; if the annotation mode is a punctuation annotation mode, when identifying the difference, only the difference between the punctuation marks in the target text and the standard text is identified, but the difference between the text characters is not identified.

[0083] Exemplarily, the marking mode can be selected by parameter control, and the parameter control can be implemented by buttons. For example, by setting a "punctuation comparison" button, when the "punctuation comparison" button is clicked, the punctuation marking mode is entered.

[0084] Accordingly, for each target text, identifying the difference between the target text and the standard text to obtain the difference element includes:

[0085] For each target text, if the marking mode is detected as a character marking mode, the difference between the target text and the standard text regarding characters is identified to obtain a difference element; if the marking mode is detected as a punctuation marking mode, the difference between the target text and the standard text regarding punctuation is identified to obtain a difference element.

[0086] It can be understood that by setting the character marking mode and the punctuation marking mode, when it is detected that the current marking mode is the character marking mode, only the differences in text characters are displayed, thereby avoiding the interference caused by the differences in punctuation marks. In this mode, it is more conducive to analyzing the differences in text content recognition; when it is detected that the current marking mode is the punctuation marking mode, only the differences in punctuation marks are displayed, so that the accuracy of sentence recognition can be analyzed more clearly.

[0087] In the related art, the diff algorithm can identify the differences between the target text and the standard text in terms of characters and punctuation, and these two types of differences are uniformly reflected, which looks confusing and is not conducive to separate analysis. The solution of this embodiment uses different annotation modes to separately display the differences in characters and punctuation, which is more conducive to separate analysis.

[0088] Optionally, in another embodiment of the present disclosure, Figure 1 Based on the example, Figure 3 As shown, the multi-text difference identification method may further include steps S301-S304:

[0089] S301, identifying each proper noun in the standard text and the corresponding index position;

[0090] In this embodiment, proper nouns are professional terms for various industries, and each proper noun can be identified from the standard text by a proper noun recognition tool. The proper noun recognition tool can be a recognition tool for each category, for example, a recognition tool for medical proper nouns and a recognition tool for chemical proper nouns. It can be understood that by identifying each proper noun and its corresponding position in the standard text, and finding out from the target text whether there are elements with incorrect recognition in the position matching the proper noun, that is, the second predetermined marked elements, the errors of the proper nouns in the target text can be analyzed, and thus, it can be analyzed in which category of proper nouns the relevant speech-to-subtitle tool has more advantages in analyzing.

[0091] S302, searching for context information for each proper noun using the index position corresponding to each proper noun;

[0092] After the corresponding index position of each proper noun in the standard text is found in step S301, elements at the preceding and succeeding index positions of the index position corresponding to each proper noun are found as the context information of each proper noun.

[0093] S303, based on the context information of each proper noun, determining a target proper noun matching the position of each proper noun from each target text;

[0094] After finding the context information of each proper noun in the standard text in step S302, a search is performed in the target text according to the context information to find an element matching the position of the context information. The element matching the position of the context information found in the target text is the context information of the target proper noun, so that the target proper noun matching the position of each proper noun can be determined.

[0095] S304: If the determined target proper noun includes the second predetermined marked element, it is determined that the target proper noun is a misrecognized target proper noun.

[0096] Since the second predetermined annotation is used to mark the difference elements between the target text and the standard text, if the determined target proper noun includes the elements of the second predetermined annotation, it means that there are misrecognized elements in the target proper noun, so it can be determined that the target proper noun is an misrecognized target proper noun.

[0097] Among them, the method for determining whether the target proper noun includes the second predetermined marked element can be: according to the index interval of the target proper noun in the target text where it is located, each element on the index interval is searched to see whether there is the second predetermined marked element; if so, the target proper noun includes the second predetermined marked element.

[0098] It is worth mentioning that the scheme of the present disclosure can also be used to evaluate the accuracy of the speech-to-subtitle tool used for each target text. Exemplarily, the evaluation method can be: calculating the ratio of the number of the second predetermined marked elements in the target text to the total number of elements as the error rate, and the speech-to-subtitle tool corresponding to the target text with a small error rate has a higher accuracy; or calculating the weighted value of the ratio of the number of the second predetermined marked elements in the target text to the total number of elements and the ratio of the number of proper nouns containing the second predetermined marked elements to the total number of proper nouns as the error rate, and the speech-to-subtitle tool corresponding to the target text with a small error rate has a higher accuracy.

[0099] It can be seen that through this solution, the errors of proper nouns in the target text can be further analyzed, so as to analyze which types of proper nouns the relevant speech-to-subtitle tools have more advantages in analyzing.

[0100] like Figure 4 As shown, evaluating the functions of different speech-to-subtitle tools can include the following steps:

[0101] (1) Obtain video content through the uploaded video link;

[0102] (2) Use ffmpeg tool to extract audio from video content;

[0103] (3) Recognize the audio through different speech-to-subtitle tools to obtain the recognized text corresponding to the different speech-to-subtitle tools (corresponding to the target text mentioned above);

[0104] (4) Performing simple manual proofreading on any recognized text, and using the result of manual proofreading as the standard text;

[0105] (5) storing the standard text and the plurality of recognition texts;

[0106] (6) using the above-mentioned multi-text difference recognition method to compare each recognized text with the standard text, marking the different elements in the recognized text relative to the standard text in red (corresponding to the second predetermined marking mentioned above), and marking the elements in the recognized text that are the same as the standard text but different from the other recognized texts in blue (corresponding to the first predetermined marking mentioned above);

[0107] (7) scoring different speech-to-subtitle tools; wherein the scoring method may be to calculate the ratio of the number of difference elements in the identified text to the total number of elements in the text, wherein the smaller the ratio, the higher the score;

[0108] (8) Generate evaluation results of different speech-to-subtitle tools.

[0109] In order to better understand the present solution, the present solution is introduced below in conjunction with a flowchart of a specific example.

[0110] like Figure 5 As shown, (1) obtain the standard text, the recognized text A and the recognized text B. For example:

[0111] Standard text: "The market has not been ideal recently and people have been asking if there are any recommended funds or what to buy, so I sorted out the things you should pay attention to before buying funds."

[0112] Recognize text A: "The market has not been ideal recently. People have been asking if there are any recommended funds or what to buy, so I sorted out the things you should pay attention to before buying funds."

[0113] Identify text B: "The market has not been ideal recently and people have been asking if there are any recommendations for chicken gold or what to buy, so I sorted out the things you should pay attention to before buying funds."

[0114] (2) The contents of the standard text, the recognized text A and the recognized text B are cut according to characters and stored in list O, list A and list B respectively.

[0115] For example, list O: ['most', 'nearest', 'city', 'market', 'not'...]

[0116] (3) Use the diff algorithm to calculate the differences between List O and List A and List B respectively, and record the difference elements in List A and List B from List O, the type of the difference elements (addition, deletion, change) and the index of the difference elements, that is, record the elements included in the difference elements and the difference type of each element.

[0117] For example, if the recognized text A has an additional character "呀" compared to the standard text, when the program traverses list A, it will record the index position and type of "呀". The specific record is as follows: {"location": 8, "type": "add", "add_chars": "呀"};

[0118] If the recognized text A has deleted the character "适" compared to the standard text, when the program traverses list A, it will record the index position and type of the element before "适". The specific record is as follows: {"location": 27, "type": "del", "del_chars": "适"};

[0119] If the recognized text A has changed the character "主" compared to the standard text, when the program traverses list A, it will record the index position and type of "主". The specific record is as follows: {"location": 44, "type": "change", "change_chars": "主"};

[0120] If the recognized text B has changed the character "鸡" compared to the recognized text A, when the program traverses list B, it will record the index position and type of "鸡". The specific record is as follows: {"location": 19, "type": "change", "change_chars": "鸡"}.

[0121] (4) Preprocess list A and list B to obtain list A-O and list B-O. Among them, the preprocessing includes: merging consecutive elements included in the different elements, and adding context information to the different elements with the type of "change".

[0122] For example, for {"location": 19, "type": "change", "change_chars": "鸡"} recorded in list B, after adding context information, list B-O is obtained: {"location": 19, "type": "change", "change_chars": "鸡", "before": "没有推荐的", "after": "金或者买什么"}.

[0123] (5) Annotate the recognized text A and the recognized text B respectively through the context information in list A-O and list B-O to obtain the recognized results with color annotation.

[0124] For example, according to the context information of the "chicken" element in list BO, that is, the above information is "no recommendation" and the following information is "gold or what to buy", the corresponding context information in list A and list O is searched, so that the target element matching the position of the "chicken" element in list A and list O can be found. The target element cut out at the same position in list A is: "base", and the target element cut out at the same position in list O is: "base", so it can be judged that the element "chicken" that is misidentified in the recognition text B is correctly identified as "base" in the recognition text A, so the position of "base" in the recognition text A is marked in blue (corresponding to the first predetermined mark above), and the "chicken" in the recognition text B is marked in red (corresponding to the second predetermined mark above).

[0125] It can be seen that through this solution, the differences in misrecognition between multiple texts to be recognized for the standard text can be intuitively marked, thereby providing an analytical basis for the functional differences between speech-to-subtitle tools.

[0126] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0127] Based on the above method embodiment, the present disclosure also provides a multi-text difference recognition device, such as Figure 6 As shown, the device comprises:

[0128] The first determination module 610 is used to determine a standard text for the audio and a plurality of target texts for which differences are to be identified; wherein each target text is a subtitle text of the audio;

[0129] A first recognition module 620 is used to identify the difference between each target text and the standard text to obtain a difference element;

[0130] A second determination module 630 is used to determine context information of a first element included in the difference element from the target text; wherein the first element is an element of a change type; a first matching module 640 is used to determine, based on the context information of the first element, target elements that match the position of the first element from the standard text and other texts, respectively; wherein the other text is a text other than the target text among multiple target texts;

[0131] The first marking module 650 is configured to perform a first predetermined marking on the determined target element in the other text if the determined target element is the same.

[0132] Optionally, the device further comprises:

[0133] A second marking module is used to perform a second predetermined marking on each element included in the difference element in the target text; wherein the second predetermined marking is different from the first predetermined marking. Optionally, the first matching module includes:

[0134] A first matching unit is used to detect, from the standard text, a first position of an element that matches the preceding information of the first element, and a second position of an element that matches the following information of the first element; and determine an element at a position between the first position and the second position as a target element that matches the position of the first element in the standard text;

[0135] The second matching unit is used to detect, from the other text, a third position of an element that matches the preceding context information of the first element, and a fourth position of an element that matches the following context information of the first element; and determine an element at a position between the third position and the fourth position as a target element in the other text that matches the position of the first element.

[0136] Optionally, the device also includes: a merging module, used to perform merging processing on the elements included in the difference element before the second determination module determines the context information of the first element included in the difference element from the target text; wherein the merging processing includes merging elements with consecutive positions.

[0137] Optionally, the device further comprises: a detection module, configured to detect a current annotation mode before the first recognition module performs, for each target text, identification of a difference between the target text and the standard text to obtain a difference element; wherein the annotation mode is a character annotation mode or a punctuation annotation mode;

[0138] The first identification module is specifically used for:

[0139] For each target text, if the marking mode is detected as a character marking mode, the difference between the target text and the standard text regarding characters is identified to obtain a difference element; if the marking mode is detected as a punctuation marking mode, the difference between the target text and the standard text regarding punctuation is identified to obtain a difference element.

[0140] Optionally, the device further comprises:

[0141] A second recognition module, used to recognize each proper noun in the standard text and the corresponding index position;

[0142] A search module, used to search for context information for each proper noun by using the index position corresponding to each proper noun;

[0143] A second matching module, configured to determine, from each target text, target proper nouns that match the positions of the respective proper nouns based on the context information of the respective proper nouns;

[0144] The determination module is configured to determine that the target proper noun is an incorrectly identified target proper noun if the determined target proper noun includes a second predetermined marked element.

[0145] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0146] An electronic device provided by the present disclosure may include:

[0147] at least one processor; and

[0148] A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the above-mentioned multi-text difference identification method.

[0149] The present disclosure provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned multiple text difference identification methods are implemented.

[0150] In another embodiment provided by the present disclosure, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute the steps of any of the multi-text difference identification methods in the above embodiments.

[0151] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0152] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0153] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0154] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the multi-text difference recognition method. For example, in some embodiments, the multi-text difference recognition method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the multi-text difference recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the multi-text difference recognition method in any other appropriate manner (e.g., by means of firmware).

[0155] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0156] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0157] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0159] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0160] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0161] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0162] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for identifying differences between multiple texts, comprising: Determine a standard text for the audio, and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio; Detecting a current marking mode; wherein the marking mode is a character marking mode or a punctuation marking mode; For each target text, if it is detected that the marking mode is a character marking mode, the difference of the target text with respect to the standard text in terms of characters is identified to obtain a difference element; if it is detected that the marking mode is a punctuation marking mode, the difference of the target text with respect to the standard text in terms of punctuations is identified to obtain a difference element; Determining context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type; Detecting, from the standard text, a first position of an element that matches the preceding context information of the first element, and a second position of an element that matches the following context information of the first element; and determining an element between the first position and the second position as a target element that matches the position of the first element in the standard text; Detecting, from other texts, a third position of an element that matches the preceding information of the first element, and a fourth position of an element that matches the following information of the first element; determining an element between the third position and the fourth position as a target element in the other text that matches the position of the first element; wherein the other text is a text other than the target text among multiple target texts; If the determined target elements are the same, a first predetermined marking is performed on the determined target elements in the other texts.

2. The method according to claim 1, further comprising: In the target text, performing a second predetermined marking on each element included in the difference element; The second predetermined mark is different from the first predetermined mark.

3. The method according to claim 1 or 2, wherein: Before determining the context information of the first element included in the difference element from the target text, the method further includes: The elements included in the difference element are merged; wherein the merging includes merging elements with consecutive positions.

4. The method according to claim 1 or 2, further comprising: Identify each proper noun in the standard text and the corresponding index position; Using the index position corresponding to each proper noun, searching for context information for each proper noun; Based on the context information of each proper noun, respectively determine a target proper noun matching the position of each proper noun from each target text; If the determined target proper noun includes the second predetermined marked element, it is determined that the target proper noun is a misrecognized target proper noun.

5. A multi-text difference recognition device, comprising: A first determination module is used to determine a standard text for the audio and a plurality of target texts whose differences are to be identified; wherein each target text is a subtitle text of the audio; A first recognition module is used to identify the difference between each target text and the standard text to obtain a difference element; A second determination module is used to determine context information of a first element included in the difference element from the target text; wherein the first element is an element of a modified type; A first matching module is used to determine, based on the context information of the first element, a target element that matches the position of the first element from the standard text and other texts, respectively; wherein the other text is a text other than the target text among multiple target texts; A first marking module, configured to perform a first predetermined marking on the determined target element in the other text if the determined target element is the same; Wherein, the first matching module includes: A first matching unit is used to detect, from the standard text, a first position of an element that matches the preceding information of the first element, and a second position of an element that matches the following information of the first element; and determine an element at a position between the first position and the second position as a target element that matches the position of the first element in the standard text; A second matching unit is used to detect, from the other text, a third position of an element that matches the preceding context information of the first element, and a fourth position of an element that matches the following context information of the first element; and determine an element at a position between the third position and the fourth position as a target element in the other text that matches the position of the first element; The device also includes: A detection module, configured to detect a current annotation mode before the first recognition module performs, for each target text, identification of a difference between the target text and the standard text to obtain a difference element; wherein the annotation mode is a character annotation mode or a punctuation annotation mode; The first identification module is specifically used for: For each target text, if the marking mode is detected as a character marking mode, the difference between the target text and the standard text regarding characters is identified to obtain a difference element; if the marking mode is detected as a punctuation marking mode, the difference between the target text and the standard text regarding punctuation is identified to obtain a difference element.

6. The device according to claim 5, further comprising: A second marking module, used for performing a second predetermined marking on each element included in the difference element in the target text; The second predetermined mark is different from the first predetermined mark.

7. The device according to claim 5 or 6, further comprising: A merging module is used to perform merging processing on the elements included in the difference element before the second determination module determines the context information of the first element included in the difference element from the target text; wherein the merging processing includes merging elements with consecutive positions.

8. The device according to claim 5 or 6, further comprising: A second recognition module, used to recognize each proper noun in the standard text and the corresponding index position; A search module, used to search for context information for each proper noun by using the index position corresponding to each proper noun; A second matching module, configured to determine, from each target text, target proper nouns that match the positions of the respective proper nouns based on the context information of the respective proper nouns; The determination module is configured to determine that the target proper noun is an incorrectly identified target proper noun if the determined target proper noun includes a second predetermined marked element.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

11. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speech recognition using multiple recognizers (selectively) applied to the same input sample

    US6122613A

  • Audio corpus screening method and device for use in speech recognition, and computer device

    WO2020224119A1