Text acquisition method and device, storage medium and computer equipment

By calculating the edit distance between the recognition texts of multiple audio recognition systems and determining the target recognition text, the problem of low speech recognition accuracy in the prior art is solved, and more efficient text acquisition is achieved.

CN120236587AActive Publication Date: 2025-07-01BEIJING XIYU JIZHI TECH CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510305443.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-01
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

The existing automatic speech recognition technology has low accuracy and has a long recognition time, making it difficult to meet the needs of high accuracy.

Method used

By acquiring the recognition text of the target audio by multiple audio recognition systems, the editing distance between different recognition texts is calculated, and the target recognition text is determined to improve the accuracy of text acquisition.

Benefits of technology

It improves the accuracy of text acquisition, reduces recognition time, and enhances the efficiency of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236587A_ABST
    Figure CN120236587A_ABST
Patent Text Reader

Abstract

The invention discloses a text acquisition method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining a recognition text set obtained through the recognition of a target audio, the recognition text set comprises a plurality of recognition texts obtained through the recognition of the target audio by each audio recognition system in an audio recognition system set, the audio recognition system set comprises at least three audio recognition systems, the audio recognition systems in the audio recognition system set are in one-to-one correspondence with the recognition texts in the recognition text set, the editing distance between the two recognition texts is obtained in the recognition text set, and based on the multiple editing distances, the audio recognition systems in the audio recognition system set correspond to the recognition texts in the recognition text set. And determining a target audio recognition system in the audio recognition system set, and determining a target recognition text corresponding to the target audio recognition system in the recognition text set as an output text of the target audio. By adopting the method and the device, the target recognition text is determined by calculating the editing distance between different recognition texts, so that the text acquisition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This case is a divisional application of a patent for invention with the application number 202210952021.8, the application date of August 9, 2022, and the invention title of a text acquisition method, device, storage medium, and computer device. This application relates to the field of computer technology, and particularly relates to a text acquisition method, device, storage medium, and computer device. Background Art

[0002] In the prior art, automatic speech recognition technology is often used for speech recognition to convert speech into text content. However, the existing automatic speech recognition technology mainly relies on manually set speech conversion rules, with low accuracy and long recognition time. There is a need to propose a speech recognition method with higher accuracy. Summary of the Invention

[0003] Embodiments of this application provide a text acquisition method, device, storage medium, and electronic device, which can determine a target recognition text by calculating the edit distance between different recognition texts, improving the accuracy of text acquisition. The technical solution is as follows:

[0004] In a first aspect, embodiments of this application provide a text acquisition method, including: obtaining a set of recognition texts obtained by recognizing a target audio, where the set of recognition texts includes multiple recognition texts obtained by each audio recognition system in an audio recognition system set respectively recognizing the target audio, the audio recognition system set includes at least three audio recognition systems, and the audio recognition systems in the audio recognition system set correspond to the recognition texts in the set of recognition texts one by one; obtaining the edit distance between two recognition texts in the set of recognition texts respectively, and based on multiple edit distances, determining a target audio recognition system in the audio recognition system set; and determining the target recognition text corresponding to the target audio recognition system in the set of recognition texts as the output text of the target audio.

[0005] Second aspect, an embodiment of the present application provides a text acquisition device, including: a text set acquisition module, configured to acquire an identification text set obtained by identifying a target audio, where the identification text set includes multiple identification texts obtained by each audio identification system in an audio identification system set respectively identifying the target audio, the audio identification system set includes at least three audio identification systems, and the audio identification systems in the audio identification system set correspond to the identification texts in the identification text set one by one; a target system determination module, configured to obtain the edit distance between two identification texts in the identification text set respectively, and determine a target audio identification system in the audio identification system set based on multiple edit distances; an output text determination module, configured to determine the target identification text corresponding to the target audio identification system in the identification text set as the output text of the target audio.

[0006] Third aspect, an embodiment of the present application provides a computer storage medium, where the computer storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the steps of the above method.

[0007] Fourth aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is suitable for being loaded and executed by the processor to perform the steps of the above method.

[0008] In an embodiment of the present application, an identification text set obtained by identifying a target audio is acquired. The identification text set includes multiple identification texts obtained by each audio identification system in an audio identification system set respectively identifying the target audio. The audio identification system set includes at least three audio identification systems, and the audio identification systems in the audio identification system set correspond to the identification texts in the identification text set one by one. The edit distance between two identification texts is acquired in the identification text set respectively, and a target audio identification system is determined in the audio identification system set based on multiple edit distances. The target identification text corresponding to the target audio identification system is determined as the output text of the target audio in the identification text set. By calculating the edit distance between different identification texts to determine the target identification text, the accuracy of text acquisition is improved. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1It is an example schematic diagram of text acquisition provided by an embodiment of the present application;

[0011] Figure 2 It is a schematic flowchart of a text acquisition method provided by an embodiment of the present application;

[0012] Figure 3 It is a schematic flowchart of a text acquisition method provided by an embodiment of the present application;

[0013] Figure 4 It is an example schematic diagram of example acquisition of editing provided by an embodiment of the present application;

[0014] Figure 5 It is a schematic flowchart of a text acquisition method provided by an embodiment of the present application;

[0015] Figure 6 It is a schematic structural diagram of a text acquisition device provided by an embodiment of the present application;

[0016] Figure 7 It is a schematic structural diagram of a target system determination module provided by an embodiment of the present application;

[0017] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0019] The text acquisition device may be a terminal device such as a mobile phone, a computer, a tablet computer, a smart watch, or a vehicle-mounted device, or may be a module in the terminal device for implementing the text acquisition method. The text acquisition device may acquire a target audio and acquire the text content corresponding to the target audio, that is, the output text of the target audio. The target audio may be an audio containing voice information input by the user for the terminal device, or may be an audio locally saved in the terminal device or obtained from the network. After the text acquisition device acquires the output text of the target audio, it may display the output text on the terminal device.

[0020] Please refer to Figure 1, This is an example schematic diagram for text acquisition provided by an embodiment of the present application. The text recognition device can acquire the target audio input by the user and use the audio recognition system to recognize the target audio. The audio recognition system is the Automatic Speech Recognition (ASR) system, which can convert human speech into text. The text recognition device can use each audio recognition system in the audio recognition system set to recognize the target audio separately to obtain multiple recognition texts, such as Figure 1 Recognition text 1, Recognition text 2,..., Recognition text n in it, where n is an integer greater than or equal to 2, and the audio recognition system set includes at least three audio recognition systems. Then, the text acquisition device can find a target recognition text among the multiple recognition texts based on the edit distance, determine the target recognition text as the output text, and then control the terminal device to display the output text. The edit distance is a quantitative measurement of the difference degree between two strings, and the measurement method is to see how many times of processing are required at least to change one string into another string.

[0021] The following will specifically describe the text acquisition method provided by the present application with reference to specific embodiments.

[0022] Please refer to Figure 2 , This is a schematic flowchart of a text acquisition method provided by an embodiment of the present application. As Figure 2 shown, the method of the embodiment of the present application may include the following steps S101 - S103.

[0023] S101, Acquire the recognition text set obtained by recognizing the target audio.

[0024] Specifically, the text acquisition device can use each audio recognition system in the audio recognition system set to recognize the target audio separately, so as to obtain multiple recognition texts and save them in the recognition text set. It can be understood that the audio recognition systems in the audio recognition system set correspond one-to-one with the recognition texts in the recognition text set, and one audio recognition system corresponds to one recognition text.

[0025] S102, Obtain the edit distance between two recognition texts in the recognition text set respectively, and determine the target audio recognition system in the audio recognition system set based on the multiple edit distances.

[0026] Specifically, the text acquisition device can obtain the edit distances between two recognition texts in the text set respectively to obtain multiple edit distances. For example, if there are three recognition texts in the text set, the text acquisition device can obtain four edit distances. The text acquisition device can determine the target audio recognition system in the audio recognition system set according to these multiple edit distances. For example, if there are a total of n audio recognition systems, then n recognition texts of the target audio can be obtained, and the edit distance between every two recognition texts is calculated. An audio recognition system has n - 1 corresponding edit distances, and the sum of the n - 1 edit distances can be obtained as the total edit distance corresponding to the audio recognition system. The edit distance reflects the difference between recognition texts. The smaller the edit distance, the more similar the two recognition texts are, and the higher the accuracy of the recognition text. The text recognition device can confirm the audio recognition system with the smallest total edit distance as the target audio recognition system. It can be understood that when performing text recognition on the target audio, the target audio recognition system is the audio recognition system with the highest accuracy in the audio recognition system set.

[0027] S103. In the recognition text set, determine the target recognition text corresponding to the target audio recognition system as the output text of the target audio.

[0028] Specifically, the text acquisition device can determine the target recognition text corresponding to the target audio recognition system in the recognition text set as the output text of the target audio, and the text acquisition device can display the output text on the terminal device.

[0029] In the embodiment of the present application, a recognition text set obtained by recognizing a target audio is acquired. The recognition text set includes multiple recognition texts obtained by each audio recognition system in the audio recognition system set respectively recognizing the target audio. The audio recognition system set includes at least three audio recognition systems. The audio recognition systems in the audio recognition system set and the recognition texts in the recognition text set are in one-to-one correspondence. The edit distances between two recognition texts are respectively acquired in the recognition text set. Based on the multiple edit distances, the target audio recognition system is determined in the audio recognition system set. In the recognition text set, the target recognition text corresponding to the target audio recognition system is determined as the output text of the target audio. By calculating the edit distances between different recognition texts to determine the target recognition text, the accuracy of text acquisition is improved.

[0030] Please refer to Figure 3 , which is a schematic flowchart of a text acquisition method provided by an embodiment of the present application. As Figure 3 shown, the method of the embodiment of the present application may include the following steps S201 - S205.

[0031] S201. Acquire a recognition text set obtained by recognizing a target audio.

[0032] Specifically, the text acquisition device can use each audio recognition system in the audio recognition system set to recognize the target audio respectively, so as to obtain multiple recognized texts and save them in the recognized text set. It can be understood that the audio recognition systems in the audio recognition system set and the recognized texts in the recognized text set are in one-to-one correspondence, and one audio recognition system corresponds to one recognized text.

[0033] S202. Obtain the edit distance between two recognized texts in the recognized text set respectively.

[0034] Specifically, the text acquisition device can obtain the edit distance between two recognized texts in the text set respectively, so as to obtain multiple edit distances. For example, if there are three recognized texts in the text set, the text acquisition device can obtain four edit distances.

[0035] Optionally, the text acquisition device can obtain the first recognized text corresponding to the first audio recognition system in the recognized text set. The first audio recognition system can be any audio recognition system in the audio recognition system set, and the first recognized text is the recognized text obtained by the first audio recognition system for recognizing the target audio. The text acquisition device can obtain the first edit distance between the first recognized text and the second recognized text in the recognized text set respectively. The second recognized text is other recognized texts in the recognized text set except the first recognized text. It can be understood that the second recognized text can be multiple, and the first edit distance can also be multiple. Please refer to Figure 4 , which provides an example schematic diagram for obtaining the edit distance in the embodiments of the present application. The text acquisition device can use the audio recognition system A to obtain the recognized text A corresponding to the target audio. If the audio recognition system A is the first audio recognition system, then the recognized text A is the recognized text A corresponding to the first audio recognition system. In addition to the recognized text A in the text set, there are also recognized text B, recognized text C,..., recognized text N, etc. Recognized text B, recognized text C,..., recognized text N are the second recognized texts in the text set. The text acquisition device can calculate the edit distance between the recognized text A and other recognized texts in the text set. For example, the edit distance AB between the recognized text A and the recognized text B, the edit distance AC between the recognized text A and the recognized text C, the edit distance AN between the recognized text A and the recognized text N, etc. The edit distances AB, AC,..., AN are the first edit distances corresponding to the first audio recognition system.

[0036] S203. Based on the multiple edit distances, determine the total edit distance corresponding to each audio recognition system respectively.

[0037] Specifically, the text acquisition device can determine the total edit distance corresponding to each audio recognition system respectively based on the multiple edit distances.

[0038] Optionally, the text acquisition device may add up all the first edit distances to obtain the total edit distance corresponding to the first audio recognition system.

[0039] S204. Determine a target audio recognition system in the audio recognition system set based on the total edit distance corresponding to each audio recognition system.

[0040] Specifically, the edit distance reflects the difference between the recognized texts. The smaller the edit distance, the more similar the two recognized texts are, and the higher the accuracy of the recognized texts. The text acquisition device may determine the target audio recognition system in the audio recognition system set according to the total edit distance corresponding to each audio recognition system. It can be understood that when performing text recognition on the target audio, the target audio recognition system is the audio recognition system with the highest accuracy in the audio recognition system set.

[0041] Optionally, the smaller the total edit distance, the higher the accuracy rate of the audio recognition system. The text acquisition device may determine the minimum sum of the edit distances among the total edit distances corresponding to each audio recognition system, and determine the second audio recognition system corresponding to the minimum sum of the edit distances as the target audio recognition system in the audio recognition system.

[0042] S205. Determine the target recognized text corresponding to the target audio recognition system in the recognized text set as the output text of the target audio.

[0043] Specifically, the text acquisition device may determine the target recognized text corresponding to the target audio recognition system in the recognized text set as the output text of the target audio, and the text acquisition device may display the output text on the terminal device.

[0044] In the embodiment of the present application, a recognized text set obtained by recognizing a target audio is acquired. The recognized text set includes multiple recognized texts obtained by respectively recognizing the target audio by each audio recognition system in the audio recognition system set. The audio recognition system set includes at least three audio recognition systems. The audio recognition systems in the audio recognition system set correspond one-to-one with the recognized texts in the recognized text set. The edit distance between two recognized texts is respectively acquired in the recognized text set, the total edit distance corresponding to each audio recognition system is respectively determined based on the multiple edit distances, the audio recognition system with the smallest total edit distance is determined as the target audio recognition system, and the target recognized text corresponding to the target audio recognition system in the recognized text set is determined as the output text of the target audio. By calculating the edit distance between different recognized texts to determine the target recognized text, the accuracy of text acquisition is improved.

[0045] Please refer to Figure 5, which provides a schematic flow chart of a text acquisition method for an embodiment of the present application. As Figure 5 shown, the method of the embodiment of the present application may include the following steps S301-S306.

[0046] S301, obtain a set of recognized texts obtained by recognizing a target audio.

[0047] Specifically, the text acquisition device may use each audio recognition system in the set of audio recognition systems to recognize the target audio respectively, so as to obtain multiple recognized texts and save them in the set of recognized texts. It can be understood that the audio recognition systems in the set of audio recognition systems correspond one-to-one with the recognized texts in the set of recognized texts, and one audio recognition system corresponds to one recognized text.

[0048] S302, obtain the first recognized text corresponding to the first audio recognition system in the set of recognized texts, and respectively obtain the first edit distance between the first recognized text and the second recognized text in the set of recognized texts.

[0049] Specifically, the text acquisition device may respectively obtain the edit distance between two recognized texts in the text set, so as to obtain multiple edit distances. For example, if there are three recognized texts in the text set, the text acquisition device may obtain four edit distances. The text acquisition device may obtain the first recognized text corresponding to the first audio recognition system in the set of recognized texts. The first audio recognition system may be any audio recognition system in the set of audio recognition systems. The first recognized text is the recognized text obtained by the first audio recognition system for recognizing the target audio. The text acquisition device may respectively obtain the first edit distance between the first recognized text and the second recognized text in the set of recognized texts. The second recognized text is other recognized texts in the set of recognized texts except the first recognized text. It can be understood that the second recognized text may be multiple, and the first edit distance may also be multiple.

[0050] S303, based on a preset threshold and multiple first edit distances, determine the first audio recognition system as a candidate audio recognition system.

[0051] Specifically, the text acquisition device may determine the first audio recognition system as a candidate audio recognition system based on a preset threshold and multiple first edit distances of the first audio recognition system. The preset threshold is used to judge the text recognition accuracy rate of the first audio recognition system and determine whether the first audio recognition system can be determined as a candidate audio recognition system.

[0052] Optionally, the edit distance reflects the difference between the recognized texts. The smaller the edit distance, the more similar the two recognized texts are, and the higher the accuracy of the recognized texts. The preset threshold can be a distance threshold. When multiple first edit distances of the first audio recognition system are all less than or equal to the distance threshold, the first audio recognition system is determined as a candidate audio recognition system, where the distance threshold can be the initial setting of the text acquisition device, or can be set by the user or relevant staff. For example, the text acquisition device can obtain the recognized text A corresponding to the target audio using the audio recognition system A. If the audio recognition system A is the first audio recognition system, then the recognized text A is the recognized text A corresponding to the first audio recognition system. If there are three audio recognition systems in the audio recognition system set, namely the audio recognition system A, the audio recognition system B, and the audio recognition system C, and the corresponding recognized texts are the recognized text A, the recognized text B, and the recognized text C respectively, the text acquisition device can calculate the first edit distance of the first audio recognition system as the edit distance AB between the recognized text A and the recognized text B, and the edit distance AC between the recognized text A and the recognized text C. Therefore, only when both the edit distance AB and the edit distance AC are less than or equal to the distance threshold, the audio recognition system A can be determined as a candidate audio recognition system. On the contrary, if the edit distance AB is greater than the distance threshold, the edit distance AC is greater than the distance threshold, or both the edit distance AB and the edit distance AC are greater than the distance threshold, it means that the recognition accuracy of the audio recognition system A for the target audio is low, and the audio recognition system A cannot be confirmed as a candidate audio recognition system.

[0053] Optionally, since the edit distance usually increases with the increase of the length of the recognized text, if the distance threshold is still used to determine the candidate audio recognition system when the recognized text is long, the accuracy of determining the candidate audio recognition system will be reduced. When the length of the recognized text is greater than the length threshold, the preset threshold can be a ratio threshold, that is, the ratio threshold is used to determine the audio recognition system. The text acquisition device can perform normalization processing on all the edit distances, can obtain the average text length of multiple recognized texts in the recognized text set, calculate the ratio of the first edit distance to the average text length, and obtain multiple length ratios corresponding to the first audio recognition system. When all the length ratios of the first audio recognition system are less than or equal to the ratio threshold, the text acquisition device can determine the first audio recognition system as a candidate audio recognition system. The ratio threshold can be the initial setting of the text acquisition device, or can be set by the user or relevant staff.

[0054] S304, determine the total edit distance corresponding to the candidate audio recognition system based on multiple first edit distances.

[0055] Specifically, the text acquisition device can add the first edit distances corresponding to the candidate audio recognition system to obtain the total edit distance corresponding to the candidate audio recognition system.

[0056] S305. Determine the target audio recognition system in the candidate audio recognition systems based on the total edit distance corresponding to the candidate audio recognition systems.

[0057] Specifically, the text acquisition device can determine the target audio recognition system among all the candidate audio recognition systems according to the total edit distance corresponding to each candidate audio recognition system. It can be understood that when performing text recognition on the target audio, the target audio recognition system is the audio recognition system with the highest accuracy in the set of audio recognition systems.

[0058] Optionally, when there are at least two candidate audio recognition systems, the text acquisition device can determine the minimum total edit distance among the total edit distances corresponding to each candidate audio recognition system, and determine the third audio recognition system corresponding to the minimum total edit distance as the target audio recognition system in the set of audio recognition systems. When there are at least two third audio recognition systems, that is, when there are two or more audio recognition systems with the minimum total edit distance, the text acquisition device can randomly determine one of the third audio recognition systems as the target audio recognition system.

[0059] When there are no candidate audio recognition systems, that is, when the first edit distances of all audio recognition systems are greater than the length threshold, or when the length ratios of all audio recognition systems are greater than the ratio threshold, it means that the accuracy of the recognized texts recognized by all audio recognition systems is not high. The text acquisition device can discard the target audio and not perform recognition on the target audio.

[0060] It can be understood that when there are no candidate audio recognition systems, the text acquisition device can send the target audio to the terminal device of the staff, and the staff can perform manual annotation on the target audio.

[0061] S306. Determine the target recognized text corresponding to the target audio recognition system in the set of recognized texts as the output text of the target audio.

[0062] Specifically, the text acquisition device can determine the target recognized text corresponding to the target audio recognition system in the set of recognized texts as the output text of the target audio, and the text acquisition device can display the output text on the terminal device.

[0063] In an embodiment of the present application, an identification text set obtained by identifying a target audio is acquired. The identification text set includes multiple identification texts obtained by each audio identification system in an audio identification system set respectively identifying the target audio. The audio identification system set includes at least three audio identification systems, and the audio identification systems in the audio identification system set correspond one-to-one with the identification texts in the identification text set. In the identification text set, a first identification text corresponding to a first audio identification system and a first edit distance are acquired. Based on a preset threshold and multiple first edit distances, the first audio identification system is determined as a candidate audio identification system. The preset threshold can be a distance threshold, and when the length of the identification text is too long, the preset threshold can be a ratio threshold. Screening the audio identification systems using the preset threshold further improves the accuracy of text acquisition. Based on multiple first edit distances, the total edit distance corresponding to each candidate audio identification system is determined respectively, and the candidate audio identification system with the smallest total edit distance is determined as the target audio identification system. In the identification text set, the target identification text corresponding to the target audio identification system is determined as the output text of the target audio. By calculating the edit distance between different identification texts to determine the target identification text, the accuracy of text acquisition is improved.

[0064] The following will combine the attached Figure 6 - attached Figure 7 , and introduce the text acquisition device provided in the embodiment of the present application in detail. It should be noted that the text acquisition device in the attached Figure 6 - attached Figure 7 is used to execute the method of the embodiment shown in the present application Figures 2 - 5 . For the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the embodiment shown in the present application Figures 2 - 5 .

[0065] Please refer to Figure 6 , which shows a schematic structural diagram of a text acquisition device provided in an exemplary embodiment of the present application. The text acquisition device can be implemented as all or part of the device through software, hardware, or a combination of both. The device 1 includes a text set acquisition module 11, a target system determination module 12, and an output text determination module 13.

[0066] The text set acquisition module 11 is configured to acquire an identification text set obtained by identifying a target audio. The identification text set includes multiple identification texts obtained by each audio identification system in an audio identification system set respectively identifying the target audio. The audio identification system set includes at least three audio identification systems, and the audio identification systems in the audio identification system set correspond one-to-one with the identification texts in the identification text set;

[0067] A target system determination module 12 is configured to obtain the edit distance between two recognized texts in the recognized text set respectively, and determine a target audio recognition system in the audio recognition system set based on multiple such edit distances.

[0068] Specifically, please refer to Figure 7 together, which provides a schematic structural diagram of a target system determination module for an embodiment of this application. Figure 7 As shown, the target system determination module 12 may include:

[0069] An edit distance acquisition unit 121 is configured to obtain the edit distance between two recognized texts in the recognized text set respectively.

[0070] Optionally, the edit distance acquisition unit 121 is specifically configured to obtain a first recognized text corresponding to a first audio recognition system in the recognized text set, where the first audio recognition system is any audio recognition system in the audio recognition system set.

[0071] Respectively obtain a first edit distance between the first recognized text and a second recognized text in the recognized text set, where the second recognized text is other recognized texts in the recognized text set except the first recognized text.

[0072] A distance sum acquisition unit 122 is configured to respectively determine the total edit distance corresponding to each audio recognition system based on multiple such edit distances.

[0073] Optionally, the distance sum acquisition unit 122 may include:

[0074] A candidate system determination subunit 124 is configured to determine the first audio recognition system as a candidate audio recognition system based on a preset threshold and multiple such first edit distances.

[0075] Optionally, the candidate system determination subunit 124 is specifically configured to obtain the average text length of the multiple recognized texts in the recognized text set.

[0076] Based on the average text length and multiple such first edit distances, obtain multiple length ratios corresponding to the first audio recognition system.

[0077] When multiple such length ratios are all less than or equal to a ratio threshold, determine the first audio recognition system as a candidate audio recognition system.

[0078] Optionally, the candidate system determination subunit 124 is specifically configured to determine the first audio recognition system as a candidate audio recognition system when multiple such first edit distances are all less than or equal to a distance threshold.

[0079] A distance sum obtaining subunit 125, configured to determine a total edit distance corresponding to the candidate audio recognition system based on the plurality of first edit distances;

[0080] A target system determining unit 123, configured to determine a target audio recognition system in the audio recognition system set based on the total edit distance corresponding to each audio recognition system;

[0081] Optionally, the target system determining unit 123 is specifically configured to determine a target audio recognition system in the candidate audio recognition systems based on the total edit distance corresponding to the candidate audio recognition system.

[0082] Optionally, when there are at least two candidate audio recognition systems, the target system determining unit 123 is specifically configured to determine a minimum total edit distance among the total edit distances corresponding to the candidate audio recognition systems, and determine the third audio recognition system corresponding to the minimum total edit distance in the audio recognition system set as the target audio recognition system;

[0083] When there is no such candidate audio recognition system, discard the target audio.

[0084] Optionally, the target system determining unit 123 is specifically configured to determine a minimum total edit distance among the total edit distances corresponding to each audio recognition system, and determine the second audio recognition system corresponding to the minimum total edit distance in the audio recognition system set as the target audio recognition system.

[0085] An output text determining module 13, configured to determine the target recognition text corresponding to the target audio recognition system in the recognition text set as the output text of the target audio.

[0086] In an embodiment of the present application, an identification text set obtained by identifying a target audio is acquired. The identification text set includes multiple identification texts obtained by each audio identification system in an audio identification system set respectively identifying the target audio. The audio identification system set includes at least three audio identification systems, and the audio identification systems in the audio identification system set and the identification texts in the identification text set are in one-to-one correspondence. In the identification text set, a first identification text corresponding to a first audio identification system and a first edit distance are acquired. Based on a preset threshold and multiple first edit distances, the first audio identification system is determined as a candidate audio identification system. The preset threshold can be a distance threshold, and when the length of the identification text is too long, the preset threshold can be a ratio threshold. Screening the audio identification systems by using the preset threshold further improves the accuracy of text acquisition. Based on the multiple first edit distances, the total edit distance corresponding to each candidate audio identification system is respectively determined, and the candidate audio identification system with the smallest total edit distance is determined as the target audio identification system. In the identification text set, the target identification text corresponding to the target audio identification system is determined as the output text of the target audio. Determining the target identification text by calculating the edit distance between different identification texts improves the accuracy of text acquisition.

[0087] It should be noted that when the text acquisition device provided in the above embodiment executes the text acquisition method, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the text acquisition device provided in the above embodiment and the text acquisition method embodiment belong to the same concept, and the implementation process is shown in detail in the method embodiment, which will not be repeated here.

[0088] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0089] An embodiment of the present application further provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded and executed by a processor to execute the text acquisition method as described in the above Figures 1 - 5 shown embodiment. The specific execution process can refer to the Figures 1 - 5 specific description of the shown embodiment and will not be elaborated here.

[0090] The present application also provides a computer program product. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to execute the text acquisition method as described in the above Figures 1 - 5 shown embodiment. The specific execution process can refer to the Figures 1 - 5 specific description of the shown embodiment and will not be elaborated here.

[0091] Please refer to Figure 8 , which shows a structural block diagram of an electronic device provided by an exemplary embodiment of the present application. The electronic device in the present application may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected through the bus 150.

[0092] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts within the entire electronic device, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, as well as calling data stored in the memory 120, it executes various functions of the terminal 100 and processes data. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user pages, and application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.

[0093] The memory 120 may include a Random Access Memory (RAM), or may also include a Read-Only Memory (ROM). Optionally, the memory 120 includes a Non-Transitory Computer-Readable Storage Medium. The memory 120 can be used to store instructions, programs, code, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above various method embodiments, etc. The operating system can be an Android system, including a system developed based on the Android system in depth, an IOS system developed by Apple Inc., including a system developed based on the IOS system in depth, or other systems.

[0094] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party application programs run in the user space. In order to ensure that different third-party application programs can achieve better running effects, the operating system allocates corresponding system resources for different third-party application programs. However, there are also differences in the system resource requirements of different application scenarios in the same third-party application program. For example, in the local resource loading scenario, the third-party application program has higher requirements for the disk read speed; in the animation rendering scenario, the third-party application program has higher requirements for the GPU performance. The operating system and the third-party application program are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application program, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application program.

[0095] In order to enable the operating system to distinguish the specific application scenarios of third-party application programs, it is necessary to establish data communication between the third-party application programs and the operating system, so that the operating system can obtain the current scenario information of the third-party application programs at any time, and then perform targeted system resource adaptation based on the current scenario.

[0096] Among them, the input device 130 is used to receive input instructions or data. The input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data. The output device 140 includes but is not limited to a display device and a speaker, etc. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen.

[0097] The touch display screen can be designed as a full-screen, a curved screen or a special-shaped screen. The touch display screen can also be designed as a combination of a full-screen and a curved screen, or a combination of a special-shaped screen and a curved screen. The embodiments of the present application do not limit this.

[0098] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above drawings does not limit the electronic device. The electronic device may include more or fewer components than shown in the drawings, or combine some components, or have different component arrangements. For example, the electronic device also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a Wireless Fidelity (WiFi) module, a power supply, a Bluetooth module, etc., which will not be elaborated here.

[0099] In Figure 8 In the electronic device shown, the processor 110 can be used to call the text acquisition application program stored in the memory 120 and specifically perform the following operations:

[0100] Obtain an identification text set obtained by identifying a target audio, where the identification text set includes multiple identification texts obtained by respectively identifying the target audio by each audio recognition system in an audio recognition system set, and the audio recognition system set includes at least three audio recognition systems, and the audio recognition systems in the audio recognition system set correspond one-to-one with the identification texts in the identification text set;

[0101] Respectively obtain the edit distance between two identification texts in the identification text set, and based on the multiple edit distances, determine a target audio recognition system in the audio recognition system set;

[0102] In the identification text set, determine the target identification text corresponding to the target audio recognition system as the output text of the target audio.

[0103] In one embodiment, when the processor 110 executes the operation of respectively obtaining the edit distance between two identification texts in the identification text set and determining a target audio recognition system in the audio recognition system set based on the multiple edit distances, it specifically performs the following operations:

[0104] Respectively obtain the edit distance between two identification texts in the identification text set;

[0105] Based on the multiple edit distances, respectively determine the total edit distance corresponding to each audio recognition system;

[0106] Based on the total edit distance corresponding to each audio recognition system, determine a target audio recognition system in the audio recognition system set.

[0107] In one embodiment, when the processor 110 executes to obtain the edit distance between two recognized texts in the recognized text set respectively, the following operations are specifically performed:

[0108] Obtain the first recognized text corresponding to the first audio recognition system in the recognized text set, where the first audio recognition system is any audio recognition system in the audio recognition system set;

[0109] Respectively obtain the first edit distance between the first recognized text and the second recognized text in the recognized text set, where the second recognized text is other recognized texts in the recognized text set except the first recognized text.

[0110] In one embodiment, when the processor 110 executes to determine the target audio recognition system in the audio recognition system set based on the sum of the edit distances corresponding to each audio recognition system, the following operations are specifically performed:

[0111] Determine the minimum sum of edit distances in the sum of the edit distances corresponding to each audio recognition system, and determine the second audio recognition system corresponding to the minimum sum of edit distances in the audio recognition system set as the target audio recognition system.

[0112] In one embodiment, when the processor 110 executes to determine the sum of the edit distances corresponding to each audio recognition system based on the multiple edit distances respectively, the following operations are specifically performed:

[0113] Based on a preset threshold and the multiple first edit distances, determine the first audio recognition system as a candidate audio recognition system;

[0114] Determine the sum of the edit distances corresponding to the candidate audio recognition system based on the multiple first edit distances;

[0115] When executing to determine the target audio recognition system in the audio recognition system set based on the sum of the edit distances corresponding to each audio recognition system, the following operations are specifically performed:

[0116] Determine the target audio recognition system among the candidate audio recognition systems based on the sum of the edit distances corresponding to the candidate audio recognition system.

[0117] In one embodiment, when the processor 110 executes to determine the first audio recognition system as a candidate audio recognition system based on a preset threshold and the multiple first edit distances, the following operations are specifically performed:

[0118] When the multiple first edit distances are all less than or equal to the distance threshold, determine the first audio recognition system as a candidate audio recognition system.

[0119] In one embodiment, when the processor 110 determines the first audio recognition system as a candidate audio recognition system based on a preset threshold and multiple first edit distances, it specifically performs the following operations:

[0120] Obtain the average text length of the multiple recognition texts in the recognition text set;

[0121] Based on the average text length and multiple first edit distances, obtain multiple length ratios corresponding to the first audio recognition system;

[0122] When multiple length ratios are all less than or equal to a ratio threshold, determine the first audio recognition system as a candidate audio recognition system.

[0123] In one embodiment, when the processor 110 determines a target audio recognition system among the candidate audio recognition systems based on the total edit distance corresponding to the candidate audio recognition system, it specifically performs the following operations:

[0124] When there are at least two candidate audio recognition systems, determine the minimum total edit distance among the total edit distances corresponding to each candidate audio recognition system, and determine the third audio recognition system corresponding to the minimum total edit distance in the audio recognition system set as the target audio recognition system;

[0125] When there is no candidate audio recognition system, discard the target audio.

[0126] In the embodiment of the present application, a recognition text set obtained by recognizing a target audio is acquired. The recognition text set includes multiple recognition texts respectively obtained by each audio recognition system in the audio recognition system set for recognizing the target audio. The audio recognition system set includes at least three audio recognition systems. The audio recognition systems in the audio recognition system set correspond one-to-one with the recognition texts in the recognition text set. In the recognition text set, obtain the first recognition text corresponding to the first audio recognition system and the first edit distance. Based on a preset threshold and multiple first edit distances, determine the first audio recognition system as a candidate audio recognition system. The preset threshold can be a distance threshold, and when the recognition text length is too long, the preset threshold can be a ratio threshold. Screening the audio recognition system using the preset threshold further improves the accuracy of text acquisition. Based on multiple first edit distances, respectively determine the total edit distance corresponding to each candidate audio recognition system, and determine the candidate audio recognition system with the minimum total edit distance as the target audio recognition system. In the recognition text set, determine the target recognition text corresponding to the target audio recognition system as the output text of the target audio. By calculating the edit distance between different recognition texts to determine the target recognition text, the accuracy of text acquisition is improved.

[0127] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0128] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A text acquisition method, characterized in that, Including: Obtain a set of recognized texts obtained by recognizing a target audio. The set of recognized texts includes multiple recognized texts obtained by respectively recognizing the target audio by each audio recognition system in a set of audio recognition systems. The set of audio recognition systems includes at least three audio recognition systems, and the audio recognition systems in the set of audio recognition systems correspond one-to-one with the recognized texts in the set of recognized texts; Obtain the edit distance between two recognized texts in the set of recognized texts respectively. Based on multiple such edit distances, determine a target audio recognition system in the set of audio recognition systems; In the set of recognized texts, determine the target recognized text corresponding to the target audio recognition system as the output text of the target audio; The obtaining the edit distance between two recognized texts in the set of recognized texts respectively includes: Obtain the first recognized text corresponding to the first audio recognition system in the set of recognized texts, where the first audio recognition system is any audio recognition system in the set of audio recognition systems; Respectively obtain the first edit distance between the first recognized text and the second recognized text in the set of recognized texts, where the second recognized text is other recognized texts in the set of recognized texts except the first recognized text; Based on multiple such edit distances, determining a target audio recognition system in the set of audio recognition systems includes: Based on multiple such edit distances, respectively determine the total edit distance corresponding to each audio recognition system; Based on the total edit distance corresponding to each audio recognition system, determine a target audio recognition system in the set of audio recognition systems; The based on multiple such edit distances, respectively determining the total edit distance corresponding to each audio recognition system includes: Based on a preset threshold and multiple such first edit distances, determine the first audio recognition system as a candidate audio recognition system; where the preset threshold is used to judge the text recognition accuracy of the first audio recognition system and determine whether the first audio recognition system can be determined as a candidate audio recognition system; Based on multiple such first edit distances, determine the total edit distance corresponding to the candidate audio recognition system; The based on the total edit distance corresponding to each audio recognition system, determining a target audio recognition system in the set of audio recognition systems includes: Based on the total edit distance corresponding to the candidate audio recognition system, determine a target audio recognition system in the candidate audio recognition systems.

2. The method according to claim 1, wherein The smaller the total edit distance, the higher the accuracy of the corresponding audio recognition system. The target audio recognition system is the audio recognition system with the highest accuracy in the set of audio recognition systems.

3. The method according to claim 1 or 2, characterized in that, The based on the total edit distance corresponding to each audio recognition system, determining a target audio recognition system in the set of audio recognition systems includes: Determine the minimum total edit distance among the total edit distances corresponding to each audio recognition system, and in the set of audio recognition systems, determine the second audio recognition system corresponding to the minimum total edit distance as the target audio recognition system.

4. The method according to claim 1, wherein Determining the first audio recognition system as a candidate audio recognition system based on a preset threshold and multiple first edit distances includes: When multiple first edit distances are all less than or equal to the distance threshold, determining the first audio recognition system as a candidate audio recognition system.

5. The method according to claim 1, wherein Determining the first audio recognition system as a candidate audio recognition system based on a preset threshold and multiple first edit distances includes: Obtaining the average text length of the multiple recognition texts in the recognition text set; Based on the average text length and multiple first edit distances, obtaining multiple length ratios corresponding to the first audio recognition system; When multiple length ratios are all less than or equal to the ratio threshold, determining the first audio recognition system as a candidate audio recognition system.

6. The method according to claim 1, wherein Determining a target audio recognition system from the candidate audio recognition systems based on the total edit distance corresponding to the candidate audio recognition system includes: When there are at least two candidate audio recognition systems, determining the minimum total edit distance among the total edit distances corresponding to each candidate audio recognition system, and determining the third audio recognition system corresponding to the minimum total edit distance in the audio recognition system set as the target audio recognition system; When there is no candidate audio recognition system, discarding the target audio.

7. The method according to claim 1, wherein When the length of the recognition text is greater than the length threshold, the preset threshold is the ratio threshold; otherwise, the preset threshold is the distance threshold.

8. A text acquisition device, characterized in that, Including: A text set acquisition module, configured to acquire a recognition text set obtained by recognizing a target audio, where the recognition text set includes multiple recognition texts obtained by respectively recognizing the target audio by each audio recognition system in an audio recognition system set, the audio recognition system set includes at least three audio recognition systems, and the audio recognition systems in the audio recognition system set correspond to the recognition texts in the recognition text set one by one; A target system determination module, configured to respectively obtain the edit distance between two recognition texts in the recognition text set, and determine a target audio recognition system in the audio recognition system set based on multiple edit distances; An output text determination module, configured to determine the target recognition text corresponding to the target audio recognition system in the recognition text set as the output text of the target audio; An edit distance acquisition unit, configured to acquire a first recognition text corresponding to a first audio recognition system in the recognition text set, where the first audio recognition system is any audio recognition system in the audio recognition system set; Respectively obtaining a first edit distance between the first recognition text and a second recognition text in the recognition text set, where the second recognition text is other recognition texts in the recognition text set except the first recognition text; A distance total acquisition unit, configured to respectively determine the total edit distance corresponding to each audio recognition system based on multiple edit distances; The distance total acquisition unit includes: A candidate system determination subunit, configured to determine the first audio recognition system as a candidate audio recognition system based on a preset threshold and a plurality of the first edit distances; wherein the preset threshold is used to judge the text recognition accuracy rate of the first audio recognition system and to judge whether the first audio recognition system can be determined as a candidate audio recognition system; A distance sum acquisition subunit, configured to determine the total edit distance corresponding to the candidate audio recognition system based on the plurality of the first edit distances; A target system determination unit, configured to determine a target audio recognition system in the set of audio recognition systems based on the total edit distance corresponding to each audio recognition system; Specifically, the target system determination unit is configured to determine a target audio recognition system in the candidate audio recognition systems based on the total edit distance corresponding to the candidate audio recognition system.

9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, and the instructions are adapted to be loaded and executed by a processor to perform the steps of the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, Comprising: A processor and a memory; Wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A method and apparatus for controlling audio and text synchronization

    CN109275009A

  • Speech recognition method and device, electronic equipment and readable storage medium

    CN112992125A

  • Text error correction model generation method and system, text error correction method and system, equipment and medium

    CN113066494A

  • Sound and text realignment and information presentation method and device, electronic equipment and storage medium

    CN113761865A

  • Video voice recognition method and device, equipment and storage medium

    CN113838460A