A verification method for speech recognition, a computing device, and a storage medium

By generating candidate information and calculating similarity for speech error correction, the problem of errors in recognition results in voice interaction is solved and the user experience is improved.

CN113971952BActive Publication Date: 2025-07-11ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010724038.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-24
Publication Date
2025-07-11
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

In voice interaction services, errors in voice recognition results lead to poor user experience, making it difficult for users to understand the inconsistency between voice input and actual results.

Method used

By obtaining voice information, generating candidate information and splitting it, the similarity is calculated to determine whether to correct errors, and using the correlation information and pinyin similarity to accurately correct errors, reducing dependence on complex models and entity thesaurus.

Benefits of technology

It improves the accuracy of speech recognition, reduces users' misunderstandings about recognition results, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971952B_ABST
    Figure CN113971952B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a verification method, a computing device, and a storage medium for speech recognition. In the embodiment of the present application, the speech text to be verified corresponding to the speech information is determined; at least one candidate information is generated according to the associated information of the speech text to be verified; at least one split information is generated by splitting at least one candidate information according to the speech text to be verified; the similarity between each split information and the speech text to be verified is determined, and whether to correct the speech text to be verified is determined according to the similarity. Thus, in the case of correction, the speech text to be verified can be corrected, the speech recognition ability can be improved, and a better experience can be brought to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for verifying speech recognition, a computing device, and a storage medium. Background Art

[0002] Voice interaction refers to an intelligent voice response mode, where users can interact with machines through voice, and the machines can receive the user's voice and make responses. In voice interaction services, when users input voice, the voice input results may be displayed in real time. When the input results are incorrect, it will bring a bad experience to users and easily cause confusion to users. Summary of the Invention

[0003] Multiple aspects of this application provide a method for verifying speech recognition, a computing device, and a storage medium, which can perform speech recognition more accurately and improve the user experience.

[0004] An embodiment of this application provides a method for verifying speech recognition, including: obtaining speech information, and determining the speech text to be verified corresponding to the speech information; generating at least one candidate information according to the associated information of the speech text to be verified; splitting the at least one candidate information according to the speech text to be verified to generate at least one split information; determining the similarity between each split information and the speech text to be verified, and determining whether to correct the speech text to be verified according to the similarity.

[0005] An embodiment of this application also provides a method for verifying speech recognition, including: obtaining speech information, and determining the speech text to be verified corresponding to the speech information; determining the associated information of the speech text to be verified; in the case where the length of the speech text to be verified is less than or equal to the length of the associated information, splitting the speech text to be verified according to the length of each associated information to generate at least one split information; determining the similarity between each split information and the corresponding associated information, and determining whether to correct the speech text to be verified according to the similarity.

[0006] An embodiment of this application also provides a method for correcting speech recognition errors, the method further including: obtaining speech information, and determining the speech text to be corrected corresponding to the speech information; generating at least one candidate information according to the associated information of the speech text to be corrected; determining the similarity between the pinyin of the speech text to be corrected and the pinyin of each candidate information, and selecting the candidate information corresponding to the maximum similarity; modifying the speech text to be corrected according to the selected candidate information to perform error correction.

[0007] An embodiment of the present application further provides a method for verifying speech recognition, including: based on a navigation map, obtaining a query speech provided by a user, and determining navigation text to be verified corresponding to the query speech; generating at least one candidate name for a destination name according to the destination name in the navigation text to be verified; splitting the at least one candidate name according to the navigation text to be verified to generate at least one split name; determining the similarity between each split name and the navigation text to be verified, and determining whether to correct the navigation text to be verified according to the similarity.

[0008] An embodiment of the present application further provides a method for verifying speech recognition, including: receiving speech text to be verified; generating at least one candidate information according to the associated information of the speech text to be verified; splitting the at least one candidate information according to the speech text to be verified to generate at least one split information; determining the similarity between each split information and the speech text to be verified, and determining whether to correct the speech text to be verified according to the similarity.

[0009] An embodiment of the present application further provides a method for verifying order-taking speech recognition, including: obtaining order-taking speech, and determining order-taking text to be verified corresponding to the order-taking speech; generating at least one candidate name for a food item name according to the food item name in the order-taking text to be verified; splitting the at least one candidate name according to the order-taking text to be verified to generate at least one split name; determining the similarity between each split name and the order-taking text to be verified, and determining whether to correct the order-taking text to be verified according to the similarity.

[0010] An embodiment of the present application further provides a method for verifying ticket-purchasing speech recognition, including: obtaining ticket-purchasing speech, and determining ticket-purchasing text to be verified corresponding to the ticket-purchasing speech; generating at least one candidate name for a ticket-purchasing name according to the ticket-purchasing name in the ticket-purchasing text to be verified; splitting the at least one candidate name according to the ticket-purchasing text to be verified to generate at least one split name; determining the similarity between each split name and the ticket-purchasing text to be verified, and determining whether to correct the ticket-purchasing text to be verified according to the similarity.

[0011] An embodiment of the present application further provides a method for verifying on-demand speech recognition, including: obtaining on-demand speech, and determining on-demand text to be verified corresponding to the on-demand speech; generating at least one candidate name for a video name according to the video name in the on-demand text to be verified; splitting the at least one candidate name according to the on-demand text to be verified to generate at least one split name; determining the similarity between each split name and the on-demand text to be verified, and determining whether to correct the on-demand text to be verified according to the similarity.

[0012] An embodiment of the present application also provides a method for correcting errors in speech recognition, including: obtaining speech information, and determining the speech text to be corrected corresponding to the speech information; generating at least one candidate information according to the associated information of the speech text to be corrected; splitting the at least one candidate information according to the speech text to be corrected to generate at least one split information; determining the similarity between each split information and the speech text to be corrected, and when the similarity is greater than a threshold, determining to correct the speech text to be corrected.

[0013] An embodiment of the present application also provides a computing device, including: a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program to: obtain speech information, and determine the speech text to be verified corresponding to the speech information; generate at least one candidate information according to the associated information of the speech text to be verified; split the at least one candidate information according to the speech text to be verified to generate at least one split information; determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0014] An embodiment of the present application also provides a computing device, including: a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program to: obtain speech information, and determine the speech text to be verified corresponding to the speech information; determine the associated information of the speech text to be verified; in the case where the length of the speech text to be verified is less than or equal to the length of the associated information, split the speech text to be verified according to the length of each associated information to generate at least one split information; determine the similarity between each split information and the corresponding associated information, and determine whether to correct the speech text to be verified according to the similarity.

[0015] An embodiment of the present application also provides a computing device, including: a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program to: obtain speech information, and determine the speech text to be corrected corresponding to the speech information; generate at least one candidate information according to the associated information of the speech text to be corrected; determine the similarity between the pinyin of the speech text to be corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; correct the speech text to be corrected according to the selected candidate information.

[0016] An embodiment of the present application further provides a computing device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to: based on a navigation map, obtain a query voice provided by a user, and determine verified query text corresponding to the query voice; generate at least one candidate name of a destination name according to the destination name in the verified query text; split the at least one candidate name according to the verified query text to generate at least one split name; determine the similarity between each split name and the verified query text, and determine whether to correct the verified query text according to the similarity.

[0017] An embodiment of the present application further provides a computing device, including: a memory, a processor, and a communication component; the memory is used to store a computer program; the processor is used to execute the computer program to: generate at least one candidate information according to the associated information of the verified voice text; split the at least one candidate information according to the verified voice text to generate at least one split information; determine the similarity between each split information and the verified voice text, and determine whether to correct the verified voice text according to the similarity; the communication component is used to receive the verified voice text.

[0018] An embodiment of the present application further provides a computing device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to: obtain an ordering voice, and determine verified ordering text corresponding to the ordering voice; generate at least one candidate name of a food item name according to the food item name in the verified ordering text; split the at least one candidate name according to the verified ordering text to generate at least one split name; determine the similarity between each split name and the verified ordering text, and determine whether to correct the verified ordering text according to the similarity.

[0019] An embodiment of the present application further provides a computing device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to: obtain a ticket-purchasing voice, and determine verified ticket-purchasing text corresponding to the ticket-purchasing voice; generate at least one candidate name of a ticket-purchasing name according to the ticket-purchasing name in the verified ticket-purchasing text; split the at least one candidate name according to the verified ticket-purchasing text to generate at least one split name; determine the similarity between each split name and the verified ticket-purchasing text, and determine whether to correct the verified ticket-purchasing text according to the similarity.

[0020] An embodiment of the present application further provides a computing device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to: obtain an on-demand voice, and determine the to-be-verified on-demand text corresponding to the on-demand voice; generate at least one candidate name of the video name according to the video name in the to-be-verified on-demand text; split the at least one candidate name according to the to-be-verified on-demand text to generate at least one split name; determine the similarity between each split name and the to-be-verified on-demand text, and determine whether to correct the to-be-verified on-demand text according to the similarity.

[0021] An embodiment of the present application further provides a computing device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to: obtain voice information, and determine the to-be-corrected voice text corresponding to the voice information; generate at least one candidate information according to the associated information of the to-be-corrected voice text; split the at least one candidate information according to the to-be-corrected voice text to generate at least one split information; determine the similarity between each split information and the to-be-corrected voice text, and when the similarity is greater than a threshold, determine to correct the to-be-corrected voice text.

[0022] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by one or more processors, the one or more processors are caused to implement the steps in the above method.

[0023] In an embodiment of the present application, determine the to-be-verified voice text corresponding to the voice information; generate at least one candidate information according to the associated information of the to-be-verified voice text; split the at least one candidate information according to the to-be-verified voice text to generate at least one split information; determine the similarity between each split information and the to-be-verified voice text, and determine whether to correct the to-be-verified voice text according to the similarity. Thus, in the case of correction, the to-be-verified voice text can be corrected, the voice recognition ability can be improved, and a better experience can be brought to the user.

[0024] At the same time, since complex methods such as constructing an entity word library and training a model can be abandoned, the implementation cost is low and the implementation effect is good. Description of the Drawings

[0025] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0026] Figure 1Schematic structural diagram of a voice recognition verification system according to an exemplary embodiment of the present application;

[0027] Figure 2 Schematic flowchart of a voice recognition verification method according to an exemplary embodiment of the present application;

[0028] Figure 3 Schematic diagram of candidate information according to another exemplary embodiment of the present application;

[0029] Figure 4 Schematic diagram of candidate information segmentation according to another exemplary embodiment of the present application;

[0030] Figure 5 Schematic diagram of the segmentation of the voice text to be verified according to another exemplary embodiment of the present application;

[0031] Figure 6 Schematic diagram of correcting the recognition result of the online map according to another exemplary embodiment of the present application;

[0032] Figure 7 Schematic flowchart of a voice recognition verification method according to an exemplary embodiment of the present application;

[0033] Figure 8 Schematic flowchart of a voice recognition error correction method according to an exemplary embodiment of the present application;

[0034] Figure 9 Schematic flowchart of a voice recognition verification method according to an exemplary embodiment of the present application;

[0035] Figure 10 Schematic flowchart of a voice recognition verification method according to an exemplary embodiment of the present application;

[0036] Figure 11 Schematic structural diagram of a voice recognition verification device according to another exemplary embodiment of the present application;

[0037] Figure 12 Schematic structural diagram of a voice recognition verification device according to another exemplary embodiment of the present application;

[0038] Figure 13 Schematic structural diagram of a voice recognition error correction device according to another exemplary embodiment of the present application;

[0039] Figure 14 Schematic structural diagram of a voice recognition verification device according to another exemplary embodiment of the present application;

[0040] Figure 15 Schematic structural diagram of a voice recognition verification device according to another exemplary embodiment of the present application;

[0041] Figure 16 Schematic structural diagram of a computing device provided by an exemplary embodiment of the present application;

[0042] Figure 17 Schematic structural diagram of a computing device provided by another exemplary embodiment of the present application;

[0043] Figure 18 Schematic structural diagram of a computing device provided by another exemplary embodiment of the present application;

[0044] Figure 19 Schematic structural diagram of a computing device provided by another exemplary embodiment of the present application;

[0045] Figure 20 Schematic structural diagram of a computing device provided by another exemplary embodiment of the present application. Detailed implementation manners

[0046] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0047] According to the foregoing, voice interaction generally can be divided into the following links: speech recognition, semantic understanding, dialogue management, natural language generation, and speech synthesis. In a voice interaction service, when a user uses voice input, the speech recognition result will be displayed in real time. When the recognition result is incorrect, it may cause the content in the voice search box to be inconsistent with the final service search result, which will cause confusion to the user.

[0048] Embodiments of the present application can rewrite the echoed content based on the search result with a relatively small difference from the actual voice of the user, improve the consistency of the result, and weaken the user's incorrect perception of the recognition result.

[0049] The following will describe in detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.

[0050] Figure 1 Schematic structural diagram of a verification system for speech recognition provided by an exemplary embodiment of the present application. As Figure 1 shown, the system 100 may include: a first device 101 and a second device 102.

[0051] Among them, the first device 101 can be a device with certain computing capabilities, which can implement the function of sending data to the second device 102 and obtaining response data from the second device 102. The basic structure of the first device 101 may include: at least one processor. The number of processors may depend on the configuration and type of the device with certain computing capabilities. The device with certain computing capabilities may also include a memory, which may be volatile, such as RAM, or non-volatile, such as read-only memory (ROM), flash memory, etc., or may include both types at the same time. Generally, an operating system (OS), one or more application programs, and program data may be stored in the memory. In addition to the processing unit and the memory, the device with certain computing capabilities also includes some basic configurations, such as a network card chip, an IO bus, a display component, and some peripheral devices. Optionally, some peripheral devices may include, for example, a keyboard, a stylus, etc. Other peripheral devices are well known in the art and will not be elaborated here. Optionally, the first device 101 may be a smart terminal, such as a mobile phone, a desktop computer, a notebook, a tablet computer, etc.

[0052] The second device 102 refers to a device that can provide computing and processing services in a network virtual environment, and can be a device that uses the network to correct speech and text errors. Physically, the second device 102 can be any device that can provide computing services, respond to service requests, and return data processing results. For example, it can be a cloud server, a cloud host, a virtual center, a conventional server, etc. The composition of the second device 102 mainly includes a processor, a hard disk, a memory, a system bus, etc., which is similar to a general computer architecture.

[0053] In the embodiment of the present application, the first device 101 receives the user's speech information and sends the speech to the second device 102. The second device 102 obtains the speech information from the first device 101, determines the speech text to be verified corresponding to the speech information; generates at least one candidate information according to the associated information of the speech text to be verified; splits at least one candidate information according to the speech text to be verified to generate at least one split information; determines the similarity between each split information and the speech text to be verified, and determines whether to correct the speech text to be verified according to the similarity.

[0054] In addition, when the second device 102 determines to correct the speech text to be verified, it corrects the speech text to be verified, and then sends the corrected speech text to be verified to the first device 101 for display.

[0055] It should be noted that in addition to the first device 101 being able to send voice messages, the first device 101 can also directly recognize the voice messages and send the recognition results to the second device 102. After the second device 102 receives the recognition results, that is, the voice text to be verified, it determines whether to perform error correction according to the subsequent steps. If error correction is performed, the corrected voice text to be verified is sent to the first device 101.

[0056] In the application scenario of the embodiment of the present application, such as in the online map scenario. The user can obtain online map information from the second device 102, such as a server, through the map APP (Application) installed on the first device, such as a mobile phone, and display the online map through the map APP on the mobile phone and indicate the user's current geographical location. The user can click the voice input button on the map APP, such as the "search box", to perform voice input. The APP obtains the voice in response to the voice input operation through the mobile phone, and can send the voice to the server. At the same time, the mobile phone can recognize the voice and display the recognition result. Or the mobile phone can directly recognize the voice, display the recognition result, and send the recognition result to the server so that the server determines whether to perform error correction according to the recognition result.

[0057] If the server receives the voice, it performs voice recognition, obtains the recognition result, that is, the voice text to be verified, or directly receives the recognition result sent by the mobile phone. The server obtains the destination therein according to the recognition result. It can also obtain its starting point. The server searches for the corresponding destination according to the destination. For example, if the recognized destination is "Traverse Supermarket", and the searched one is "Convenience Supermarket (Tianjin Airport Store)", that is, the associated information. The server can generate candidate information according to the search result, such as Convenience Supermarket, Tianjin Airport Store, Convenience Supermarket Tianjin Airport Store, Tianjin Airport Store Convenience Supermarket. The server then segments the candidate information according to the word length of the recognition result, and can segment it into "Convenience Supermarket" and "Tianjin Airport Store", etc. The server can calculate the similarity between the segmented words and the recognition result respectively. When the similarity exceeds the threshold, the corresponding segmented word is used to replace the recognition result. For example, if the similarity between "Convenience Supermarket" and "Traverse Supermarket" is greater than the threshold, then "Convenience Supermarket" is used to replace "Traverse Supermarket" as the final voice text and sent to the mobile phone for display.

[0058] In addition, if it is determined according to the foregoing that the speech text to be verified does not need to be corrected, the server can continue with the correction according to the following method. For example, when the number of words in the search result is less than the number of words in the recognition result, the recognition result can be directly segmented according to the number of words in the search result, and the similarity between the segmented words and the search result can be calculated. The corresponding words with a similarity greater than the threshold are selected, and these words can be replaced by the corresponding search results. The specific process is similar to that described above and will not be elaborated here.

[0059] If it is determined according to the foregoing that the speech text to be verified does not need to be corrected, the server can continue with the correction according to the following method. For example, calculate the similarity between the pinyin string of the recognition result and the pinyin string of each of the foregoing candidate information, and obtain the corresponding longest common subsequence. According to the text corresponding to the longest common subsequence, replace the corresponding text in the recognition result to complete the correction.

[0060] It should also be noted that, in addition to being applied to online maps, the embodiments of the present application can also be applied to application scenarios such as smart homes and mobile phone assistants. In the smart home scenario, users can perform voice interactions with smart home appliances, and the voice recognition results can be displayed on smart home appliances such as TV screens. In the mobile phone assistant scenario, users can perform voice interactions with mobile phones and tablets, and the voice recognition results can be displayed on their screens. The specific implementation methods are similar to those described above and will not be elaborated here. In these scenarios, the voice recognition results of users can be corrected and the echo content can be modified to improve the user experience.

[0061] In the above-mentioned embodiment of the present application, the first device 101 and the second device 102 are network-connected, and this network connection can be a wireless connection. If the first device 101 and the second device 102 are communicatively connected, the network mode of this mobile network can be any one of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), WiMax, 5G, etc.

[0062] Next, in combination with the method embodiments, the verification process of speech recognition will be described in detail.

[0063] Figure 2 It is a schematic flowchart of a method for verifying speech recognition according to an exemplary embodiment of the present application. The method 200 provided by the embodiments of the present application is executed by a computing device, such as a smart terminal or a server, and more specifically, it can be a mobile phone or a server. The method 200 includes the following steps:

[0064] 201: Obtain speech information and determine the speech text to be verified corresponding to the speech information.

[0065] 202: Generate at least one candidate information according to the associated information of the speech text to be verified.

[0066] 203: Split at least one candidate information according to the speech text to be verified to generate at least one split information.

[0067] 204: Determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0068] The following elaborates on the above steps in detail:

[0069] 201: Obtain speech information and determine the speech text to be verified corresponding to the speech information.

[0070] Among them, in the case where the execution entity is an intelligent terminal, the acquisition method of obtaining speech information may include: responding to the user's speech input operation to obtain speech information.

[0071] For the server as the execution entity, the method of obtaining speech information may be to receive the speech information sent by the intelligent terminal.

[0072] The method of determining the speech text to be verified may be to use an ASR (Automatic Speech Recognition) model at the same time for speech recognition to obtain the corresponding text.

[0073] Among them, the ASR model takes speech as the research object, and through speech signal processing and pattern recognition, enables the machine to automatically recognize and understand the speech spoken by humans.

[0074] The recognition principle of the ASR model is to preprocess the input speech, then extract the features of the speech, and on this basis, establish the templates required for speech recognition. This model can be obtained through training. The specific training process is prior art and will not be elaborated here. After generation, it can be deployed at the intelligent terminal for speech recognition.

[0075] In addition, this model can also be deployed on the server side. When deployed on the server side, the speech information obtained can be sent to the server side through the application program deployed on the intelligent terminal, so that this model in the server side performs speech recognition, so that the server side obtains the corresponding speech text to be verified.

[0076] For example, as described above, a user can perform a route search through a map APP installed on a mobile phone. A voice information input button can be provided on the interface of the map APP, and the user clicks the voice input button to perform voice input. The APP obtains the user's voice information through a sound collection device configured on the mobile phone, such as a microphone. And it can perform voice recognition through the ASR model provided in the APP to obtain the voice text to be verified.

[0077] It should be noted that after obtaining the voice text to be verified, the intelligent terminal can also send the voice text to be verified to the server.

[0078] 202: Generate at least one candidate information according to the associated information of the voice text to be verified.

[0079] Among them, the associated information refers to the information related to the voice text to be verified, which can be the information that can represent the semantics of the real voice text. The so-called real voice text refers to the voice text that the user really wants to express. For example, the voice text to be verified can be "traverse the supermarket store", and the associated information can be "convenient supermarket store".

[0080] The acquisition method of the associated information can be obtained in the form of search.

[0081] Specifically, the method 200 further includes: searching for the search result corresponding to the voice text to be verified according to the voice text to be verified as the associated information.

[0082] For example, as described above, for the map APP installed on the user's mobile phone, after obtaining the voice text to be verified, such as "traverse the supermarket store", it can send the voice text to be verified to the server so that the server performs an address search according to the voice text to be verified. Obtain multiple pieces of associated information related to the voice text to be verified, such as other addresses.

[0083] It should also be noted that for the map APP, when searching for associated information, it can directly call the corresponding search API (Application Programming Interface) or search function of the server to obtain the corresponding search result from the server.

[0084] When the user's voice information is a paragraph, or rather, the user's voice information is not only the destination, it is necessary to determine the keyword from the voice text to be verified corresponding to the voice information. For different application scenarios, the keyword can represent the user's purpose. For an online map, the keyword can be the destination. For online song listening, the song name is the keyword. In the smart home scenario, for a smart TV, the TV program is the keyword.

[0085] Specifically, according to the speech text to be verified, search for the search results corresponding to the speech text to be verified, including: obtaining the semantics therein according to the speech text to be verified, and determining the keywords of the speech text to be verified; obtaining the search results according to the keywords.

[0086] For example, as described above, whether it is a mobile phone or a server, after obtaining the speech text to be verified, such as "Navigate to Traverse Supermarket", according to the semantics, the keyword can be determined as "Traverse Supermarket", and the search results are obtained according to "Traverse Supermarket".

[0087] In addition, in addition to searching, relevant information corresponding to the speech text to be verified can also be directly obtained locally. For example, a set of preset words belonging to the same attribute as the speech text to be verified is stored locally. This attribute can represent the type of the word, such as all being addresses, foods, fruits, songs, etc.

[0088] For the determination method of the attribute of the speech text to be verified, it can be determined according to the service currently provided to the user. For example, for a map service, the attribute can be an address, and for a song service, the attribute can be a song. Or, a service, such as the service provided by an APP installed on a smart terminal, can have a set of preset words, and it is also possible to directly search for similar words from the corresponding set of words without knowing the attribute.

[0089] In addition, the attribute can also be determined according to the semantics of the speech text to be verified. It should be noted that although the speech text to be verified may have problems with inaccurate recognition, the semantics and attributes can be determined through the entire speech text to be verified.

[0090] For the server side, the corresponding relevant information can be directly searched according to the speech text to be verified, or the corresponding relevant information can be obtained according to the set of preset words, which will not be elaborated here.

[0091] Among them, the candidate information refers to the information obtained by splitting and recombining according to the relevant information. The specific generation method is as follows:

[0092] Specifically, according to the relevant information of the speech text to be verified, generate at least one candidate information, including: splitting the relevant information to generate word segmentation information; generating at least one candidate information according to the word segmentation information.

[0093] Figure 3 Fig. 300 shows a schematic diagram of the candidate information. For example, as described above, such as Figure 3As shown, for a mobile phone, when the map APP obtains a search result, such as, "Convenience Supermarket (Tianjin Airport Store)". First, the punctuation marks can be removed from the search result. If there are English letters or pinyin in it, the case can be unified. Then, "Convenience Supermarket Tianjin Airport Store" can be obtained. It can be split according to punctuation marks. For example, according to the parentheses, the search result can be split into "Convenience Supermarket" and "Tianjin Airport Store". The word segmentation information can be directly combined with these two pieces of word segmentation information to obtain candidate information, such as 1. Convenience Supermarket, 2. Tianjin Airport Store, 3. Convenience Supermarket Tianjin Airport Store, and 4. Tianjin Airport Store Convenience Supermarket, etc.

[0094] It should be noted that for the server, it is also implemented as described above, and the specific implementation method will not be elaborated here.

[0095] 203: According to the speech text to be verified, split at least one candidate information to generate at least one split information.

[0096] Among them, the splitting method can include: according to the length of the speech text to be verified, split each candidate information to generate at least one split information as the split information.

[0097] For example, as described above, when the number of characters in the speech text to be verified is less than or equal to the number of characters in the candidate information, each candidate information can be split according to the number of characters in the speech text to be verified.

[0098] Among them, Figure 4 Figure 400 showing the splitting of candidate information is shown. As Figure 4 shown, after the map APP on the mobile phone obtains the candidate information, since the number of characters in the speech text to be verified is 5, which is less than or equal to the number of characters in the candidate information. Therefore, the candidate information is split according to 5 characters. For example, "Convenience Supermarket" can be directly split into "Convenience Supermarket" (because the candidate information is also 5 characters, and the candidate information can be directly used as the split information). "Tianjin Airport Store" can also be directly split into "Tianjin Airport Store". "Convenience Supermarket Tianjin Airport Store" can be split into "Li Supermarket Store Tian", "Supermarket Store Tianjin", and "Store Tianjin Airport", etc.

[0099] It should be noted that in order to improve the segmentation effect, after obtaining the segmentation information or the splitting information, redundant splitting information can be removed. It is also possible to remove unreasonable or invalid splitting information based on the positions of punctuation marks in the search results. For example, from the positions of punctuation marks in the search result "Convenient Supermarket (Tianjin Airport Store)", it can be known that the most appropriate and effective segmentation method is to split "Convenient Supermarket" and "Tianjin Airport Store". If it is split into "Supermarket Tian" and "Store Tianjin Air", it is not very reasonable, so these unreasonable splitting information can be removed.

[0100] For the server side, it is also implemented as described above, and the specific implementation method will not be elaborated here.

[0101] It should also be noted that when the number of characters in the speech text to be verified is greater than or equal to the number of characters in the candidate information, the speech text to be verified can be segmented according to the number of characters in the candidate information to generate segmentation information, so that the number of characters in the segmentation information is the same as that of the candidate information. Then, the similarity between the segmentation information and the candidate information can be determined. When the similarity is greater than the threshold, the corresponding candidate information is used to replace the corresponding segmentation information to complete error correction and obtain the error-corrected speech text to be verified.

[0102] 204: Determine the similarity between each splitting information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0103] Among them, the determination method of similarity can be determined through a similarity algorithm. The similarity algorithm can include: Euclidean algorithm, Manhattan distance, cosine similarity, etc.

[0104] In addition, the similarity can also be determined according to the number of characters, pinyin, and tones, etc. For example, if both the splitting information and the speech text to be verified have 5 characters and they differ by 1 character, then their similarity can be 4 / 5 = 0.8. Or if the pinyins are the same but the tones are different, each character's corresponding pinyin can be regarded as a whole, and the tone can be regarded as half of a character's pinyin, or one-third of a character's pinyin, etc., which can be determined according to requirements. Then when a splitting information has exactly the same pinyin as the speech text to be verified, but the tone of the pinyin of one character in the speech text to be verified is different from the tone of the corresponding character's pinyin in the splitting information, such as the speech text to be verified "Bianli Supermarket Store" and the splitting information "Convenient Supermarket Store". At this time, the tone of the character "Li" in the speech text to be verified is the third tone, and the tone of the character "Li" in the splitting information is the fourth tone. At this time, the tone can be set as half of a character's pinyin. For the pinyin of a five-character speech text to be verified (or splitting information), there are a total of 5 pinyins + 5 * 0.5 tones = 7.5 pinyins. Then the similarity can be 7 / 7.5 = 0.93.

[0105] Therefore, both the intelligent terminal and the server can determine the similarity between the voice text to be verified and each split information according to the above method. After determining the similarity, it can be determined whether to perform error correction on the voice text to be verified.

[0106] Specifically, whether to correct the voice and text to be verified is determined based on the similarity, including: when the similarity is greater than a threshold, determining to correct the voice and text to be verified.

[0107] For example, according to the above, after determining the similarity, whether it is a mobile phone or a server, it is determined whether the similarity is greater than a threshold, such as 0.9. If it is greater than 0.9, it is determined that the voice text to be verified needs to be corrected. If it is less than or equal to the threshold, the voice text to be verified does not need to be corrected.

[0108] The specific method of error correction may be: determining the split information corresponding to the similarity greater than a threshold, and replacing the voice text to be verified with the determined split information for display.

[0109] For example, according to the above, whether it is a mobile phone or a server, after determining that the voice and text to be verified need to be corrected, the corresponding split information will be replaced with the voice and text to be verified, such as replacing the voice and text to be verified "Traversal Supermarket Store" with the split information "Convenience Supermarket Store". At this time, if it is a mobile phone, the map APP can directly display the corrected voice and text to be verified (also called corrected voice and text), that is, "Convenience Supermarket Store" through the mobile phone.

[0110] If it is a server, the server will send the corrected voice text to be verified to the map APP of the mobile phone for display.

[0111] It should be noted that about 60% of the errors in the speech and text to be verified can be corrected by the above method. In order to improve the accuracy of error correction, if it is determined that the speech and text to be verified do not need to be corrected by the above method, the following method can be used to correct the errors to improve accuracy.

[0112] Specifically, when it is determined that no error correction is required for the voice text to be verified, the method also includes: when the length of the voice text to be verified is less than or equal to the length of the associated information, segmenting the voice text to be verified according to the length of each associated information to generate at least one segmentation information; determining the similarity between each segmentation information and the corresponding associated information, and determining whether to correct the voice text to be verified based on the similarity.

[0113] It should be noted that when the associated information obtained is a short fragment, that is, the number of characters is small, less than or equal to the speech text to be verified, error correction needs to be performed according to this method. For the associated information in the form of short fragments, there may be multiple independent short fragments in this associated information. Taking each short fragment as a unit, error correction is performed.

[0114] Among them, Figure 5 FIG. 500 shows a schematic diagram of the segmentation of the speech text to be verified for short fragments. For example, as described above, whether it is a mobile phone or a server, search results can be obtained, such as three independent short fragments: 1-Hugong Road, 1-Hugong Road, and 1-Songjiang District, each of which is three characters. At this time, whether it is a mobile phone or a server, the speech text to be verified "Songjiang District Hugong Road" is segmented according to 3 characters, and segmentation information such as "Songjiang District" and "Hugong Road" can be obtained (the specific segmentation method is similar to that described above and will not be elaborated here). Then, the similarity between each segmentation information and the short fragment is calculated to determine whether to perform error correction on the speech text to be verified.

[0115] It should be noted that the calculation of similarity has been described above and will not be elaborated here.

[0116] In addition, the error correction process may include: when the similarity is greater than the threshold, the associated information is used to replace the segmentation information in the speech text to be verified to perform error correction on the speech text to be verified.

[0117] Since the error correction process has been described above, it will not be elaborated here. Only to note: when the similarity is greater than the threshold, the corresponding short fragment can be used to replace the segmentation information for error correction. For example, the short fragment "Songjiang District" is used to replace the segmentation information "Songjiang District", and the short fragment "Hugong Road" is used to replace the segmentation information "Hugong Road", etc. As Figure 5 shown, the speech text to be verified "Songjiang District Hugong Road" is corrected to "Songjiang District Hugong Road".

[0118] Figure 6Figure 600 shows the view of the map APP presenting the corrected voice text to be verified. In this view 600, the user can perform voice input by clicking the voice input button "605" on the interface 601 provided by the map APP. While the user is inputting voice, the mobile phone can directly perform voice translation through the local ASR module and display the translated voice text to be verified on the interface 601, that is, "Songjiang District Hugong Road" in the display box 602. Through the above error correction method, the voice text to be verified in the display box 602 can be corrected to "Songjiang District Hugong Road". At this time, for this error correction result, that is, the corrected voice text to be verified, an online map path search can be performed. That is, the map APP provides the interface 602 for displaying the search path and simultaneously shows this error correction result, that is, "Songjiang District Hugong Road" in the display box 606. Meanwhile, the specific address of this error correction result can also be shown below the interface 602, such as "Songjiang District Hugong Road" in the display position 604.

[0119] It should be noted that about 10% of the errors in the voice text to be verified can be corrected through the above short segment method. To further improve the error correction accuracy, in the case where it is determined not to correct the voice text to be verified through the above two methods, the following method can be continued to perform error correction to provide accuracy.

[0120] Specifically, in the case where it is determined not to correct the voice text to be verified, the method 200 further includes: determining the similarity between the pinyin of the voice text to be verified and the pinyin of each candidate information, and selecting the candidate information corresponding to the maximum similarity; modifying the voice text to be verified according to the selected candidate information to perform error correction.

[0121] Among them, the method of determining the similarity between the pinyin of the voice text to be verified and the pinyin of each candidate information refers to determining the longest common subsequence of the pinyin strings, and this longest common subsequence can be continuous pinyin or discontinuous pinyin. For example, the pinyin "bianlichaoshidian" of the voice text to be verified "bianlichaoshidian", the pinyin "bianlichaoshidian" of a candidate information "bianlichaoshidian", and the pinyin "tianjinkonggangdian" of another candidate information "tianjinkonggangdian". Then the longest common subsequence can be determined as "bianlichaoshidian". Then, whether it is a smart terminal, such as a mobile phone, or a server, after determining the similarity, according to the longest common subsequence "bianlichaoshidian", the corresponding candidate information is used to replace the voice text to be verified.

[0122] It should be noted that if the longest common subsequence pinyin string or pinyin is part of the pinyin of the candidate information, for example, "bianli" is part of the pinyin "bianlichaoshidian" of the candidate information, then replace the corresponding characters in the speech text to be verified with the characters corresponding to this part of the pinyin string or pinyin. For example, replace "bianli" in the candidate information with "bianli" in the speech text to be verified.

[0123] In addition, if there are two longest common subsequences and they are the same, the similarity can be further determined by tones or a tone comparison can be made. For example, compare the tones of the pinyin strings or pinyins of these two identical longest common subsequences with the tones of the pinyin string or pinyin of the speech text to be verified respectively, and determine the longest common subsequence whose tone of the pinyin string or pinyin is the same as or most similar to that of the speech text to be verified. The determination method is similar to the method for determining the similarity of characters in the previous text and will not be elaborated here.

[0124] In addition to determining by tones, it can also be determined according to the characters corresponding to the pinyin. For example, compare the characters of the pinyin strings or pinyins of these two identical longest common subsequences with the characters of the pinyin string or pinyin of the speech text to be verified respectively, and determine the longest common subsequence whose characters of the pinyin string or pinyin are the same as or most similar to those of the speech text to be verified. The determination method is similar to the method for determining the similarity of characters in the previous text and will not be elaborated here.

[0125] It should also be noted that when the two longest common subsequences are not completely different but have the same length. For example, the pinyin of the speech text to be verified "bianlichaoshidian" is "bianlichaoshidian". And the two longest common subsequences are "bianlichaoshi" and "bianlishidian" respectively. It can also be determined which longest common subsequence to select for error correction according to the above method and will not be elaborated here. If the characters corresponding to the same parts in the two longest common subsequences are also the same, the different parts of the characters in the speech text to be verified can be directly corrected by these two longest common subsequences.

[0126] When the two longest common subsequences are completely different but have the same length. It can also be determined which longest common subsequence to select for error correction according to the above method. This will not be elaborated here.

[0127] Since the embodiments of the present application can perform error correction without relying on a model and constructing a complex entity word library, and perform error correction in a low-cost manner, it brings a large improvement in accuracy.

[0128] In addition, in the embodiments of the present application, the similarity can be calculated by using a model. That is, a similarity model is trained and constructed through the similarity algorithm described above to determine the similarity.

[0129] In addition to correcting the speech text to be verified, the above keywords can also be corrected, that is, the user's intention is corrected to make the search results more accurate.

[0130] Regardless of which of the above correction methods, they can be applied to smart terminals and servers, and will not be elaborated here.

[0131] Specifically, the method 200 further includes: obtaining a plurality of words associated with the keyword according to the keyword; determining the similarity between the keyword and each word, and selecting the words with similarity greater than the threshold to correct the keyword.

[0132] Among them, the way to obtain a plurality of words associated with the keyword may include:

[0133] 1) Obtain all entity words in the entity word list of the same type as the keyword, and determine whether to correct the keyword according to the similarity.

[0134] For example, as described above, since in the online map scenario, the type of the keyword is the destination. Then, the entity word list of the destination can be used to determine whether the keyword needs to be corrected, and the similarity between the entity word and the keyword is determined respectively. In the case where the similarity is the largest and higher than the threshold, the corresponding entity word can be used to replace the keyword.

[0135] 2) According to the principle of method 1), a translation model can be trained. The corrected words and keywords in method 1) can be used as training data to train the translation model. Whether the keyword needs to be corrected is determined through the trained model, and the corrected word is output if there is correction.

[0136] It should be noted that the model can be for multiple scenarios or for one scenario. For multiple scenarios, the training data needs to involve data of multiple scenarios, that is, the corrected words and keywords corresponding to multiple scenarios.

[0137] The above modification of the keyword can be applied to smart terminals and servers, and will not be elaborated here.

[0138] After correcting the keyword, search results are obtained according to the corrected keyword, that is, according to the modified keyword. Since this has been elaborated above, it will not be elaborated here.

[0139] It should be noted that, in addition to the intelligent terminal or the server being the execution subject of this method 200, step 201 can also be executed by the intelligent terminal, and steps 202-204 can be executed by the server. At this time, the intelligent terminal can send the speech text to be verified to the server, and then the server performs the steps of determining whether to correct the error. The process of whether to correct the error is as described above and will not be elaborated here.

[0140] Based on the same technical problems described above, Figure 7 FIG. shows a schematic flowchart of a verification method for speech recognition provided by another exemplary embodiment of the present application. The method 700 provided by the embodiment of the present application is executed by the above-mentioned intelligent terminal or server, and more specifically, it can be a mobile phone or a server. As Figure 7 shown, the method 700 includes the following steps:

[0141] 701: Obtain speech information and determine the speech text to be verified corresponding to the speech information.

[0142] 702: Determine the associated information of the speech text to be verified.

[0143] 703: When the length of the speech text to be verified is less than or equal to the length of the associated information, segment the speech text to be verified according to the length of each associated information to generate at least one segmentation information.

[0144] 704: Determine the similarity between each segmentation information and the corresponding associated information, and determine whether to correct the speech text to be verified according to the similarity.

[0145] Since the specific implementation manners of steps 701-704 have been elaborated in detail above, they will not be elaborated here.

[0146] In addition, when it is determined not to correct the speech text to be verified, the method 700 further includes: generating at least one candidate information according to the associated information of the speech text to be verified; determining the similarity between the pinyin of the speech text to be verified and the pinyin of each candidate information, and selecting the candidate information corresponding to the maximum similarity; modifying the speech text to be verified according to the selected candidate information to perform error correction.

[0147] In addition, when it is determined not to correct the speech text to be verified, the method 700 further includes: generating at least one candidate information according to the associated information of the speech text to be verified; splitting at least one candidate information according to the speech text to be verified to generate at least one split information; determining the similarity between each split information and the speech text to be verified, and determining whether to correct the speech text to be verified according to the similarity.

[0148] Since this has been elaborated above, it will not be elaborated here.

[0149] In addition, for the content not described in detail in this method 700, reference can also be made to the respective steps in the above-mentioned method 200.

[0150] Based on the same technical problems described above, Figure 8 The flowchart of an error correction method for speech recognition provided by another exemplary embodiment of the present application is shown. The method 800 provided by the embodiments of the present application is executed by the above-mentioned intelligent terminal or server, and more specifically, it can be a mobile phone or a server. As Figure 8 shown, the method 800 includes the following steps:

[0151] 801: Obtain speech information and determine the speech text corresponding to the speech information to be error-corrected.

[0152] 802: Generate at least one candidate information according to the associated information of the speech text to be error-corrected.

[0153] 803: Determine the similarity between the pinyin of the speech text to be error-corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity.

[0154] 804: Modify the speech text to be error-corrected according to the selected candidate information to perform error correction.

[0155] Since the specific implementation manners of steps 801-804 have been elaborated in detail above, they will not be repeated here. Only note that the speech text to be error-corrected is the speech text to be verified in the previous text.

[0156] In addition, in the case of determining not to perform error correction on the speech text to be error-corrected, the method 800 further includes: splitting at least one candidate information according to the speech text to be error-corrected to generate at least one split information; determining the similarity between each split information and the speech text to be error-corrected, and determining whether to perform error correction on the speech text to be error-corrected according to the similarity.

[0157] In addition, in the case of determining not to perform error correction on the speech text to be error-corrected, the method 800 further includes: when the length of the speech text to be error-corrected is less than or equal to the length of the associated information, splitting the speech text to be error-corrected according to the length of each associated information to generate at least one split information; determining the similarity between each split information and the corresponding associated information, and determining whether to perform error correction on the speech text to be error-corrected according to the similarity.

[0158] Since this has been elaborated above, it will not be repeated here.

[0159] In addition, for the content not described in detail in this method 800, reference can also be made to the respective steps in the above-mentioned method 200.

[0160] Based on the same technical problems described above, Figure 9The flowchart shows a verification method for speech recognition provided by another exemplary embodiment of the present application. The method 900 provided by the embodiments of the present application is executed by the above-mentioned intelligent terminal or server, and more specifically, it can be a mobile phone or a server. As Figure 9 shown, the method 900 includes the following steps:

[0161] 901: Based on the navigation map, obtain the query speech provided by the user, and determine the query text to be verified corresponding to the query speech.

[0162] 902: Generate at least one candidate name for the destination name according to the destination name in the query text to be verified.

[0163] 903: Split at least one candidate name according to the query text to be verified to generate at least one split name.

[0164] 904: Determine the similarity between each split name and the query text to be verified, and determine whether to correct the query text to be verified according to the similarity.

[0165] Since the specific implementation manners of steps 901 - 904 have been elaborated in detail above, they will not be repeated here. Only note that: the query text to be verified is the query speech text in the previous context. The split name is the split information in the previous context. The candidate name is the candidate information in the previous context. The destination name is the associated information in the previous context.

[0166] In addition, the method 900 further includes: displaying the query text to be verified; in the case of determining to correct the query text to be verified, correcting the query text to be verified according to the corresponding split name, and displaying the corrected query text to be verified.

[0167] In addition, the method 900 further includes: obtaining a plurality of name words representing destinations; determining the similarity between the destination name and each name word, and selecting the name word with a similarity greater than the threshold to correct the destination name; searching for a navigation path according to the corrected destination name, and providing the navigation path.

[0168] In addition, the method 900 further includes: obtaining the semantics therein according to the query text to be verified, and determining the destination name of the query text to be verified.

[0169] Since this has been elaborated above, it will not be repeated here. Only note that: the navigation text can be displayed through the front end or the intelligent terminal.

[0170] In addition, the method 900 further includes: displaying the query text to be verified; when it is determined to correct the query text to be verified, displaying the corresponding split name and prompting the user whether to correct the query text to be verified; receiving a correction instruction, and correcting the query text to be verified according to the split name.

[0171] Based on the foregoing, after the query text to be verified is displayed, a prompt message can also be displayed through the intelligent terminal, and this prompt message is used to provide whether the user corrects the query text to be verified. When the execution entity is an intelligent terminal, such as a mobile phone, the query text to be verified and the prompt message can be directly displayed through the mobile phone. This prompt message can be a pop-up window, which has two options: correct and not correct. The mobile phone can receive the user's correction instruction, such as a correction operation, click on this correction option, and in response to this correction operation, the mobile phone performs correction. The specific correction process is as described above and will not be elaborated here.

[0172] If the execution entity is a server, it can be displayed through the intelligent terminal, such as displaying navigation text and a prompt message, and then performing a correction operation according to the foregoing content.

[0173] If the user does not perform correction, a non-correction instruction can be issued, such as a non-correction operation, and no correction will be performed.

[0174] In addition, for the content not described in detail in the method 900, reference can also be made to each step in the above method 200.

[0175] Based on the same technical problems described above, Figure 10 shows a schematic flowchart of a verification method for speech recognition provided by another exemplary embodiment of the present application. The method 1000 provided by the embodiment of the present application is executed by the above-mentioned server, and more specifically, it can be a server. As Figure 10 shown, the method 1000 includes the following steps:

[0176] 1001: Receive the speech text to be verified.

[0177] 1002: Generate at least one candidate information according to the associated information of the speech text to be verified.

[0178] 1003: Split at least one candidate information according to the speech text to be verified to generate at least one split information.

[0179] 1004: Determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0180] Since the specific implementation manners of steps 1001-1004 have been elaborated in detail above, they will not be elaborated here.

[0181] In addition, for the content not described in detail in this method 1000, reference can also be made to the respective steps in the above-mentioned method 200.

[0182] Figure 11 This is a schematic structural framework diagram of a verification device for speech recognition provided by an exemplary embodiment of the present application. The device 1100 can be applied to a smart terminal or a server, and more specifically, it can be a mobile phone or a server. The device 1100 includes an acquisition module 1101, a generation module 1102, a splitting module 1103, and a determination module 1104; the functions of each module are elaborated in detail below:

[0183] The acquisition module 1101 is configured to acquire speech information and determine the speech text to be verified corresponding to the speech information.

[0184] The generation module 1102 is configured to generate at least one candidate information according to the associated information of the speech text to be verified.

[0185] The splitting module 1103 is configured to split at least one candidate information according to the speech text to be verified to generate at least one split information.

[0186] The determination module 1104 is configured to determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0187] In addition, the device 1100 further includes: a search module, configured to search for a search result corresponding to the speech text to be verified according to the speech text to be verified as the associated information.

[0188] Specifically, the generation module 1102 includes: a generation unit, configured to segment the associated information to generate word segmentation information; and generate at least one candidate information according to the word segmentation information.

[0189] Specifically, the splitting module 1103 is configured to segment each candidate information according to the length of the speech text to be verified to generate at least one segmentation information as the split information.

[0190] Specifically, the determination module 1104 is configured to determine to correct the speech text to be verified when the similarity is greater than the threshold.

[0191] In addition, the determination module 1104 is further configured to determine the split information corresponding to the similarity greater than the threshold, and replace the speech text to be verified with the determined split information for display.

[0192] In addition, when it is determined not to correct the speech text to be verified, the device 1100 further includes: a segmentation module, configured to segment the speech text to be verified according to the length of each associated information to generate at least one segmentation information when the length of the speech text to be verified is less than or equal to the length of the associated information; a determination module 1104, further configured to determine the similarity between each segmentation information and the corresponding associated information, and determine whether to correct the speech text to be verified according to the similarity.

[0193] In addition, the device 1100 further includes: a replacement module, configured to replace the segmentation information in the speech text to be verified with the associated information when the similarity is greater than a threshold value, so as to correct the speech text to be verified.

[0194] In addition, when it is determined not to correct the speech text to be verified, the determination module 1104 is further configured to determine the similarity between the pinyin of the speech text to be verified and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; the device 1100 further includes: a modification module, configured to modify the speech text to be verified according to the selected candidate information so as to perform correction.

[0195] Specifically, the search module includes: an acquisition unit, configured to acquire the semantics therein according to the speech text to be verified and determine the keywords of the speech text to be verified; and acquire search results according to the keywords.

[0196] In addition, the acquisition module 1101 is further configured to acquire a plurality of words associated with the keywords according to the keywords; the determination module is further configured to determine the similarity between the keywords and each word, and select the words with a similarity greater than the threshold value to correct the keywords.

[0197] In addition, the acquisition module 1101 is further configured to acquire search results according to the modified keywords.

[0198] Figure 12 This is a schematic structural framework diagram of a verification device for speech recognition provided by an exemplary embodiment of the present application. The device 1200 can be applied to an intelligent terminal or a server, and more specifically, can be a mobile phone or a server. The device 1200 includes an acquisition module 1201, a determination module 1202, and a generation module 1203; the functions of each module are elaborated in detail below:

[0199] The acquisition module 1201 is configured to acquire speech information and determine the speech text to be verified corresponding to the speech information.

[0200] The determination module 1202 is configured to determine the associated information of the speech text to be verified.

[0201] A generation module 1203, configured to, when the length of the speech text to be verified is less than or equal to the length of the associated information, segment the speech text to be verified according to the length of each piece of associated information, and generate at least one segmentation information.

[0202] A determination module 1202, configured to determine the similarity between each segmentation information and the corresponding associated information, and determine whether to correct the speech text to be verified according to the similarity.

[0203] In addition, when it is determined not to correct the speech text to be verified, the generation module 1203 is further configured to generate at least one candidate information according to the associated information of the speech text to be verified; the determination module 1202 is further configured to determine the similarity between the pinyin of the speech text to be verified and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; the apparatus 1200 further includes: a modification module, configured to modify the speech text to be verified according to the selected candidate information for error correction.

[0204] In addition, when it is determined not to correct the speech text to be verified, the generation module 1203 is further configured to generate at least one candidate information according to the associated information of the speech text to be verified; split at least one candidate information according to the speech text to be verified to generate at least one split information; the determination module 1202 is further configured to determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0205] It should be noted that for the parts not mentioned in the apparatus 1200, reference may be made to the content of the above-mentioned apparatus 1100.

[0206] Figure 13 This is a schematic structural framework diagram of an error correction apparatus for speech recognition provided by an exemplary embodiment of the present application. The apparatus 1300 can be applied to a smart terminal or a server, and more specifically, can be a mobile phone or a server. The apparatus 1300 includes an acquisition module 1301, a generation module 1302, a selection module 1303, and a modification module 1304; the functions of each module are elaborated in detail below:

[0207] An acquisition module 1301, configured to acquire speech information and determine the speech text to be corrected corresponding to the speech information.

[0208] A generation module 1302, configured to generate at least one candidate information according to the associated information of the speech text to be corrected.

[0209] A selection module 1303, configured to determine the similarity between the pinyin of the speech text to be corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity.

[0210] The modification module 1304 is used to modify the speech text to be corrected according to the selected candidate information for error correction.

[0211] In addition, in the case of determining not to correct the speech text to be corrected, the device 1300 further includes: a splitting module, configured to split at least one candidate information according to the speech text to be corrected to generate at least one split information; a determining module, configured to determine the similarity between each split information and the speech text to be corrected, and determine whether to correct the speech text to be corrected according to the similarity.

[0212] In addition, in the case of determining not to correct the speech text to be corrected, the device 1300 further includes: a segmentation module, configured to segment the speech text to be corrected according to the length of each associated information to generate at least one segmented information when the length of the speech text to be corrected is less than or equal to the length of the associated information; the determining module is further configured to determine the similarity between each segmented information and the corresponding associated information, and determine whether to correct the speech text to be corrected according to the similarity.

[0213] It should be noted that for the parts not mentioned in the device 1300, reference may be made to the content of the above device 1100.

[0214] Figure 14 This is a schematic structural framework diagram of a speech recognition verification device provided in an exemplary embodiment of the present application. The device 1400 can be applied to a smart terminal or a server, and more specifically, it can be a mobile phone or a server. The device 1400 includes an acquisition module 1401, a generation module 1402, a splitting module 1403, and a determining module 1404; the functions of each module are elaborated in detail below:

[0215] The acquisition module 1401 is configured to obtain the speech to be verified query provided by the user based on the navigation map, and determine the text to be verified query corresponding to the query speech.

[0216] The generation module 1402 is configured to generate at least one candidate name of the destination name according to the destination name in the text to be verified query.

[0217] The splitting module 1403 is configured to split at least one candidate name according to the text to be verified query to generate at least one split name.

[0218] The determining module 1404 is configured to determine the similarity between each split name and the text to be verified query, and determine whether to correct the text to be verified query according to the similarity.

[0219] In addition, the device 1400 further includes: a display module for displaying the query text to be verified; the determination module 1404 is further configured to, when determining that the query text to be verified needs to be corrected, correct the query text to be verified according to the corresponding split name and display the corrected query text to be verified.

[0220] In addition, the acquisition module 1401 is further configured to acquire a plurality of name words representing destinations; the determination module 1404 is further configured to determine the similarity between the destination name and each name word, and select the name words with similarity greater than the threshold to correct the destination name; the device 1400 further includes: a provision module for searching for a navigation path according to the corrected destination name and providing the navigation path.

[0221] In addition, the determination module 1404 is further configured to: obtain the semantics therein according to the query text to be verified and determine the destination name of the query text to be verified.

[0222] In addition, the display module is further configured to: display the query text to be verified; the device 1400 further includes: a prompt module for, when determining that the query text to be verified needs to be corrected, displaying the corresponding split name and prompting the user whether to correct the query text to be verified; a correction module for receiving a correction instruction and correcting the query text to be verified according to the split name.

[0223] It should be noted that for the parts not mentioned in the device 1400, reference may be made to the content of the device 1100 above.

[0224] Figure 15 This is a schematic structural framework diagram of a verification device for speech recognition provided by an exemplary embodiment of the present application. The device 1500 can be applied to a server side, and more specifically, it can be a server. The device 1500 includes a receiving module 1501, a generating module 1502, a splitting module 1503, and a determining module 1504; the functions of each module are elaborated in detail below:

[0225] The receiving module 1501 is configured to receive the speech text to be verified.

[0226] The generating module 1502 is configured to generate at least one candidate information according to the associated information of the speech text to be verified.

[0227] The splitting module 1503 is configured to split at least one candidate information according to the speech text to be verified to generate at least one split information.

[0228] The determining module 1504 is configured to determine the similarity between each split information and the speech text to be verified and determine whether to correct the speech text to be verified according to the similarity.

[0229] It should be noted that for the parts not mentioned in the device 1500, the content of the above-mentioned device 1100 can be referred to.

[0230] The above describes Figure 11 the internal functions and structures of the device 1100 shown. In a possible design, Figure 11 the structure of the device 1100 shown can be implemented as a computing device, such as a server or a mobile phone. As Figure 16 shown, the device 1600 may include: a memory 1601 and a processor 1602;

[0231] The memory 1601 is used to store computer programs.

[0232] The processor 1602 is used to execute the computer program to: obtain voice information, determine the voice text to be verified corresponding to the voice information; generate at least one candidate information according to the associated information of the voice text to be verified; split at least one candidate information according to the voice text to be verified to generate at least one split information; determine the similarity between each split information and the voice text to be verified, and determine whether to correct the voice text to be verified according to the similarity.

[0233] In addition, the processor 1602 is further used to: search for the search result corresponding to the voice text to be verified according to the voice text to be verified as the associated information.

[0234] Specifically, the processor 1602 is specifically used to: split the associated information to generate word segmentation information; generate at least one candidate information according to the word segmentation information.

[0235] Specifically, the processor 1602 is specifically used to: split each candidate information according to the length of the voice text to be verified to generate at least one split information as the split information.

[0236] Specifically, the processor 1602 is specifically used to: when the similarity is greater than the threshold, determine to correct the voice text to be verified.

[0237] In addition, the processor 1602 is further used to: determine the split information corresponding to the similarity greater than the threshold, and replace the voice text to be verified with the determined split information for display.

[0238] In addition, in the case of determining not to correct the voice text to be verified, the processor 1602 is further used to: when the length of the voice text to be verified is less than or equal to the length of the associated information, split the voice text to be verified according to the length of each associated information to generate at least one split information; determine the similarity between each split information and the corresponding associated information, and determine whether to correct the voice text to be verified according to the similarity.

[0239] In addition, the processor 1602 is further configured to: when the similarity is greater than the threshold, replace the segmentation information in the speech text to be verified with the associated information to correct the speech text to be verified.

[0240] In addition, when it is determined not to correct the speech text to be verified, the processor 1602 is further configured to: determine the similarity between the pinyin of the speech text to be verified and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; modify the speech text to be verified according to the selected candidate information to perform error correction.

[0241] Specifically, the processor 1602 is specifically configured to: obtain the semantics therein according to the speech text to be verified, and determine the keywords of the speech text to be verified; obtain search results according to the keywords.

[0242] In addition, the processor 1602 is further configured to: obtain multiple words associated with the keyword according to the keyword; the determination module is further configured to determine the similarity between the keyword and each word, and select the words with similarity greater than the threshold to correct the keyword.

[0243] In addition, the processor 1602 is further configured to: obtain search results according to the modified keyword.

[0244] In addition, an embodiment of the present invention provides a computer storage medium, when a computer program is executed by one or more processors, causing the one or more processors to implement Figure 2 the steps of a voice recognition verification method in the method embodiment.

[0245] The above describes Figure 12 the internal functions and structures of the device 1200 shown. In a possible design, Figure 12 the structure of the device 1200 shown can be implemented as a computing device, such as a server or a mobile phone. As Figure 17 shown, the device 1700 may include: a memory 1701 and a processor 1702;

[0246] The memory 1701 is used to store computer programs.

[0247] The processor 1702 is configured to execute a computer program to: obtain voice information, determine the speech text to be verified corresponding to the voice information; determine the associated information of the speech text to be verified; when the length of the speech text to be verified is less than or equal to the length of the associated information, segment the speech text to be verified according to the length of each associated information to generate at least one segmentation information; determine the similarity between each segmentation information and the corresponding associated information, and determine whether to correct the speech text to be verified according to the similarity.

[0248] In addition, when it is determined not to correct the speech text to be verified, the processor 1702 is further configured to: generate at least one candidate information according to the associated information of the speech text to be verified; determine the similarity between the pinyin of the speech text to be verified and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; modify the speech text to be verified according to the selected candidate information for error correction.

[0249] In addition, when it is determined not to correct the speech text to be verified, the processor 1702 is further configured to: generate at least one candidate information according to the associated information of the speech text to be verified; split at least one candidate information according to the speech text to be verified to generate at least one split information; determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0250] It should be noted that for the parts not mentioned in the device 1700, reference can be made to the content of the above-mentioned device 1600.

[0251] In addition, an embodiment of the present invention provides a computer storage medium, when a computer program is executed by one or more processors, it causes the one or more processors to implement Figure 7 the steps of a verification method for speech recognition in the method embodiment.

[0252] The above describes Figure 13 the internal functions and structures of the device 1300 shown. In a possible design, Figure 13 the structure of the device 1300 shown can be implemented as a computing device, such as a server or a mobile phone. As Figure 18 shown, the device 1800 may include: a memory 1801 and a processor 1802;

[0253] The memory 1801 is used to store a computer program.

[0254] The processor 1802 is used to execute the computer program to: obtain speech information, and determine the speech text to be error-corrected corresponding to the speech information; generate at least one candidate information according to the associated information of the speech text to be error-corrected; determine the similarity between the pinyin of the speech text to be error-corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; modify the speech text to be error-corrected according to the selected candidate information for error correction.

[0255] In addition, when it is determined not to correct the speech text to be corrected, the processor 1802 is further configured to: split at least one candidate information according to the speech text to be corrected to generate at least one split information; determine the similarity between each split information and the speech text to be corrected, and determine whether to correct the speech text to be corrected according to the similarity.

[0256] In addition, when it is determined not to correct the speech text to be corrected, the processor 1802 is further configured to: when the length of the speech text to be corrected is less than or equal to the length of the associated information, segment the speech text to be corrected according to the length of each associated information to generate at least one segmented information; determine the similarity between each segmented information and the corresponding associated information, and determine whether to correct the speech text to be corrected according to the similarity.

[0257] It should be noted that for the parts not mentioned in the device 1800, reference may be made to the content of the above-mentioned device 1600.

[0258] In addition, an embodiment of the present invention provides a computer storage medium. When a computer program is executed by one or more processors, one or more processors are caused to implement Figure 8 the steps of an error correction method for speech recognition in the method embodiment.

[0259] The above describes Figure 14 the internal functions and structures of the device 1400 shown. In a possible design, Figure 14 the structure of the device 1400 shown can be implemented as a computing device, such as a server or a mobile phone. As Figure 19 shown, the device 1900 may include: a memory 1901 and a processor 1902;

[0260] The memory 1901 is used to store a computer program.

[0261] The processor 1902 is configured to execute the computer program to: based on a navigation map, obtain a query speech provided by a user, and determine a query text to be verified corresponding to the query speech; generate at least one candidate name for the destination name according to the destination name in the query text to be verified; split at least one candidate name according to the query text to be verified to generate at least one split name; determine the similarity between each split name and the navigation text, and determine whether to correct the query text to be verified according to the similarity.

[0262] In addition, the processor 1902 is further configured to: display the query text to be verified; when it is determined to correct the navigation text, correct the query text to be verified according to the corresponding split name and display the corrected query text to be verified.

[0263] In addition, the processor 1902 is further configured to: obtain multiple name words representing destinations; determine the similarity between the destination name and each name word, and select the name words with a similarity greater than the threshold to correct the destination name; search for a navigation path based on the corrected destination name, and provide the navigation path.

[0264] In addition, the processor 1902 is further configured to: obtain the semantics therein according to the query text to be verified, and determine the destination name of the query text to be verified.

[0265] In addition, the processor 1902 is further configured to: display the query text to be verified; when it is determined that the query text to be verified is to be corrected, display the corresponding split name, and prompt the user whether to correct the query text to be verified; receive a correction instruction, and correct the query text to be verified according to the split name.

[0266] It should be noted that for the parts not mentioned in the device 1900, reference may be made to the content of the above device 1600.

[0267] In addition, an embodiment of the present invention provides a computer storage medium. When a computer program is executed by one or more processors, one or more processors are caused to implement Figure 9 the steps of a verification method for speech recognition in the method embodiment.

[0268] The above describes Figure 15 the internal functions and structures of the device 1500 shown. In a possible design, Figure 15 the structure of the device 1500 shown can be implemented as a computing device, such as a server or a mobile phone. As Figure 20 shown, the device 2000 may include: a memory 2001, a processor 2002, and a communication component 2003;

[0269] The memory 2001 is used to store a computer program.

[0270] The processor 2002 is configured to execute the computer program to: generate at least one candidate information according to the associated information of the speech text to be verified; split at least one candidate information according to the speech text to be verified to generate at least one split information; determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

[0271] The communication component 2003 is used to receive the speech text to be verified.

[0272] It should be noted that for the parts not mentioned in the device 2000, reference may be made to the content of the above device 1600.

[0273] In addition, an embodiment of the present invention provides a computer storage medium. When a computer program is executed by one or more processors, one or more processors are caused to implement Figure 10 the steps of a voice recognition verification method in the method embodiment.

[0274] Based on the similar technical problems described above, a verification method for order-taking voice recognition provided by another exemplary embodiment of the present application. The method 2100 provided by the embodiment of the present application is executed by the above-mentioned intelligent terminal or server, and more specifically, it can be a mobile phone or a server. The method 2100 includes the following steps:

[0275] 2101: Obtain the order-taking voice and determine the order-taking text to be verified corresponding to the order-taking voice.

[0276] 2102: Generate at least one candidate name for the food item name according to the food item name in the order-taking text to be verified.

[0277] 2103: Split at least one candidate name according to the order-taking text to be verified to generate at least one split name.

[0278] 2104: Determine the similarity between each split name and the order-taking text to be verified, and determine whether to correct the order-taking text to be verified according to the similarity.

[0279] Since the specific implementation manners of steps 2101-2104 have been elaborated in detail above, they will not be repeated here. It is only noted that for order-taking, it can be done on an external APP (application), or in a physical restaurant, and order-taking can be done through a tablet computer. If the food item name is corrected, online ordering can be performed based on the corrected food item name, and then the merchant can receive the order and deliver the food for takeout, or serve the food in a physical restaurant.

[0280] In addition, for the content not described in detail in this method 2100, reference can also be made to each step in the above method 200.

[0281] An order-taking voice recognition verification device provided by an exemplary embodiment of the present application. The device 2200 can be applied to an intelligent terminal or a server, and more specifically, it can be a mobile phone or a server. The device 2200 includes an acquisition module 2201, a generation module 2202, a splitting module 2203, and a determination module 2204; the functions of each module are elaborated in detail below:

[0282] The acquisition module 2201 is used to obtain the order-taking voice and determine the order-taking text to be verified corresponding to the order-taking voice.

[0283] A generation module 2202, configured to generate at least one candidate name for a food item name according to the food item name in the to-be-verified ordering text.

[0284] A splitting module 2203, configured to split at least one candidate name according to the to-be-verified ordering text to generate at least one split name.

[0285] A determination module 2204, configured to determine the similarity between each split name and the to-be-verified ordering text, and determine whether to correct the to-be-verified ordering text according to the similarity.

[0286] It should be noted that for the parts not mentioned in the apparatus 2200, reference can be made to the content of the above-mentioned apparatus 1100.

[0287] The internal functions and structures of the apparatus 2200 are described above. In a possible design, the structure of the apparatus 2200 can be implemented as a computing device, such as a server or a mobile phone. The device 2300 may include: a memory 2301 and a processor 2302;

[0288] A memory 2001, configured to store computer programs.

[0289] A processor 2002, configured to execute the computer program to: obtain an ordering voice, determine the to-be-verified ordering text corresponding to the ordering voice; generate at least one candidate name for a food item name according to the food item name in the to-be-verified ordering text; split at least one candidate name according to the to-be-verified ordering text to generate at least one split name; determine the similarity between each split name and the to-be-verified ordering text, and determine whether to correct the to-be-verified ordering text according to the similarity.

[0290] It should be noted that for the parts not mentioned in the device 2300, reference can be made to the content of the above-mentioned device 1600.

[0291] In addition, an embodiment of the present invention provides a computer storage medium. When the computer program is executed by one or more processors, it causes the one or more processors to implement the steps of a verification method for ordering voice recognition in an embodiment of the method 2100.

[0292] Based on the similar technical problems described above, a verification method for ticket-purchasing voice recognition provided in another exemplary embodiment of the present application. The method 2400 provided in the embodiment of the present application is executed by the above-mentioned intelligent terminal or server, and more specifically, it can be a mobile phone or a server. The method 2400 includes the following steps:

[0293] 2401: Obtain a ticket-purchasing voice, and determine the to-be-verified ticket-purchasing text corresponding to the ticket-purchasing voice.

[0294] 2402: Generate at least one candidate name for the ticket purchase name according to the ticket purchase name in the ticket purchase text to be verified.

[0295] 2403: Split at least one candidate name according to the ticket purchase text to be verified to generate at least one split name.

[0296] 2404: Determine the similarity between each split name and the ticket purchase text to be verified, and determine whether to correct the ticket purchase text to be verified according to the similarity.

[0297] Since the specific implementation manners of steps 2401 - 2404 have been elaborated in detail above, they will not be repeated here. Only note that ticket purchase can be for movie tickets, sports events, concerts, bus tickets, train tickets, etc. Users can purchase tickets online through a smart terminal, such as a ticket purchase APP installed on a mobile phone. When a user inputs a ticket purchase voice in a voice manner. And when the ticket purchase name in the ticket purchase text to be verified, such as xx movie ticket, needs to be corrected, the corrected movie ticket is displayed, then the user can purchase the ticket online for the corrected movie ticket, and thus tickets are issued to the user.

[0298] A verification device for ticket purchase voice recognition provided by an exemplary embodiment of the present application. The device 2500 can be applied to a smart terminal or a server, and more specifically, can be a mobile phone or a server. The device 2500 includes an acquisition module 2501, a generation module 2502, a splitting module 2503, and a determination module 2504; the functions of each module are elaborated in detail below:

[0299] The acquisition module 2501 is used to: acquire the ticket purchase voice and determine the ticket purchase text to be verified corresponding to the ticket purchase voice.

[0300] The generation module 2502 is used to: generate at least one candidate name for the ticket purchase name according to the ticket purchase name in the ticket purchase text to be verified.

[0301] The splitting module 2503 is used to: split at least one candidate name according to the ticket purchase text to be verified to generate at least one split name.

[0302] The determination module 2504 is used to: determine the similarity between each split name and the ticket purchase text to be verified, and determine whether to correct the ticket purchase text to be verified according to the similarity.

[0303] It should be noted that for the parts not mentioned in the device 2500, the content of the above device 1100 can be referred to.

[0304] The internal functions and structures of the device 2500 have been described above. In a possible design, the structure of the device 2500 can be implemented as a computing device, such as a server or a mobile phone. The device 2600 may include: a memory 2601 and a processor 2602;

[0305] The memory 2601 is used to store computer programs.

[0306] The processor 2602 is used to execute the computer program to: obtain a ticket-purchasing voice, and determine the to-be-verified ticket-purchasing text corresponding to the ticket-purchasing voice; generate at least one candidate name for the ticket-purchasing name according to the ticket-purchasing name in the to-be-verified ticket-purchasing text; split at least one candidate name according to the to-be-verified ticket-purchasing text to generate at least one split name; determine the similarity between each split name and the to-be-verified ticket-purchasing text, and determine whether to correct the to-be-verified ticket-purchasing text according to the similarity.

[0307] It should be noted that for the parts not mentioned in the device 2600, reference can be made to the content of the device 1600 above.

[0308] In addition, an embodiment of the present invention provides a computer storage medium. When the computer program is executed by one or more processors, one or more processors are caused to implement the steps of a verification method for ticket-purchasing voice recognition in the embodiment of the method 2400.

[0309] Based on the similar technical problems described above, an embodiment of the present application provides a verification method for on-demand voice recognition. The method 2700 provided by the embodiment of the present application is executed by the above intelligent terminal or server, and more specifically, it can be a mobile phone or a server. The method 2700 includes the following steps:

[0310] 2701: Obtain an on-demand voice, and determine the to-be-verified on-demand text corresponding to the on-demand voice.

[0311] 2702: Generate at least one candidate name for the video name according to the video name in the to-be-verified on-demand text.

[0312] 2703: Split at least one candidate name according to the to-be-verified on-demand text to generate at least one split name.

[0313] 2704: Determine the similarity between each split name and the to-be-verified on-demand text, and determine whether to correct the to-be-verified on-demand text according to the similarity.

[0314] Since the specific implementation manners of steps 2701 - 2704 have been elaborated in detail above, they will not be repeated here. Only to note that video on demand can be performed through a smart terminal, such as a video - on - demand APP installed on a mobile phone. This video on demand can play video - on - demand content of TV programs or websites (which can be displayed on a computer), as well as the switching of video or TV content. Users can input the on - demand voice in a voice manner, such as xx movie, xx program, etc. They can also switch channels, such as switching to xx TV station, etc. If the video name or TV station name in the to - be - verified on - demand text needs to be corrected, then based on the corrected video name or TV station name, the corresponding video content or TV content, etc. will be directly played for the user to watch.

[0315] A verification device for on - demand voice recognition provided by an exemplary embodiment of the present application. The device 2800 can be applied to a smart terminal or a server, more specifically, a mobile phone or a server. The device 2800 includes an acquisition module 2801, a generation module 2802, a splitting module 2803, and a determination module 2804; the functions of each module will be elaborated in detail below:

[0316] The acquisition module 2801 is configured to acquire an on - demand voice and determine the to - be - verified on - demand text corresponding to the on - demand voice.

[0317] The generation module 2802 is configured to generate at least one candidate name of a video name according to the video name in the to - be - verified on - demand text.

[0318] The splitting module 2803 is configured to split at least one candidate name according to the to - be - verified on - demand text to generate at least one split name.

[0319] The determination module 2804 is configured to determine the similarity between each split name and the to - be - verified on - demand text, and determine whether to correct the to - be - verified on - demand text according to the similarity.

[0320] It should be noted that for parts not mentioned in the device 2800, the content of the above - mentioned device 1100 can be referred to.

[0321] The internal functions and structures of the device 2800 have been described above. In a possible design, the structure of the device 2800 can be implemented as a computing device, such as a server or a mobile phone. The device 2900 may include: a memory 2901, a processor 2902;

[0322] The memory 2901 is used to store computer programs.

[0323] A processor 2902 for executing a computer program for: obtaining on-demand voice, determining on-demand text to be verified corresponding to the on-demand voice; generating at least one candidate name for the video name according to the video name in the on-demand text to be verified; splitting at least one candidate name according to the on-demand text to be verified to generate at least one split name; determining the similarity between each split name and the on-demand text to be verified, and determining whether to correct the on-demand text to be verified according to the similarity.

[0324] It should be noted that for the parts not mentioned in the device 2900, reference may be made to the content of the above device 1600.

[0325] In addition, an embodiment of the present invention provides a computer storage medium, when a computer program is executed by one or more processors, causing the one or more processors to implement the steps of a verification method for on-demand voice recognition in the method 2700 embodiment.

[0326] Based on the similar technical problems described above, another exemplary embodiment of the present application provides an error correction method for speech recognition. The method 3000 provided by the embodiment of the present application is executed by the above intelligent terminal or server, and more specifically, it can be a mobile phone or a server. The method 3000 includes the following steps:

[0327] 3001: Obtain voice information, and determine the voice text to be corrected corresponding to the voice information.

[0328] 3002: Generate at least one candidate information according to the associated information of the voice text to be corrected.

[0329] 3003: Split at least one candidate information according to the voice text to be corrected to generate at least one split information.

[0330] 3004: Determine the similarity between each split information and the voice text to be corrected. When the similarity is greater than the threshold, it is determined to correct the voice text to be corrected.

[0331] Since the specific implementation manners of steps 3001 - 3004 have been elaborated in detail above, they will not be repeated here.

[0332] In addition, in the case of determining not to correct the voice text to be corrected, the method 3000 further includes: when the length of the voice text to be corrected is less than or equal to the length of the associated information, splitting the voice text to be corrected according to the length of each associated information to generate at least one split information; determining the similarity between each split information and the corresponding associated information, and determining whether to correct the voice text to be corrected according to the similarity.

[0333] In addition, in the case of determining not to correct the speech text to be corrected, the method 3000 further includes: determining the similarity between the pinyin of the speech text to be corrected and the pinyin of each candidate information, and selecting the candidate information corresponding to the maximum similarity; modifying the speech text to be corrected according to the selected candidate information for error correction.

[0334] Since it has been described above, it will not be elaborated here.

[0335] An error correction device for speech recognition provided by an exemplary embodiment of the present application. The device 3100 can be applied to a smart terminal or a server, and more specifically, it can be a mobile phone or a server. The device 3100 includes an acquisition module 3101, a generation module 3102, a splitting module 3103, and a determination module 3104; the functions of each module are elaborated in detail below:

[0336] The acquisition module 3101 is configured to acquire speech information and determine the speech text to be corrected corresponding to the speech information.

[0337] The generation module 3102 is configured to generate at least one candidate information according to the associated information of the speech text to be corrected.

[0338] The splitting module 3103 is configured to split at least one candidate information according to the speech text to be corrected to generate at least one split information.

[0339] The determination module 3104 is configured to determine the similarity between each split information and the speech text to be corrected. When the similarity is greater than the threshold, it is determined to correct the speech text to be corrected.

[0340] In addition, in the case of determining not to correct the speech text to be corrected, the device 3100 further includes: a segmentation module, configured to segment the speech text to be corrected according to the length of each associated information to generate at least one segmentation information when the length of the speech text to be corrected is less than or equal to the length of the associated information; the determination module 3104 is further configured to determine the similarity between each segmentation information and the corresponding associated information, and determine whether to correct the speech text to be corrected according to the similarity.

[0341] In addition, in the case of determining not to correct the speech text to be corrected, the device 3100 further includes: a selection module, configured to determine the similarity between the pinyin of the speech text to be corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; a modification module, configured to modify the speech text to be corrected according to the selected candidate information for error correction.

[0342] The internal functions and structures of apparatus 3100 have been described above. In one possible design, the structure of apparatus 3200 may be implemented as a computing device, such as a server or a mobile phone. The apparatus 3200 may include: a memory 3201 and a processor 3202;

[0343] The memory 3201 is used to store computer programs.

[0344] The processor 3202 is used to execute the computer program to: obtain voice information, and determine the voice text to be corrected corresponding to the voice information; generate at least one candidate information according to the associated information of the voice text to be corrected; split at least one candidate information according to the voice text to be corrected to generate at least one split information; determine the similarity between each split information and the voice text to be corrected, and when the similarity is greater than the threshold, determine to correct the voice text to be corrected.

[0345] In addition, in the case of determining not to correct the voice text to be corrected, the processor 3202 is further used to: when the length of the voice text to be corrected is less than or equal to the length of the associated information, segment the voice text to be corrected according to the length of each associated information to generate at least one segmented information; the determining module 3104 is further used to determine the similarity between each segmented information and the corresponding associated information, and determine whether to correct the voice text to be corrected according to the similarity.

[0346] In addition, in the case of determining not to correct the voice text to be corrected, the processor 3202 is further used to: determine the similarity between the pinyin of the voice text to be corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; modify the voice text to be corrected according to the selected candidate information to perform error correction.

[0347] It should be noted that for the parts not mentioned in the apparatus 3200, reference may be made to the content of the above apparatus 1600.

[0348] In addition, an embodiment of the present invention provides a computer storage medium. When the computer program is executed by one or more processors, one or more processors are caused to implement the steps of an error correction method for speech recognition in an embodiment of method 3000.

[0349] In addition, in some of the processes described in the above embodiments and the accompanying drawings, which include a plurality of operations that appear in a specific order, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 201, 202, 203, etc., are only used to distinguish the different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0350] The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0351] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0352] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable multimedia data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable multimedia data processing device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0353] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable multimedia data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the acts Figure 1 or acts and / or boxes Figure 1 or boxes specified in one or more of the boxes.

[0354] These computer program instructions can also be loaded onto a computer or other programmable multimedia data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the acts Figure 1 or acts and / or boxes Figure 1 or boxes specified in one or more of the boxes.

[0355] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0356] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0357] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0358] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A verification method for speech recognition, characterized in that, Including: Obtain voice information and determine the voice text to be verified corresponding to the voice information; Segment the associated information of the voice text to be verified to obtain at least one segmented information, and recombine the at least one segmented information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user truly wants to express; According to the voice text to be verified, split the at least one candidate information to generate at least one split information; Determine the similarity between each split information and the voice text to be verified, and determine whether to correct the voice text to be verified based on the similarity.

2. The method according to claim 1, wherein The method further includes: According to the voice text to be verified, search for the search result corresponding to the voice text to be verified as the associated information.

3. The method according to claim 1, characterized in that The splitting the at least one candidate information according to the voice text to be verified to generate at least one split information includes: According to the length of the voice text to be verified, segment each candidate information to generate at least one segmented information as the split information.

4. The method according to claim 1, characterized in that The determining whether to correct the voice text to be verified according to the similarity includes: When the similarity is greater than the threshold, it is determined to correct the voice text to be verified.

5. The method according to claim 4, wherein The method further includes: Determine the split information corresponding to the similarity greater than the threshold, and replace the voice text to be verified with the determined split information for display.

6. The method according to claim 1, wherein In the case where it is determined not to correct the voice text to be verified, the method further includes: In the case where the length of the voice text to be verified is less than or equal to the length of the associated information, segment the voice text to be verified according to the length of each associated information to generate at least one segmented information; Determine the similarity between each segmented information and the corresponding associated information, and determine whether to correct the voice text to be verified based on the similarity.

7. The method according to claim 6, characterized in that, The method further includes: When the similarity is greater than the threshold, replace the segmented information in the voice text to be verified with the associated information to correct the voice text to be verified.

8. The method according to claim 1, wherein In the case where it is determined not to correct the voice text to be verified, the method further includes: Determine the similarity between the pinyin of the voice text to be verified and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; Modify the voice text to be verified according to the selected candidate information for error correction.

9. The method according to claim 2, characterized in that, The searching for the search result corresponding to the voice text to be verified according to the voice text to be verified includes: According to the voice text to be verified, obtain the semantics therein and determine the keywords of the voice text to be verified; According to the keywords, obtain the search result.

10. The method according to claim 9, wherein The method further includes: According to the keywords, obtain multiple words associated with the keywords; Determine the similarity between the keywords and each word, and select the words with similarity greater than the threshold to correct the keywords.

11. The method according to claim 10, wherein The method further includes: According to the modified keywords, obtain the search result.

12. A verification method for speech recognition, characterized in that, Including: Obtain voice information and determine the voice text to be verified corresponding to the voice information; Determine the associated information of the voice text to be verified; When the length of the speech text to be verified is less than or equal to the length of the associated information, segment the speech text to be verified according to the length of each piece of associated information to generate at least one segmentation information; Determine the similarity between each piece of segmentation information and the corresponding associated information, and determine whether to correct the speech text to be verified according to the similarity; When it is determined not to correct the speech text to be verified, segment the associated information of the speech text to be verified to obtain at least one word segmentation information, and recombine the at least one word segmentation information to generate at least one candidate information, where the associated information represents the semantic information of the speech text that the user actually wants to express; Split the at least one candidate information according to the speech text to be verified to generate at least one split information; Determine the similarity between each piece of split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

13. The method according to claim 12, wherein When it is determined not to correct the speech text to be verified, the method further includes: Generate at least one candidate information according to the associated information of the speech text to be verified; Determine the similarity between the pinyin of the speech text to be verified and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; Modify the speech text to be verified according to the selected candidate information to perform error correction.

14. An error correction method for speech recognition, characterized in that, The method further includes: Obtain speech information, and determine the speech text to be error-corrected corresponding to the speech information; Segment the associated information of the speech text to be error-corrected to obtain at least one word segmentation information, and recombine the at least one word segmentation information to generate at least one candidate information, where the associated information represents the semantic information of the speech text that the user actually wants to express; Determine the similarity between the pinyin of the speech text to be error-corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; Modify the speech text to be error-corrected according to the selected candidate information to perform error correction; When it is determined not to correct the speech text to be error-corrected, split the at least one candidate information according to the speech text to be error-corrected to generate at least one split information; determine the similarity between each piece of split information and the speech text to be error-corrected, and determine whether to correct the speech text to be error-corrected according to the similarity.

15. The method according to claim 14, wherein When it is determined not to correct the speech text to be error-corrected, the method further includes: When the length of the speech text to be error-corrected is less than or equal to the length of the associated information, segment the speech text to be error-corrected according to the length of each piece of associated information to generate at least one segmentation information; Determine the similarity between each piece of segmentation information and the corresponding associated information, and determine whether to correct the speech text to be error-corrected according to the similarity.

16. A verification method for speech recognition, characterized in that, Includes: Based on a navigation map, obtain a query speech provided by a user, and determine the speech text to be verified corresponding to the query speech; Recombine the destination names in the speech text to be verified to generate at least one candidate name for the destination name; Split the at least one candidate name according to the query text to be verified, generating at least one split name; Determine the similarity between each split name and the query text to be verified, and determine whether to correct the query text to be verified according to the similarity.

17. The method according to claim 16, wherein The method further includes: Display the query text to be verified; When it is determined to correct the query text to be verified, correct the query text to be verified according to the corresponding split name, and display the corrected query text to be verified.

18. The method according to claim 16, wherein The method further includes: Obtain multiple name words representing destinations; Determine the similarity between the destination name and each name word, and select the name words with similarity greater than the threshold to correct the destination name; Search for a navigation path according to the corrected destination name and provide the navigation path.

19. The method according to claim 16, wherein The method further includes: According to the query text to be verified, obtain the semantics therein and determine the destination name of the query text to be verified.

20. The method according to claim 16, characterized in that The method further includes: Display the query text to be verified; When it is determined to correct the query text to be verified, display the corresponding split name and prompt the user whether to correct the query text to be verified; Receive a correction instruction and correct the query text to be verified according to the split name.

21. A verification method for speech recognition, characterized in that, Includes: Receive the speech text to be verified; Segment the associated information of the speech text to be verified to obtain at least one word segmentation information, and recombine the at least one word segmentation information to generate at least one candidate information, where the associated information represents the semantic information of the speech text that the user truly wants to express; Split the at least one candidate information according to the speech text to be verified, generating at least one split information; Determine the similarity between each split information and the speech text to be verified, and determine whether to correct the speech text to be verified according to the similarity.

22. A verification method for voice recognition of ordering food, characterized in that, Includes: Obtain an ordering speech, and determine the ordering text to be verified corresponding to the ordering speech; Recombine the food item names in the ordering text to be verified to generate at least one candidate name for the food item names; Split the at least one candidate name according to the ordering text to be verified, generating at least one split name; Determine the similarity between each split name and the ordering text to be verified, and determine whether to correct the ordering text to be verified according to the similarity.

23. A verification method for ticket purchase voice recognition, characterized in that, Includes: Obtain a ticket-purchasing speech, and determine the ticket-purchasing text to be verified corresponding to the ticket-purchasing speech; Recombine the ticket-purchasing names in the ticket-purchasing text to be verified to generate at least one candidate name for the ticket-purchasing names; Split the at least one candidate name according to the ticket-purchasing text to be verified, generating at least one split name; Determine the similarity between each split name and the ticket-purchasing text to be verified, and determine whether to correct the ticket-purchasing text to be verified according to the similarity.

24. A verification method for on-demand speech recognition, characterized in that, Includes: Obtain an on-demand speech, and determine the on-demand text to be verified corresponding to the on-demand speech; Recombine the video names in the on-demand text to be verified to generate at least one candidate name for the video names; Split the at least one candidate name according to the on-demand text to be verified, generating at least one split name; Determine the similarity between each split name and the on-demand text to be verified, and determine whether to correct the on-demand text to be verified according to the similarity.

25. An error correction method for speech recognition, characterized in that, Including: Obtain voice information and determine the voice text to be corrected corresponding to the voice information; Segment the associated information of the voice text to be corrected to obtain at least one word segmentation information, and recombine the at least one word segmentation information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user actually wants to express; Split the at least one candidate information according to the voice text to be corrected, generating at least one split information; Determine the similarity between each split information and the voice text to be corrected. When the similarity is greater than the threshold, determine to correct the voice text to be corrected.

26. The method according to claim 25, wherein In the case of determining not to correct the voice text to be corrected, the method further includes: In the case where the length of the voice text to be corrected is less than or equal to the length of the associated information, segment the voice text to be corrected according to the length of each associated information, generating at least one segmented information; Determine the similarity between each segmented information and the corresponding associated information, and determine whether to correct the voice text to be corrected according to the similarity.

27. The method according to claim 26, wherein In the case of determining not to correct the voice text to be corrected, the method further includes: Determine the similarity between the pinyin of the voice text to be corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; Modify the voice text to be corrected according to the selected candidate information to perform error correction.

28. A computing device, comprising: A memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program for: Obtain voice information and determine the voice text to be verified corresponding to the voice information; Segment the associated information of the voice text to be verified to obtain at least one word segmentation information, and recombine the at least one word segmentation information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user actually wants to express; Split the at least one candidate information according to the voice text to be verified, generating at least one split information; Determine the similarity between each split information and the voice text to be verified, and determine whether to correct the voice text to be verified according to the similarity.

29. A computing device, comprising: A memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program for: Obtain voice information and determine the voice text to be verified corresponding to the voice information; Determine the associated information of the voice text to be verified; In the case where the length of the voice text to be verified is less than or equal to the length of the associated information, segment the voice text to be verified according to the length of each associated information, generating at least one segmented information; Determine the similarity between each segmented information and the corresponding associated information, and determine whether to correct the voice text to be verified according to the similarity; In the case of determining not to correct the voice text to be verified, segment the associated information of the voice text to be verified to obtain at least one segmented information, and recombine the at least one segmented information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user actually wants to express; According to the voice text to be verified, split the at least one candidate information to generate at least one split information; Determine the similarity between each split information and the voice text to be verified, and determine whether to correct the voice text to be verified according to the similarity.

30. A computing device, comprising: A memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program for: Obtain voice information and determine the voice text to be corrected corresponding to the voice information; Segment the associated information of the voice text to be corrected to obtain at least one segmented information, and recombine the at least one segmented information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user actually wants to express; Determine the similarity between the pinyin of the voice text to be corrected and the pinyin of each candidate information, and select the candidate information corresponding to the maximum similarity; Modify the voice text to be corrected according to the selected candidate information to perform error correction.

31. A computing device, comprising: A memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program for: Based on a navigation map, obtain a query voice provided by a user and determine the query text to be verified corresponding to the query voice; Recombine the destination names in the query text to be verified to generate at least one candidate name for the destination name; According to the query text to be verified, split the at least one candidate name to generate at least one split name; Determine the similarity between each split name and the query text to be verified, and determine whether to correct the query text to be verified according to the similarity.

32. A computing device, comprising: A memory, a processor, and a communication component; The memory is used to store a computer program; The processor is used to execute the computer program for: Receive the voice text to be verified; Segment the associated information of the voice text to be verified to obtain at least one segmented information, and recombine the at least one segmented information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user actually wants to express; According to the voice text to be verified, split the at least one candidate information to generate at least one split information; Determine the similarity between each split information and the voice text to be verified, and determine whether to correct the voice text to be verified according to the similarity; The communication component is used to receive the voice text to be verified.

33. A computing device, comprising: A memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program for: Obtain an ordering voice and determine the ordering text to be verified corresponding to the ordering voice; Recombine the food item names in the to-be-verified order-taking text to generate at least one candidate name for the food item name; Split the at least one candidate name according to the to-be-verified order-taking text to generate at least one split name; Determine the similarity between each split name and the to-be-verified order-taking text, and determine whether to correct the to-be-verified order-taking text according to the similarity.

34. A computing device, comprising: A memory and a processor; The memory is used for storing a computer program; The processor is used for executing the computer program for: Obtain a ticket-purchasing voice, and determine the to-be-verified ticket-purchasing text corresponding to the ticket-purchasing voice; Recombine the ticket-purchasing names in the to-be-verified ticket-purchasing text to generate at least one candidate name for the ticket-purchasing name; Split the at least one candidate name according to the to-be-verified ticket-purchasing text to generate at least one split name; Determine the similarity between each split name and the to-be-verified ticket-purchasing text, and determine whether to correct the to-be-verified ticket-purchasing text according to the similarity.

35. A computing device, comprising: A memory and a processor; The memory is used for storing a computer program; The processor is used for executing the computer program for: Obtain an on-demand voice, and determine the to-be-verified on-demand text corresponding to the on-demand voice; Recombine the video names in the to-be-verified on-demand text to generate at least one candidate name for the video name; Split the at least one candidate name according to the to-be-verified on-demand text to generate at least one split name; Determine the similarity between each split name and the to-be-verified on-demand text, and determine whether to correct the to-be-verified on-demand text according to the similarity.

36. A computing device, comprising: A memory and a processor; The memory is used for storing a computer program; The processor is used for executing the computer program for: Obtain voice information, and determine the to-be-corrected voice text corresponding to the voice information; Segment the associated information of the to-be-corrected voice text to obtain at least one segmented information, and recombine the at least one segmented information to generate at least one candidate information, where the associated information represents the semantic information of the voice text that the user actually wants to express; Split the at least one candidate information according to the to-be-corrected voice text to generate at least one split information; Determine the similarity between each split information and the to-be-corrected voice text. When the similarity is greater than the threshold, determine to correct the to-be-corrected voice text.

37. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by one or more processors, the one or more processors are caused to implement the steps in the method according to any one of claims 1-27.

Citation Information

Patent Citations

  • Speech recognition error correction method and speech recognition error correction device used for intelligent hardware equipment

    CN106710592A

  • Character error correction method and device, electronic equipment and storage medium

    CN110472701A