Speech recognition method and apparatus, storage medium, and electronic device
Patent Information
- Application Number
- CN202211193604.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2042-09-28
AI Technical Summary
但是,当前的技术在进行语音识别时由于使用者的口音、同音字、算法局限等问题,很容易出现语音识别错误,从而智能设备就可能无法按照用户期望的意图进行工作,在这种情况下,用户只有通过手动的对语音识别结果进行修正或者编辑,才能使得智能设备继续运行,给用户带来了很不好的使用体验
[0015] In the embodiments of the present application, an error correction intention is recognized from a target voice to be converted into text, wherein the error correction intention is used to describe a character in text converted from the voice. In a case that the error correction intention is recognized from the target voice, the target voice is divided into a first paragraph and a second paragraph, wherein the first paragraph carries the target error correction intention in the target voice. The second paragraph is converted into text to obtain a candidate text. The candidate text is corrected by using the first paragraph to obtain a target text corresponding to the target voice, i.e. in the target voice, the first paragraph carrying the target error correction intention and the second paragraph to be converted into text content are included. In a case that the error correction intention is recognized from the target voice, the candidate text of the second paragraph is corrected by using the content of the first paragraph carrying the error correction intention, and then the target text is output, so that the target text can match the intention of the target voice. By using the above technical solution, the problem of low voice recognition efficiency in the related art is solved, and the technical effect of improving the efficiency of voice recognition is achieved.
Smart Images

Figure CN115810359B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of smart home, in particular to a speech recognition method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the increasing maturity of artificial intelligence related technologies, more and more smart devices begin to enter people's lives. These devices can interact with people and constantly provide convenience for people's production and life. In the interaction process, one of the most commonly used interaction methods is voice interaction. In the field of voice interaction, the speech recognition is performed on the user's voice, the voice is converted into corresponding text, and thus the key information carried in the voice or the user's intention information is obtained. Subsequently, the running state of the smart device can be directly controlled through the voice recognition result, so that the smart device can work according to the intention contained in the user's voice. However, in the current technology, due to the user's accent, homophonic words, algorithm limitations and other problems, voice recognition errors are likely to occur, so that the smart device may not work according to the user's expected intention. In this case, the user can only manually correct or edit the voice recognition result to make the smart device continue to run, which brings a very poor user experience.
[0003] An effective solution has not been proposed for the problem of low voice recognition efficiency in the related art. SUMMARY
[0004] Embodiments of the present application provide a speech recognition method and device, a storage medium and an electronic device to at least solve the problem of low voice recognition efficiency in the related art.
[0005] According to one of the embodiments of the present application, a speech recognition method is provided, including: recognizing a correction intention from a target speech to be converted into text, wherein the correction intention is used to describe a word in the text converted by using the speech; in the case that the correction intention is recognized from the target speech, dividing the target speech into a first paragraph and a second paragraph, wherein the first paragraph carries a target correction intention in the target speech; performing text conversion on the second paragraph to obtain a candidate text; using the first paragraph to correct the candidate text to obtain a target text corresponding to the target speech.
[0006] Optionally, the recognizing a correction intention from a target speech to be converted into text includes: receiving a text conversion request, wherein the text conversion request is used to request to convert the target speech into text; and responding to the text conversion request, recognizing a statement description of a target format from the target speech, wherein the target format is a language expression format used to describe a word style of a word to be corrected.
[0007] Optionally, the identifying the sentence description of the target format from the target speech comprises at least one of: retrieving a sentence description including the to-be-corrected text structure from the target speech; and retrieving a sentence description including a word using the to-be-corrected text from the target speech.
[0008] Optionally, the retrieving the sentence description including the word using the to-be-corrected text from the target speech comprises: obtaining a target string corresponding to the target speech; in a case where a first string corresponding to a reference keyword exists in the target string, detecting whether a second string located before the first string in the target string includes a third string located after the first string; in a case where the second string includes the third string, matching the second string with a word string included in a target dictionary; and in a case where part or all of the word string in the second string matches the word string in the target dictionary, determining that the sentence description including the word using the to-be-corrected text exists in the target speech.
[0009] Optionally, the dividing the target speech into the first paragraph and the second paragraph comprises: dividing the target speech into a plurality of speech segments according to semantic meanings; extracting a target speech segment expressing the correction intention from the plurality of speech segments as the first paragraph, and determining other speech segments in the plurality of speech segments except the target speech segment as the second paragraph.
[0010] Optionally, the correcting the candidate text using the first paragraph to obtain the target text corresponding to the target speech comprises: converting the first paragraph into a target correction text; and correcting the corresponding text in the candidate text using the target correction text to obtain the target text.
[0011] Optionally, the converting the first paragraph into the target correction text comprises: converting the first paragraph into a target string, wherein the target string is used to indicate pronunciation of the first paragraph; extracting a key string expressing the target correction intention from the target string; and obtaining the target correction text corresponding to the key string.
[0012] Optionally, the extracting the key string expressing the target error correction intention from the target string comprises: in a case that a language expression format of the target string is used to describe a structure of the target error correction character, determining the target string as the key string, wherein the key string corresponds to the target error correction character in a string corresponding to a character; in a case that the language expression format of the target string is used to describe a target word using the target error correction character, determining a string corresponding to the target word as the key string, wherein the key string corresponds to the target error correction character in a string corresponding to a character.
[0013] According to a further aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores a computer program. The computer program is configured to perform the voice recognition method when executed.
[0014] According to a further aspect of the embodiments of the present application, an electronic device is provided, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor. The processor performs the voice recognition method by executing the computer program.
[0015] In the embodiments of the present application, an error correction intention is recognized from a target voice to be converted into text, wherein the error correction intention is used to describe a character in text converted from the voice. In a case that the error correction intention is recognized from the target voice, the target voice is divided into a first paragraph and a second paragraph, wherein the first paragraph carries the target error correction intention in the target voice. The second paragraph is converted into text to obtain a candidate text. The candidate text is corrected by using the first paragraph to obtain a target text corresponding to the target voice, i.e. in the target voice, the first paragraph carrying the target error correction intention and the second paragraph to be converted into text content are included. In a case that the error correction intention is recognized from the target voice, the candidate text of the second paragraph is corrected by using the content of the first paragraph carrying the error correction intention, and then the target text is output, so that the target text can match the intention of the target voice. By using the above technical solution, the problem of low voice recognition efficiency in the related art is solved, and the technical effect of improving the efficiency of voice recognition is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles behind the present application.
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, for those skilled in the field, other drawings can also be obtained based on these drawings without any creative effort.
[0018] Figure 1 Fig. 1 is a hardware environment schematic diagram of a voice recognition method according to an embodiment of the present application;
[0019] Figure 2 Fig. 2 is a flow chart of a voice recognition method according to an embodiment of the present application;
[0020] Figure 3 Fig. 3 is an optional voice segment division schematic diagram according to an embodiment of the present application;
[0021] Figure 4 Fig. 4 is an optional character correction flow according to an embodiment of the present application; Figure 1
[0022] Figure 5 Fig. 5 is an optional character correction flow according to an embodiment of the present application; Figure 2
[0023] Figure 6 Fig. 6 is a structural block diagram of a voice recognition device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the person skilled in the art better understand the present application, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should be within the scope of protection of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to one aspect of the embodiments of this application, a voice recognition method is provided. This method is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and smart house ecosystems. Optionally, Figure 1 This is a schematic diagram of the hardware environment for a speech recognition method according to an embodiment of this application. In this embodiment, the above method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0028] This embodiment provides a speech recognition method, applied to the aforementioned device terminal. Figure 2 This is a flowchart of a speech recognition method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0029] Step S202: Identify the error correction intent from the target speech to be converted into text, wherein the error correction intent is used to describe the text in the speech-to-text conversion.
[0030] Step S204, in the case that the error correction intention is recognized from the target voice, the target voice is divided into a first paragraph and a second paragraph, wherein the first paragraph carries the target error correction intention in the target voice;
[0031] Step S206, text conversion is performed on the second paragraph to obtain a candidate text;
[0032] Step S208, the candidate text is corrected using the first paragraph to obtain a target text corresponding to the target voice.
[0033] Through the above steps, the target voice includes the first paragraph carrying the target error correction intention and the second paragraph to be converted into text content. In the case that the error correction intention is recognized from the target voice, the candidate text of the second paragraph is corrected using the content of the first paragraph carrying the error correction intention, and then the target text is output. Therefore, the output target text can match the intention of the target voice. The above technical solution solves the problem of low voice recognition efficiency in related technologies, and achieves the technical effect of improving the efficiency of voice recognition.
[0034] In the technical solution provided in the above step S202, the error correction intention in the target voice can be recognized by recognizing the target keyword in the voice for indicating the error correction intention, and then the voice content including the target keyword is determined as the voice content corresponding to the error correction intention. For example, when the keywords such as "change to", "correct to", "change" are detected in the target voice, it is determined that the target voice has an error correction intention, and the voice content including the keywords is used as the voice content for error correction.
[0035] Optionally, in the present embodiment, the error correction intention in the target voice can be recognized by recognizing the format of the voice content in the target voice. When the voice content in the recognized voice includes a target format, it is determined that the part of the voice content is the voice content corresponding to the error correction intention. The target format is a language expression format for describing the modified text, such as a well-known idiom or a poem. The idiom or sentence includes one or more characters with the same pronunciation as the characters included in other parts of the content in the target voice.
[0036] Optionally, in the present embodiment, the error correction intention can be described by describing the composition structure of the character, the stroke order of the character, the commonly used word combination of the character, or the well-known sentence, such as "Wen Dao Liu" to describe the structure of the Chinese character Liu, so as to indicate that the character "Liu" needs to be replaced in the voice, such as "one stroke and one nodule" to describe the stroke order of the Chinese character Ren, so as to indicate that the character "Ren" needs to be replaced in the voice, and such as the idiom "Sheng Yuan Ji Di" to describe the character "Di".
[0037] In the technical solution provided in step S204, the order of the first paragraph and the second paragraph in the target speech can be any order, such as the first paragraph first and the second paragraph second, the second paragraph first and the first paragraph second, or the second paragraph being located among multiple first paragraphs, such as the target speech including one second paragraph to be converted into text and multiple first paragraphs, the second paragraph being at the beginning of the target speech, at the end of the target speech, or being inserted among the multiple first paragraphs, and the present solution does not make a limitation on this.
[0038] Optionally, in the present embodiment, the second paragraph can be one continuous speech content in the target speech, or can also be obtained by splicing multiple discontinuous speech contents in the target speech, such as, when dividing, the target speech can be processed by punctuation according to the short interval between speeches, multiple speech segments are obtained, the target speech segment expressing the correction intention is taken as the first paragraph, and the speech segments other than the target speech segment in the multiple speech segments are taken as the second paragraph, or after identifying the multiple speech segments, it is found that all the multiple speech segments are the first speech segment, the correction intention of the first speech segment is identified, the speech content corresponding to the to-be-corrected text in the correction intention is extracted, and the second paragraph is obtained by splicing in the order. For example, the user wants to output the text content of “I want to go to Mount Tai”, the user can first say “I want to go to Mount Tai” and then say the target correction intention corresponding to the text to be corrected, the sentence is punctuated, and the target speech segment corresponding to the correction intention is taken as the first paragraph, and “I want to go to Mount Tai” is taken as the second speech paragraph, or the user can input the speech of 5 target correction intentions in turn, and the 5 correction intentions are used to describe “I”, “want”, “go”, “Tai”, and “mountain” respectively, and then the second paragraph is obtained by extracting the part of the 5 correction intentions corresponding to the text to be modified.
[0039] In the technical solution provided in step S206, the candidate text can be a text, or can also be a pinyin text corresponding to the target speech.
[0040] In the technical solution provided in step S208, the way of correcting the candidate text using the first paragraph can include, but is not limited to, changing the text in the candidate text or adding text content in the candidate text, such as determining the target text to be corrected according to the first paragraph, and replacing the text with the same pronunciation as the target text in the candidate text with the target text, or determining the target text to be added in the candidate text and the position of the target text in the candidate text according to the first paragraph, and adding the target text in the candidate text to obtain the target text.
[0041] As an optional embodiment, identifying a correction intention from the target speech to be converted into text includes:
[0042] Receiving a text conversion request, where the text conversion request is used to request converting the target speech into text;
[0043] Responding to the text conversion request, and identifying a statement description in a target format from the target speech, where the target format is a language expression format for describing the text style of the text to be corrected.
[0044] Optionally, in this embodiment, identifying a statement description in a target format from the target speech in a target format may, but is not limited to, be based on the pinyin text corresponding to the target speech, and then identifying based on the pinyin text, so as to obtain a statement description in a target format. <000010-seven>Optionally, in this embodiment, the statement description in a target format may, but is not limited to, include a description statement of the writing structure of the text, a description statement of the writing stroke order of the text, the word formation method of the text, or a poem containing the text, etc. For example, "Wen Dao Liu" is a description statement of the writing structure of the Chinese character "Liu", and through this statement, the text style of the text to be corrected with the pronunciation of "liu" in the candidate text is described, or the description statement "A person, one stroke down and one stroke across" describes the writing stroke order of the character "person", and through this statement, the text style of the text to be corrected with the pronunciation of "ren" in the candidate text is described. Another example is that "The 'Di' in '状元及第' " describes the Chinese character "Di", and through this statement, the text style of the text to be corrected with the pronunciation of "di" in the candidate text is described.
[0046] As an optional embodiment, identifying a statement description in a target format from the target speech includes at least one of the following:
[0047] Retrieving a statement description including the structure of the text to be corrected from the target speech;<OO00112>
[0048] Retrieving a statement description including a word using the text to be corrected from the target speech.
[0049] Optionally, in this embodiment, the method of retrieving a statement description including the structure of the text to be corrected from the target speech may be by matching the target speech with a preset text structure description statement. For example, converting the target speech into a pinyin text, and matching the pinyin text of the pre-trial structure description statement with the pinyin text corresponding to the target speech, so as to identify a statement description in a first format from the target speech.<OO00116>
[0050] Optionally, in the embodiment, the manner of retrieving the sentence description including the word using the to-be-corrected character from the target voice can be a manner of recognizing the target voice by using a preset character description word. For example, the second format sentence description included in the target voice is "zhuang yuan ji di de di", which describes the to-be-corrected character "di" by the word "zhuang yuan ji di". Therefore, the target voice is matched by using the preset description word, and in a case where the pronunciation of the preset word and the word "zhuang yuan ji di" included in the target voice are matched, it is determined that "zhuang yuan ji di de di" is the second format sentence description.
[0051] Figure 3 An optional voice segment division schematic diagram according to an embodiment of the present application, as shown in FIG. 3, can include but is not limited to the following contents: Figure 3
[0052] S301, a target voice is acquired, the target voice includes two parts of contents, i.e., a second paragraph to be converted into text and a first paragraph used for correcting the content in the second paragraph;
[0053] S302, the target voice is matched by using a preset character structure description sentence, a sentence description including a to-be-corrected character structure is obtained, and the sentence description including the to-be-corrected character structure is determined as the first paragraph;
[0054] S303, the target voice is matched by using a preset character description word, a sentence description including a word using the to-be-corrected character is obtained, and the sentence description including the word using the to-be-corrected character is determined as the first paragraph;
[0055] S304, the content of the target voice except the first paragraph is determined as the second paragraph.
[0056] As an optional embodiment, the retrieving the sentence description including the word using the to-be-corrected character from the target voice includes:
[0057] acquiring a target string corresponding to the target voice;
[0058] in a case where a first string corresponding to a reference keyword exists in the target string, detecting whether a second string located before the first string in the target string includes a third string located after the first string;
[0059] in a case where the second string includes the third string, matching the second string with a word string included in a target dictionary;
[0060] In the case that part or all of the matched character strings in the second character string exist in the target dictionary, it is determined that the target speech includes a sentence description using the character string with the to-be-corrected character.
[0061] Optionally, in the embodiment, the reference keyword is a keyword in a sentence description using the character string with the to-be-corrected character, and the reference keyword may, but is not limited to, include "of", "xth character in", for example, the sentence description "of the 4th character in" is a sentence description using the character string (zhuangyuan jidi) with the to-be-corrected character "4th", and the keyword "of the 4th character in" in the sentence description is used to indicate that the character "4th" in the four characters of "zhuangyuan jidi" is the to-be-corrected character.
[0062] Optionally, in the embodiment, the second character string includes a third character string used to indicate that part or all of the character strings in the second character string are the same as the third character string, for example, for the sentence description "of the 4th character in", the reference keyword is "of", the second character string is the pinyin text corresponding to "zhuangyuan jidi", the third character string is the pinyin text of the character "4th" after the reference keyword "of", and the character string corresponding to "4th" in the second character string is the same as the third character string.
[0063] Optionally, in the embodiment, the character strings in the target dictionary are determined according to the recognition result of the reference speech in the historical time period, or can also be preset common idioms or character strings, etc., for example, in the case that the reference text is recognized from the historical time reference speech, the character strings or sentence descriptions in the reference text are used as target character strings for subsequent text recognition.
[0064] As an optional embodiment, the dividing the target speech into a first paragraph and a second paragraph includes:
[0065] dividing the target speech into a plurality of speech segments according to semantic meanings;
[0066] extracting a target speech segment expressing a correction intention from the plurality of speech segments as the first paragraph, and determining other speech segments in the plurality of speech segments except the target speech segment as the second paragraph.
[0067] Optionally, in the embodiment, the target speech segment can be extracted from the plurality of speech segments by detecting whether the speech segment has a target keyword corresponding to the error correction intention, the target keyword can include but is not limited to "delete", "modify", "change" and the like, the target keyword can be determined according to the historical speech habits of the user, for example, the user often uses the keywords "modify" and "change" in the past period of time, so "modify" and "change" can be directly used as the target keyword, or the keywords can be determined according to the text error correction in the past period of time, the reference word vector of the two keywords is calculated, and the speech segment is divided by word vector calculation, and the keywords matching the reference word vector in the speech segment are used as the target keyword.
[0068] Optionally, in the embodiment, the target speech segment can be extracted from the plurality of speech segments by detecting whether the speech segment has an error correction template corresponding to the error correction intention, the error correction template is used to describe the character style of the target error correction character, for example, the target error correction character is described by describing the writing order of the target error correction character, the target error correction character is described by describing the character structure of the target error correction character, or the target error correction character is described by including the reference word of the target error correction character, for example, the error correction template "Mu Zi Li" is used to describe the character structure of the Chinese character "Li", and the error correction template "Zhuang Yuan Ji Di's Di" is used to describe the character style of the Chinese character "Di" by including the well-known Chinese idiom including the Chinese character "Di". The error correction template can be determined by recognizing the speech habits of the user, or can be manually set by the user according to the needs of the user.
[0069] As an optional embodiment, the using the first paragraph to correct the candidate text to obtain the target text corresponding to the target speech includes:
[0070] Converting the first paragraph into a target error correction character;
[0071] Using the target error correction character to correct the corresponding character in the candidate text to obtain the target text.
[0072] Optionally, in the embodiment, the target error correction character can be generated according to the target description statement included in the first paragraph for describing the character style of the to-be-corrected character, or the target error correction character corresponding to the target description statement can be determined from the description statement and the error correction word having a corresponding relationship, for example, the description statement is a statement for describing the stroke order of the character or the composition structure of the character, and then the character generation model is used to generate the target error correction character corresponding to the target description statement.
[0073] In the above embodiment, the first paragraph carries the target error correction intention in the target voice, the first paragraph can record the sentence description in the first format for describing the structure of the to-be-corrected character, or can also record the sentence description in the second format for describing the word using the to-be-corrected character, and then the target error correction character can be determined through the first paragraph, Figure 4 is an optional character correction process according to the application embodiment Figure 1 , as shown in the following figure, Figure 4 may include, but not limited to, the following steps:
[0074] S401, obtaining a target voice, the target voice including a quantity part, i.e. a second paragraph to be converted into text, and a first paragraph for correcting the content in the second paragraph, and converting the target text into a pinyin string in pinyin format;
[0075] S402, the preset dictionary stores a sentence description for describing a to-be-corrected character structure, the sentence description including the to-be-corrected character structure, the description including a Chinese character component and a word composed of the component, and a corresponding pinyin string of the sentence description, such as "early chapter (lizaozhang)", "bow long Zhang (gongchangzhang)", "white spoon (baishaode / baishaodi)", "soil also (tuyede / tuyedi)";
[0076] S403, determining whether there is a sentence description including the to-be-corrected character structure in the target voice by matching the pinyin string corresponding to the target text and the pinyin string corresponding to the sentence description including the to-be-corrected character structure;
[0077] S404, in the case that there is a sentence description including the to-be-corrected character structure, determining a target error correction character according to the sentence description, i.e. "chapter", "Zhang", "of" and the like.
[0078] S405, detecting whether there is a to-be-corrected character with the same pronunciation as the target error correction character in the text (candidate text) corresponding to the second paragraph of the target voice;
[0079] S406, in the case that there is a to-be-corrected character, replacing the to-be-corrected character in the candidate text corresponding to the target voice with the target error correction character;
[0080] S407, outputting the target text after text correction.
[0081] Figure 5 is an optional character correction process according to the application embodiment Figure 2 , as shown in the following figure, Figure 5 may include, but not limited to, the following steps:
[0082] S501, obtaining a target voice, the target voice including a quantity part, i.e., a second paragraph to be converted into text, and a first paragraph for correcting content in the second paragraph, and converting the target text into a pinyin string in a pinyin format;
[0083] S502, detecting whether there is a character with the same pronunciation as the keyword "de" in the target text;
[0084] S503, in the case where there is a character with the same pronunciation as the keyword "de" in the target text, detecting whether there is a target word before the keyword, the target word including a character with the same pronunciation as a character at a position adjacent to the keyword, such as "zhuangyuanjidi" including the homophonic character "di";
[0085] S504, in the case where there is the target word, detecting whether the target word is located in a preset dictionary, the preset dictionary recording public words, commonly used words of a user, and a pinyin string of the commonly used words, such as "jiabin", "zhihui", "jiayibingding", "zhuangyuanjidi", and the like;
[0086] S505, in the case where the target word is in the preset dictionary, determining a target correction character as a character in the target word with the same pronunciation as a character after the keyword;
[0087] S506, replacing a to-be-corrected character in a candidate text corresponding to the target voice with the target correction character;
[0088] S507, outputting a target text after text correction.
[0089] As an optional embodiment, the converting the first paragraph into the target correction character includes:
[0090] converting the first paragraph into a target string, wherein the target string is used to indicate pronunciation of the first paragraph;
[0091] extracting a key string expressing the target correction intention from the target string;
[0092] obtaining the target correction character corresponding to the key string.
[0093] Optionally, in the embodiment, the target character can be a pinyin string of the first paragraph, or a pinjiazi string, or a string of English words, which is not limited by the present scheme.
[0094] Optionally, in this embodiment, the key string is extracted from the target string by detecting whether the target string includes a string identical to the pronunciation of the key word or the error correction template corresponding to the target error correction intention, such as the converted pinyin format target string "lizaozhang", which is identical to the pinyin text of the error correction template "lizaozhang", so the string is determined as the key string, or the converted pin y format target string includes the pinyin string "genggai" corresponding to the key word "genggai", so the string including the pinyin string can be determined as the key string.
[0095] As an optional embodiment, the extracting the key string expressing the target error correction intention from the target string comprises:
[0096] In the case where the language expression format of the target string is used to describe the structure of the target error correction character, the target string is determined as the key string, wherein the character corresponding to the string having the corresponding relationship in the key string is the target error correction character.
[0097] In the case where the language expression format of the target string is used to describe the target word using the target error correction character, the string corresponding to the target word is determined as the key string, wherein the character corresponding to the string having the corresponding relationship in the key string is the target error correction character.
[0098] Optionally, in this embodiment, the target word can be determined according to the recognition result of the reference voice in the historical time period, or can also be a preset well-known idiom or word, etc., such as the reference text recognized by the historical time reference voice memory, the word or sentence in the reference text is used as the target word for subsequent text recognition.
[0099] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the methods of various embodiments of the present application.
[0100] Figure 6 is a structural block diagram of a voice recognition device according to an embodiment of the present application; as Figure 6As shown, the method comprises: identifying a correction intention from target speech to be converted into text, wherein the correction intention is used to describe a word in text converted from the target speech;
[0101] processing the target speech into a first paragraph and a second paragraph, wherein the first paragraph carries a target correction intention in the target speech;
[0102] converting the second paragraph into candidate text;
[0103] correcting the candidate text using the first paragraph to obtain target text corresponding to the target speech.
[0104] In the above embodiment, a correction intention is identified from target speech to be converted into text, wherein the correction intention is used to describe a word in text converted from the target speech. In the case where the correction intention is identified from the target speech, the target speech is divided into a first paragraph and a second paragraph, wherein the first paragraph carries a target correction intention in the target speech. The second paragraph is converted into candidate text. The candidate text is corrected using the first paragraph to obtain target text corresponding to the target speech. That is, the target speech includes the first paragraph carrying the target correction intention and the second paragraph to be converted into text content. In the case where the correction intention is identified from the target speech, the candidate text of the second paragraph is corrected using the content of the first paragraph carrying the correction intention, and then the target text is output. Therefore, the target text can match the intention of the target speech. The above technical solution solves the problem of low speech recognition efficiency in related technologies, and achieves the technical effect of improving the efficiency of speech recognition.
[0105] Optionally, the identifying module comprises: a receiving unit configured to receive a text conversion request, wherein the text conversion request is used to request conversion of the target speech into text; and an identifying unit configured to identify a target format sentence description from the target speech in response to the text conversion request, wherein the target format is a language expression format used to describe a word style of to-be-corrected words.
[0106] Optionally, the identifying unit is configured to perform at least one of the following operations: retrieving a sentence description including a structure of the to-be-corrected words from the target speech; and retrieving a sentence description including a word using the to-be-corrected words from the target speech.
[0107] Optionally, the identifying unit is further configured to: acquire a target character string corresponding to the target speech; in a case where a first character string corresponding to a reference keyword exists in the target character string, detect whether a second character string located before the first character string in the target character string includes a third character string located after the first character string; in a case where the second character string includes the third character string, match the second character string with a character string included in a target dictionary; and in a case where part or all of the second character string matches a consistent character string included in the target dictionary, determine that a sentence description using a character included in the to-be-corrected text exists in the target speech.
[0108] Optionally, the processing module includes: dividing the target speech into a plurality of speech segments according to semantic meanings; extracting a target speech segment expressing a correction intention in a semantic meaning from the plurality of speech segments as the first paragraph, and determining other speech segments in the plurality of speech segments except the target speech segment as the second paragraph.
[0109] Optionally, the correction module includes: a conversion unit configured to convert the first paragraph into a target correction text; and a correction unit configured to correct a corresponding text in the candidate text using the target correction text to obtain the target text.
[0110] Optionally, the conversion unit is configured to: convert the first paragraph into a target character string, where the target character string is used to indicate pronunciation of the first paragraph; extract a key character string expressing the target correction intention from the target character string; and acquire the target correction text corresponding to the key character string.
[0111] Optionally, the conversion unit is configured to: in a case where a language expression format of the target character string is used to describe a structure of the target correction text, determine the target character string as the key character string, where the key character string corresponds to the target correction text in a character string and a character having a corresponding relationship; and in a case where the language expression format of the target character string is used to describe a target character string using the target correction text, determine a character string corresponding to the target character string as the key character string, where the key character string corresponds to the target correction text in a character string and a character having a corresponding relationship.
[0112] Embodiments of the present application also provide a storage medium including a stored program, wherein the program performs the speech recognition method of any of the above when running.
[0113] Optionally, in the embodiment, the storage medium can be configured to store program code for performing the following steps: identifying a correction intention from a target voice to be converted into text, wherein the correction intention is used to describe a word in text converted from the target voice; in a case where the correction intention is identified from the target voice, dividing the target voice into a first paragraph and a second paragraph, wherein the first paragraph carries a target correction intention in the target voice; performing text conversion on the second paragraph to obtain a candidate text; and correcting the candidate text using the first paragraph to obtain target text corresponding to the target voice.
[0114] The embodiment of the present application further provides an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the speech recognition method embodiments.
[0115] Optionally, the electronic device can further include a transmission device and an input / output device, wherein the transmission device is connected with the processor, and the input / output device is connected with the processor.
[0116] Optionally, in the embodiment, the processor can be configured to perform the following steps by the computer program: identifying a correction intention from a target voice to be converted into text, wherein the correction intention is used to describe a word in text converted from the target voice; in a case where the correction intention is identified from the target voice, dividing the target voice into a first paragraph and a second paragraph, wherein the first paragraph carries a target correction intention in the target voice; performing text conversion on the second paragraph to obtain a candidate text; and correcting the candidate text using the first paragraph to obtain target text corresponding to the target voice.
[0117] Optionally, in the embodiment, the storage medium can include but is not limited to a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk and various storage program codes.
[0118] Optionally, specific examples in the embodiment can refer to examples described in the above embodiments and optional implementation manners, and the embodiment will not be described here.
[0119] It is apparent that those skilled in the art can, without departing from the spirit of the present application, make various changes and modifications of the modules or steps of the present application described above, which can be implemented by general computing devices, and can be centralized on a single computing device or distributed on a network composed of multiple computing devices, and optionally, can be implemented by program codes executable by computing devices, so that they can be stored in storage devices and executed by computing devices, and in some cases, the steps shown or described can be executed in different order, or can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module. Thus, the present application is not limited to any particular combination of hardware and software.
[0120] The above description is only the preferred embodiments of the present application, and it should be pointed out that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A speech recognition method, characterized by, The method comprises the following steps: recognizing a correction intention from a target speech to be converted into text, wherein the correction intention is used to describe a word in text converted from the target speech; in the case that the correction intention is recognized from the target speech, dividing the target speech into a first paragraph and a second paragraph, wherein the first paragraph carries a target correction intention in the target speech; performing text conversion on the second paragraph to obtain a candidate text; correcting the candidate text using the first paragraph to obtain a target text corresponding to the target speech; recognizing a correction intention from a target speech to be converted into text, comprising: obtaining a target string corresponding to the target speech; in the case that a first string corresponding to a reference keyword exists in the target string, detecting whether a second string located before the first string in the target string includes a third string located after the first string; in the case that the second string includes the third string, matching the second string with a word string included in a target dictionary; in the case that part or all of the word string in the second string matches the word string in the target dictionary, determining that a sentence in the target speech includes a word using a word to be corrected.
2. The method of claim 1, wherein, The method comprises the following steps: receiving a text conversion request, wherein the text conversion request is used to request conversion of the target speech into text; in response to the text conversion request, recognizing a sentence description in a target format from the target speech, wherein the target format is a language expression format used to describe a word style of a word to be corrected.
3. The method of claim 2, wherein, The method comprises the following steps: retrieving a sentence description including a structure of the word to be corrected from the target speech.
4. The method of claim 1, wherein, The method comprises the following steps: dividing the target speech into a plurality of speech segments according to semantic meanings; extracting a target speech segment expressing a correction intention from the plurality of speech segments as the first paragraph, and determining other speech segments in the plurality of speech segments except the target speech segment as the second paragraph.
5. The method of claim 1, wherein, The method comprises the following steps: converting the first paragraph into a target correction word; correcting a corresponding word in the candidate text using the target correction word to obtain the target text.
6. The method of claim 5, wherein, The method comprises the following steps: converting the first paragraph into a target string, wherein the target string is used to indicate pronunciation of the first paragraph; extracting a key string expressing the target correction intention from the target string; obtaining the target correction word corresponding to the key string.
7. The method of claim 6, wherein, The method comprises the following steps: In a case where the language expression format of the target string is used to describe the structure of the target error correction character, the target string is determined as the key string, wherein the key string is the target error correction character corresponding to the character and the character in the string having the corresponding relationship; In a case where the language expression format of the target string is used to describe the target word using the target error correction character, the string corresponding to the target word is determined as the key string, wherein the key string is the target error correction character corresponding to the character and the character in the string having the corresponding relationship.
8. A computer readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the program executes the method of any one of claims 1 to 7 when running. 9.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Correcting voice recognition using selective re-speak
CN106095766A
Apparatus, method, and computer program product for correcting speech recognition error
US20170270086A1
Electronic device for correcting user's voice input and method for operating same
WO2022186435A1