Speech recognition method and related device
By replacing characters with confidence below the threshold in the first character string in the ASR system, the target character string is obtained, which solves the problem of inaccurate speech recognition in the prior art and improves the accuracy of the recognition results.
Patent Information
- Application Number
- CN202510357255.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-25
AI Technical Summary
When existing ASR systems process speech input containing alphanumeric, they are affected by factors such as accent, speech speed and clarity, resulting in inaccurate recognition, especially when entering character strings such as device serial number, which makes it difficult for the system to accurately identify.
By obtaining the voice information to be recognized, the first character string is identified, and the character whose confidence is lower than the set threshold is replaced based on the confidence of each character in the first character string, the target character string is obtained, and the target character string is finally output as the recognition result.
The accuracy of speech recognition is improved, especially when processing character strings containing alphanumeric numbers, by replacing characters with low confidence, a recognition result with a higher degree of matching with the speech information to be recognized is obtained.
Smart Images

Figure CN119993131A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing, and in particular to a speech recognition method and related devices. Background Art
[0002] As a common input method, voice input has been widely used in various scenarios.
[0003] For voice input, the ASR (Automatic Speech Recognition) system is currently generally used to process and obtain recognition results.
[0004] When users input various alphanumeric information through voice, inaccurate voice recognition may occur due to factors such as accent, speaking speed, and clarity. Especially when inputting alphanumeric strings such as device serial numbers, due to the limitations of the ASR system, the system often cannot accurately identify the available strings when processing the input voice, resulting in incorrect recognition results. Summary of the invention
[0005] The first aspect of the present application provides a speech recognition method, comprising:
[0006] Obtaining voice information to be recognized;
[0007] Recognize the voice information to be recognized to obtain a first character string;
[0008] Replacing characters in the first character string whose confidences meet a replacement condition to obtain a target character string;
[0009] The target character string is output as a recognition result.
[0010] In a possible implementation, replacing characters in the first character string whose confidences meet a replacement condition to obtain a target character string includes:
[0011] Determining, based on the confidences of the characters in the first character string, that the first character in the first character string meets a replacement condition, and that the confidence of the first character is lower than a set threshold;
[0012] The first character is replaced according to the second character in the first character string to obtain a target character string, wherein the matching degree between the target character string and the voice information to be recognized is greater than the matching degree between the first character string and the voice information to be recognized.
[0013] In a possible implementation, the replacing the first character according to the second character in the first character string to obtain the target character string includes:
[0014] According to the second character in each first character string and the position of the second character, at least one candidate character string is obtained by searching in a preset database, wherein the characters in the candidate character string are consistent with the second character in the corresponding first character string, and the number of the first character string is at least one;
[0015] Determining the matching degree between each candidate character string and each first character string in turn;
[0016] According to the matching degree, one of the at least one candidate character string is selected as a first target character string, and the matching degree between the first target character string and each first character string is higher than the matching degree between non-first target character strings in the candidate character string and each first character string.
[0017] In a possible implementation, the replacing the first character according to the second character in the first character string to obtain the target character string includes:
[0018] According to the second character in each first character string and the position of the second character, no candidate character string is found in the preset database, and the number of the first character string is at least one;
[0019] According to the second character in each first character string, a preset transition probability matrix is searched to obtain at least one target character to which any second character can be transferred, wherein the target character corresponds to the first character in each first character string, and the target character is a character sorted after the second character, and the preset transition probability matrix includes a number of characters, a transfer character of each character, and a transition probability corresponding to each character and the transfer character;
[0020] A second target character string is determined according to the second characters in each first character string and the corresponding target characters.
[0021] In a possible implementation, determining the second target character string according to the second characters in each first character string and the corresponding target characters includes:
[0022] Combining the target character with a corresponding second character in the first character string to obtain at least two character strings to be selected;
[0023] Query the preset transition probability matrix to obtain the transition probability corresponding to each target character in any candidate character string;
[0024] Determining the weight of the arbitrary character string to be selected according to the transition probability corresponding to each target character in the arbitrary character string to be selected;
[0025] A second target character string is selected from each candidate character string, wherein the weight of the second target character string is higher than the weight of non-second target character strings from each candidate character string.
[0026] In a possible implementation, the selecting a second target character string from the candidate character strings includes:
[0027] Sort the candidate strings according to their weights;
[0028] According to the ranking of each candidate character string, querying from a preset database whether the candidate character string exists;
[0029] An existing target character string to be selected is queried from a preset database as a second target character string, and the weight of the target character string to be selected is higher than the weight of the non-target character string to be selected.
[0030] In a possible implementation, before obtaining the voice information to be recognized, the method further includes:
[0031] Obtaining device information provided by a user terminal;
[0032] Accordingly, outputting the target character string as a recognition result includes:
[0033] Based on the matching of the target character string and the device information, predicting the service process corresponding to the voice information to be recognized;
[0034] The service process is used to output the recognition result to a service terminal, and the service terminal is used to provide services to each user terminal.
[0035] In a possible implementation, obtaining the voice information to be recognized includes:
[0036] Based on the response to the call request, the voice information to be recognized is obtained in real time;
[0037] Accordingly, outputting the target character string as a recognition result includes:
[0038] Count the first times the target string is obtained in this call;
[0039] Based on the first number of times being less than a first number threshold, outputting the target character string as a recognition result;
[0040] Based on the fact that the first number is not less than a first number threshold, outputting the target character string as a recognition result is prohibited.
[0041] In a possible implementation, the method further includes:
[0042] Obtain relevant information of the user terminal;
[0043] Based on the relevant information of the user terminal, counting a second number of times that the user terminal obtains the target character string during the call within a preset time period;
[0044] Based on the second number being not less than a second number threshold, and according to the relevant information of the target terminal, it is prohibited to respond to the voice information to be recognized provided by the target terminal.
[0045] A second aspect of the present application provides a sound recognition device, comprising:
[0046] Interface, used to obtain the voice information to be recognized;
[0047] A processor is used to identify the speech information to be recognized to obtain a first character string; replace the characters in the first character string whose confidence meets the replacement condition to obtain a target character string; and output the target character string as a recognition result
[0048] The third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the speech recognition method of the first aspect or any implementation of the first aspect.
[0049] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0050] The memory is used to store computer programs;
[0051] The processor is used to execute the computer program so that the electronic device can implement the speech recognition method of the above-mentioned first aspect or any implementation manner of the first aspect.
[0052] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the speech recognition method of the above-mentioned first aspect or any implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0054] Figure 1 It is a flowchart of a speech recognition method provided by an embodiment of the present application;
[0055] Figure 2 is a flowchart of replacing characters in the first character string whose confidences meet the replacement condition to obtain a target character string provided by an embodiment of the present application;
[0056] Figure 3is a schematic diagram of a first character string provided in an embodiment of the present application;
[0057] Figure 4 is a flowchart of replacing characters in the first character string whose confidences meet the replacement condition to obtain a target character string provided by an embodiment of the present application;
[0058] Figure 5 is a flow chart of obtaining a target character string by replacing the first character with the second character in the first character string provided by an embodiment of the present application;
[0059] Figure 6 It is a flowchart of determining a second target character string according to the second character in each first character string and the corresponding target character provided by an embodiment of the present application;
[0060] Figure 7 is a schematic diagram of a flow chart of selecting a second target character string from among the candidate character strings provided in an embodiment of the present application;
[0061] Figure 8 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0062] Fig. 9 is a flow chart of outputting the target character string as a recognition result provided by an embodiment of the present application;
[0063] Fig.10 is another flow chart of outputting the target character string as a recognition result provided by an embodiment of the present application;
[0064] Fig.11 It is a flowchart of the speech recognition method provided in the embodiment of the present application;
[0065] Fig.12 is a structural schematic diagram of a speech recognition device provided in an embodiment of the present application;
[0066] Fig.13 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0068] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0069] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0070] Reference Figure 1 , Figure 1 is a flow chart of a speech recognition method provided in an embodiment of the present application, such as Figure 1 As shown, a speech recognition method provided in an embodiment of the present application may include steps 101 to 104, and these steps are described in detail below.
[0071] 101. Obtain voice information to be recognized;
[0072] A speech recognition method provided in an embodiment of the present application is applied to an electronic device.
[0073] The voice information to be recognized may be voice information obtained from other devices, or voice information collected by the device.
[0074] As an example, the electronic device acts as a server and is connected to a terminal wirelessly or wiredly. The terminal has an audio acquisition device. After the terminal acquires voice information, it sends the voice information to the electronic device. The electronic device recognizes the voice information and obtains its corresponding character string.
[0075] As an example, the electronic device is provided with an audio collection device, and after collecting voice information, the audio collection device recognizes the character string in the voice information.
[0076] The voice information to be recognized may include multiple voice segments. During the recognition process, each voice segment may be recognized separately to obtain multiple corresponding target character strings.
[0077] The voice information to be recognized may be obtained in real time, and recognition is performed on the voice information obtained in real time to obtain a target character string.
[0078] The voice information to be recognized may also be obtained from a storage location, in which a segment of voice information is stored. The electronic device recognizes the segment of voice information obtained from the storage location to obtain a target character string.
[0079] 102. Recognize the voice information to be recognized to obtain a first character string;
[0080] The first character string may represent one character string or multiple character strings.
[0081] Herein, a recognition system may be used to recognize the voice information to be recognized to obtain a first character string.
[0082] As an example, the recognition system may employ ASR.
[0083] The first character string includes a plurality of characters, each of which has a confidence level, and the confidence level represents a quantitative reflection of the reliability of the recognition result of the recognition system.
[0084] For example, a higher confidence level indicates that the recognition system has a higher degree of reliability in the recognition result, the input speech features have a higher degree of match with the speech patterns in the model, and the speech is more recognizable. This means that the speech information to be recognized is more consistent with the patterns trained by the system at both the acoustic and linguistic levels, and is more likely to be an accurate recognition result; a low confidence level indicates that there may be greater uncertainty in the recognition result, which may be due to poor speech quality, the presence of uncommon vocabulary or language structures, or the speech features have a certain similarity with multiple model patterns, making it difficult for the system to make an accurate judgment. At this point, further analysis of the speech or other measures are needed to verify the accuracy of the recognition result.
[0085] Due to problems such as the user's non-standard accent, background noise or other factors, the first character string recognized by the ASR may contain characters with low confidence.
[0086] Therefore, the first character string needs to be further processed to improve the recognition accuracy of the voice information to be recognized.
[0087] 103. Replace characters in the first character string whose confidences meet a replacement condition to obtain a target character string;
[0088] The replacement condition may be that the confidence level is less than a set threshold.
[0089] In a possible implementation, each character in the first character string whose confidence level is less than a set threshold is replaced to obtain a target character string.
[0090] As an example, when the first character string represents a character string, characters in the character string whose confidence level is less than a set threshold may be replaced to obtain a corresponding target character string.
[0091] As an example, when the first character string represents multiple character strings, characters in each character string whose confidence is less than a set threshold may be replaced to obtain corresponding character strings, and subsequently one of the corresponding multiple character strings may be selected as a target character string.
[0092] 104. Output the target character string as a recognition result.
[0093] Among them, the target character string can be composed of only numbers, only letters, letters and numbers, or Chinese characters, letters and numbers. The letters can be one or more English letters or Greek letters. This application does not limit the format of the characters contained in the target character string.
[0094] The target character string is obtained by replacing characters whose confidences meet the replacement condition, and the target character string has a higher matching degree with the speech information to be recognized.
[0095] The recognition result may be output in the form of voice or display.
[0096] In a possible implementation, the target character string is converted into a speech form and outputted through an audio output device.
[0097] The audio may be outputted through an audio output device provided on the electronic device, or may be outputted through an audio output device on a terminal connected to the electronic device for communication.
[0098] In a possible implementation, the target character string is displayed and outputted in a setting display interface, and is audio-outputted through an audio output device.
[0099] The audio may be outputted through an audio output device provided on the electronic device, or may be outputted through an audio output device on a terminal connected to the electronic device for communication.
[0100] In this embodiment, after obtaining the voice information to be recognized, the voice information to be recognized is recognized to obtain a first character string; the characters in the first character string whose confidence meets the replacement condition are replaced to obtain a target character string; and the target character string is output as a recognition result. In this process, the recognition result is optimized by replacing the characters in the first character string whose confidence meets the replacement condition, and a target character string with higher accuracy is obtained, thereby improving the accuracy of voice information recognition. In the voice recognition scenario or telephone customer service scenario for the device SN (Serial Number), the recognition result can be optimized. Even if the surrounding noise is large or the user's accent causes the first character string to appear, the recognition result can be optimized and corrected through subsequent processes. The user does not need to input a string of characters multiple times by raising the voice or changing the accent, thereby optimizing the user experience, and does not need the program in the device or the telephone customer service program / personnel to remind multiple times, thereby simplifying the workflow.
[0101] Figure 2 This is a flowchart of replacing characters in the first character string whose confidences meet the replacement condition to obtain a target character string provided by an embodiment of the present application, which may include steps 201 to 202, and these steps are described in detail below.
[0102] 201. Determine, based on the confidences of the characters in the first character string, that a first character in the first character string satisfies a replacement condition, and that the confidence of the first character is lower than a set threshold;
[0103] After recognizing the voice information to be recognized, the recognition system obtains a first character string and the confidence level of each character in the first character string.
[0104] Among them, the recognition system converts the input voice information into text characters and generates a confidence level for each character.
[0105] Among them, if the recognition system has a higher degree of confidence in the recognition result, a higher confidence is set for the converted character; if the recognition system has a lower degree of confidence in the recognition result, a lower confidence is set for the converted character.
[0106] If the confidence of a character is lower than a set threshold, it can be determined that the character meets the replacement condition; otherwise, the character does not meet the replacement condition.
[0107] In a possible implementation, the confidence level may be expressed as a percentage, such as 90%, 52%, and so on.
[0108] As an example, the set threshold may be 70%. If the confidence of a character is less than 70%, it can be determined that the character meets the replacement condition. The character with a confidence of less than 70% in the first character string is the first character, and the character with a confidence of not less than 70% is the second character.
[0109] In a possible implementation, the confidence level may be represented by a decimal, such as 0.95, 0.76, etc.
[0110] As an example, the set threshold may be 0.75. If the confidence of a character is less than 0.75, it can be determined that the character meets the replacement condition. The character with a confidence less than 0.75 in the first character string is the first character, and the character with a confidence not less than 0.75 is the second character.
[0111] As an example, for a speech segment, the first character string obtained by recognition is: “3A75C24”. The confidence of each character in the first character string is shown in Table 1 below.
[0112] Table 1
[0113]
[0114] If the threshold is set to 80%, in the above Table 1, the first character in the first character string includes "C" and "4", and the remaining characters are the second characters.
[0115] The arrangement order of the first character and the second character in the first character string may be any order, for example, the first character and the second character may be spaced in sequence, or a plurality of second characters may be spaced apart from a first character, or a plurality of first characters may be spaced apart from a second character. The present application does not limit the arrangement order of the first character and the second character in the first character string.
[0116] Figure 3 is a schematic diagram of a first character string provided in an embodiment of the present application, wherein the first character string includes a first character and a second character. Figure 3 In the diagram, the first character is marked with a dotted line and the second character is marked with a solid line. In the string "3A75C24", "3", "A", "7", "5", and "2" are the second characters, and "C" and "4" are the first characters; in the string "352791", "3", "5", "9", and "1" are the second characters, and "2" and "7" are the first characters; in the string "672CD8Y0", "6", "7", "2", "Y", and "0" are the second characters, and "C", "D", and "8" are the first characters.
[0117] In a possible implementation, multiple first character strings may be recognized for the same voice information to be recognized.
[0118] As an example, when recognizing the pronunciation of the third character in the voice information, it is recognized as multiple characters, such as "3" (confidence 0.35), "5" (confidence 0.4) or "8" (confidence 0.25). Accordingly, the "3", "5" or "8" can be combined with other recognized characters to form a first character string, and the third characters in the three first character strings are "3", "5" and "8" respectively.
[0119] Accordingly, the process in this embodiment is performed for each first character string to obtain a character string corresponding to each first character string, and then one of the character strings corresponding to each first character string is determined as a target character string.
[0120] 202. Replace the first character with the second character in the first character string to obtain a target character string, wherein the matching degree between the target character string and the voice information to be recognized is greater than the matching degree between the first character string and the voice information to be recognized.
[0121] Among them, since the second character in the first character string has a higher confidence, the recognition system has a higher degree of confidence in it, and replaces the first character according to the second character in the first character string to replace the character with a low confidence in the first character string.
[0122] In a possible implementation, a character required to replace the first character may be determined according to the second character in the first character string, and the first character in the first character string may be replaced with the required character to obtain a target character string.
[0123] Among them, since the characters required for the replacement are determined based on the second character, it combines the second character and refers to the influence of more dimensions relative to the first character, so the determined target character string has a higher degree of matching with the voice information to be recognized.
[0124] As an example, combined with the example in Table 1 above, the characters required for replacement are determined based on "3A75" and "2", which are "4" and "S", and the target character string "3A7542S" is obtained.
[0125] In this embodiment, based on the confidence of each character in the first string, it is determined that the first character in the first string meets the replacement condition, and the confidence of the first character is lower than the set threshold; the first character is replaced by the second character in the first string to obtain a target string, and the degree of matching between the target string and the voice information to be recognized is greater than the degree of matching between the first string and the voice information to be recognized. The second character with a higher confidence in the first string is used to replace the first character with a lower confidence, and a target string with a higher degree of matching with the voice information to be recognized is obtained, thereby optimizing the recognition result and improving the accuracy of voice information recognition.
[0126] Figure 4This is a flowchart of replacing characters in the first character string whose confidences meet the replacement condition to obtain a target character string provided by an embodiment of the present application, which may include steps 401 to 403, and these steps are described in detail below.
[0127] 401. According to the second character in each first character string and the position of the second character, query in a preset database to obtain at least one candidate character string, the characters in the candidate character string are consistent with the second character in the corresponding first character string, and the number of the first character string is at least one;
[0128] The voice information to be recognized is a voice input for a specific type of character string. Accordingly, a preset database can be set for the character string, and the preset database can store possible character string forms of the character string.
[0129] As an example, the voice information to be recognized is the voice information collected when the user inputs the SN code of the electronic device by voice. The SN code is a set of codes composed of letters and numbers, which is used to uniquely identify a product.
[0130] Accordingly, an SN code database may be set for the electronic device SN code, and the SN code database may pre-store SN codes of various models produced by various manufacturers. Accordingly, a candidate character string matching the second character and the position of the second character in the first character string may be searched in the SN code database based on the second character and the position of the second character.
[0131] As an example, the voice information to be recognized is voice information collected when a telephone number is input by voice, and the telephone number is a code consisting of 11 digits.
[0132] Accordingly, a telephone number database may be set up for the telephone number, and the telephone number data may pre-store telephone numbers provided by various operators. Accordingly, a candidate character string matching the second character and the position of the second character in the first character string may be queried in the telephone number database.
[0133] The preset database is queried to see whether there is a candidate character string having the same second character and position as the second character in the first character string.
[0134] In a possible implementation, the number of digits with high confidence in the first character string is fixed as the second characters, the number of digits of these second characters is fixed as the correct recognition value, and candidate character strings matching these second characters are screened out from a preset database.
[0135] Among them, the characters in the candidate string are consistent with the second character in the corresponding first string, that is, some characters in the candidate string are the same as the second character in the first string and the position of the second character is the same, and the characters at the position of the first character in the candidate string and the corresponding first string are different.
[0136] In the query process, the number of characters in the first string can also be used as a query condition, the preset database is first screened according to the number of characters, and then the candidate string that is consistent with the second character and the position of the second character in the first string is searched.
[0137] As an example, for the first character string "PF3748D2" obtained by recognizing the voice information to be recognized, the first character is "P" (confidence 0.95), the second character is "F" (confidence 0.92), the third character is: "3" (confidence 0.88), the fourth character is: "7" (confidence 0.93), the fifth character is: "4" (confidence 0.65), the sixth character is: "8" (confidence 0.72), the seventh character is: "D" (confidence 0.75), and the eighth character is: "2" (confidence 0.70). The threshold is set to 0.8, and the confidence of the first four characters ("PF37") is determined to be higher, and they are used as the second character, while the confidence of the last four characters ("48D2") is lower, and they are used as the first character. Based on the first four characters PF37, the candidate character strings are queried in the preset database, which are PF3748D2, PF3749D2, PF3728D2, PF3748A2, and PF3748D2.
[0138] When multiple first character strings are identified for a piece of voice information to be recognized, a corresponding candidate character string may be searched in a preset database for each character string.
[0139] 402. Determine the matching degree between each candidate character string and each first character string in turn;
[0140] After searching the preset database for each first character string to obtain the corresponding candidate character string, the matching degree between each candidate character string and each first character string is determined.
[0141] In a possible implementation, respectively calculating the degree of matching between any candidate string and each first string may include: selecting any first string and calculating the degree of matching between the any first string and the candidate string; similarly calculating the degree of matching between each first string and the candidate string, accumulating the degree of matching between the candidate string and each first string, and obtaining the degree of matching between the any candidate string and each first string.
[0142] As an example, there are three first character strings, and two candidate character strings are determined for each first character string, so a total of six candidate character strings are queried. The matching degrees of the six candidate character strings and the three first character strings are calculated respectively, and the matching degrees of each candidate character string and each first character string are accumulated to obtain the matching degrees of the candidate character string and each first character string.
[0143] 403. Select one of the at least one candidate character string as a first target character string according to the matching degree, wherein the matching degree between the first target character string and each first character string is higher than the matching degree between non-first target character strings in the candidate character string and each first character string.
[0144] According to the value of the matching degree, a candidate character string with the highest matching degree is selected from the multiple candidate character strings as the first target character string.
[0145] The matching degree between the candidate string and each first string represents the overall matching degree between the candidate string and each first string. Therefore, the candidate string with the highest matching degree with each first string is selected, and the first target string is determined to be the best match for the speech information to be recognized from an overall perspective.
[0146] In this embodiment, according to the second character in each first string and the position of the second character, at least one candidate string is obtained by querying in a preset database, the characters in the candidate string are consistent with the second character in the corresponding first string, and the number of the first string is at least one; the matching degree of each candidate string and each first string is determined in turn; according to the matching degree, one is selected from the at least one candidate string as a first target string, and the matching degree of the first target string with each first string is higher than the matching degree of the non-first target string in the candidate string with each first string. According to the confidence level, the characters in the first string are processed in layers, and the second character with high confidence is fixed, and the candidate string matching the second character is screened out in the preset database by using the fixed second character with high confidence, and the one with the highest matching degree with the first string is selected from the subsequent string as the first target string, and the first target string is determined by combining the recognition result and the existing string in the preset database, so as to improve the accuracy of speech information recognition compared with only using the recognition system for recognition.
[0147] Figure 5 This is a flow chart of replacing the first character according to the second character in the first character string to obtain a target character string provided by an embodiment of the present application, which may include steps 501 to 503, and these steps are described in detail below.
[0148] 501. No candidate character string is found in a preset database according to the second character in each first character string and the position of the second character, and the number of the first character string is at least one;
[0149] Wherein, based on the second character in each first character string and the position of the second character, if no candidate character string is found in the preset database, it indicates that there is no candidate character string matching the second character in the first character string in the preset database.
[0150] Then, it is necessary to adopt other methods to replace the first character in the first character string to achieve correction of the first character string.
[0151] In a possible implementation, the recognition system recognizes the voice information to be recognized and obtains multiple first character strings, wherein the first part of the first character strings can be queried in a preset database to obtain candidate character strings, and the second part of the first character strings cannot be queried in the preset database to obtain candidate character strings, and the second part of the first character strings adopts the process of obtaining the target character string in this embodiment.
[0152] In a possible implementation, the recognition system recognizes the speech information to be recognized and obtains multiple first character strings. All the first character strings cannot be queried in the preset database to obtain candidate character strings. All the first character strings are subjected to the process of obtaining the target character string in this embodiment.
[0153] In a possible implementation, a Markov chain model may be used for correction. In this embodiment, the process of using the Markov chain model for correction is described.
[0154] Among them, the Markov chain is a random process, which assumes that given the current state, the future state depends only on the current state and has nothing to do with the past history. This property is called the Markov property.
[0155] The Markov chain model is a statistical method used to capture the transfer rules between elements in a sequence. For strings of the same type, the order of characters is not random, but follows a certain rule. The Markov chain model can use this rule through training to predict the characters that may appear after a character in the string, predict the possibility of low-confidence characters, and do not need to blindly arrange all combinations, thereby improving processing efficiency.
[0156] As an example, for device serial numbers, the order of characters has a certain regularity. For example, a certain character is usually followed by another specific character. This regularity can be learned and utilized through the Markov chain model.
[0157] 502. According to the second character in each first character string, a preset transition probability matrix is searched to obtain at least one target character to which any second character can be transferred, the target character corresponds to the first character in each first character string, the target character is a character sorted after the second character, the preset transition probability matrix includes a number of characters, a transfer character of each character, and a transition probability corresponding to each character and the transfer character;
[0158] Among them, in an electronic device that executes a speech recognition method provided by an embodiment of the present application, a Markov chain model can be used. The transition probability matrix is the core of the Markov chain, which records the transition probability between each character.
[0159] In the present application, the Markov chain model calculates the transition frequencies between different character pairs through historical data to obtain a preset transition probability matrix, in which the probability of each character transferring to another character is recorded.
[0160] Table 2 below is a list of preset transition probability matrices provided in an embodiment of the present application.
[0161] Table 2
[0162]
[0163] Table 2 above lists the characters that may be transferred after each character and the probability of transfer.
[0164] For example, when the current character is "P", the characters that may be transferred are "F", "4", and "G".
[0165] Since the preset transition probability matrix is obtained by counting the order relationship between strings and characters in real data, and the probability of character transition is calculated by counting the order before and after each character (letter, number, etc.) in the same type of string, the preset transition probability matrix can discard the arrangement order that has never appeared, and reduce the calculation amount of predicting the characters that may appear after any character.
[0166] When generating the preset transition probability matrix, the following formula can be used to calculate the transition probability of each character in the string:
[0167] (1)
[0168] Among them, P(a,b) represents the probability of transitioning from state a (that is, character a) to state b (that is, character b); F(a,b) represents the actual number of transitions from state a to state b (frequency); ∑cF(a,c) represents the total number of transitions from state a to all other states (including a itself). It is the sum of the transition frequencies of a to each possible state, representing the total number of times state a appears. In the above formula, a, b, and c represent different characters.
[0169] The numerator F(a,b) in the above formula (1) represents the number of transitions from state a to state b in the historical data. For example, in the process of recognizing the device serial number, if character a is "P" and character b is "F", then F(P,F) represents the number of occurrences of "P" followed by "F" in the historical data.
[0170] The denominator ∑cF(a,c) in the above formula (1) is the total number of times that the state a (character a) can be transferred to all possible states after it appears. For example, the character "P" may be followed by "F", "3", "Q", etc., so the denominator is the total number of times all the characters after "P" appear.
[0171] By dividing the number of a particular transition (such as "P" → "F") by the total number of all possible transitions, we get the transition probability of "P" followed by "F", that is, P(P,F).
[0172] According to the above process, the transition character and the probability of transition of each character in the character string belonging to the same category are determined, and the preset transition probability matrix as shown in Table 2 above is obtained.
[0173] According to the second character in the first character string, a preset transition probability matrix is searched to obtain a target character to which each second character string can be transferred.
[0174] The target character to which the second character can be transferred may be a character with the largest transfer probability corresponding to the second character, or a character with a transfer probability corresponding to the second character greater than a set threshold.
[0175] In a possible implementation, the above-mentioned query process needs to be performed for each first character string to obtain a target character to which each second character in each first character string can be transferred.
[0176] In a possible implementation, when a first character is behind a second character, a target character to which the second character can be transferred is searched in a preset transfer probability matrix.
[0177] In the preset transition probability matrix, when there are multiple characters that may appear after the second character is queried, one or more characters can be selected as target characters to which the second character can be transferred according to the transition probability.
[0178] The selection basis may be that the transition probability is greater than a set transition probability threshold, or may be a set number of transition probabilities that are in front when the transition probabilities are arranged from large to small.
[0179] In one possible implementation, when there are two or more consecutive first characters after a second character, the target character to which the second character can be transferred is queried in the preset transition probability matrix, and then the target character to which the target character can be transferred is queried in the preset transition probability matrix, and so on, to obtain the target characters corresponding to the first characters in the first character string.
[0180] As an example, when the voice information to be recognized is the SN code of an electronic device, a Markov chain is trained for a large number of device SN codes to calculate the transition frequencies between different character pairs, and a preset transition probability matrix for the device SN code is obtained.
[0181] As an example, the first four digits "PF37" of the voice information to be recognized obtained by the recognition system have been correctly recognized, the low confidence bit is the 5th character, and the first four digits are "PF37". In the preset transition probability matrix, the most likely character after "PF37" is analyzed. It is more likely to be "4" rather than other characters after "PF37". "4" is recommended as the 5th character, and it can be inferred that "8" is the most suitable character, and so on.
[0182] As an example, the first four digits "PF37" and the last two digits "50" of the voice information to be recognized obtained by the recognition system have been correctly recognized, the low confidence bit is the 5th character, and the first four digits are "PF37". In the preset transition probability matrix, the most likely characters to appear after "PF37" are analyzed. After "PF37", there may be "4" and "5" instead of other characters. "4" is recommended as the 5th character, and it may be speculated that "8" is the most suitable character. "5" is recommended as the 5th character, and "7" may be recommended as the most suitable character, and so on.
[0183] 503. Determine a second target character string according to the second characters in each first character string and the corresponding target characters.
[0184] After the target character corresponding to the second character is determined, the second character and the corresponding target character are combined into a character string to be selected, and the second target character string is determined from the multiple character strings to be selected.
[0185] As an example, the first four digits "PF37" and the last two digits "50" of the voice information to be recognized obtained by the recognition system have been correctly recognized, the low confidence bit is the 5th character, and the first four digits are "PF37". In the preset transfer probability matrix, the most likely characters after "PF37" are analyzed. After "PF37", there may be "4" and "5" instead of other characters. "4" is recommended as the 5th character, and it may be speculated that "8" is the most suitable character. "5" is recommended as the 5th character, and "7" may be recommended as the most suitable character. The first four digits "PF37", the fifth digit "4", the sixth digit "8" and the last two digits "50" are combined into the candidate character string "PF374850", and the first four digits "PF37", the fifth digit "5", the sixth digit "7" and the last two digits "50" are combined into the candidate character string "PF375750", and two candidate character strings are obtained. The second target character string is determined from the two candidate character strings.
[0186] In this embodiment, according to the second character in each first string and the position of the second character, no candidate string is queried in the preset database, and the number of the first string is at least one; according to the second character in each first string, a preset transition probability matrix is queried to obtain at least one target character that any second character can transfer to, the target character corresponds to the first character in each first string, the target character is a character sorted after the second character, and the preset transition probability matrix contains a number of characters, the transfer characters of each character, and the corresponding transition probabilities of each character and the transfer character; according to the second character in each first string and the corresponding target character, the second target string is determined. The preset transition probability matrix is obtained by processing a large number of strings belonging to the same category by a Markov chain model, and the target characters that may appear after the second character are predicted by using the preset transition probability matrix, and the second string is determined by using the second character in the first string and the corresponding target character. Since the preset transition probability matrix can predict the possibility of the character appearing after any character, when the preset transition probability matrix is used to predict the target character after the second character, it is not necessary to list all possible character combinations, but to directly predict and optimize according to the probability matrix. This method not only reduces the invalid combinations that the system needs to process, effectively reduces the computational complexity, but also dynamically optimizes the recognition process.
[0187] Figure 6 This is a flow chart of determining a second target character string according to the second characters in each first character string and the corresponding target characters provided by an embodiment of the present application, which may include steps 601 to 604, and these steps are described in detail below.
[0188] 601. Combine the target character with a corresponding second character in the first character string to obtain at least two character strings to be selected;
[0189] After the target character that may appear after the second character is determined by using the preset transition probability matrix, the target character and the corresponding second character are combined in order to obtain a character string to be selected. Similarly, the second character corresponding to each first character string and the target character that may appear after it are combined in order to obtain multiple character strings to be selected.
[0190] 602. Query the preset transition probability matrix to obtain the transition probability corresponding to each target character in any character string to be selected;
[0191] The preset transition probability matrix is searched to determine the transition probability of each second character in the candidate character string to the target character.
[0192] As an example, as shown in Table 2, the current character is "P", its target character is "F", and the transition probability is 0.70; the current character is "F", its target character is "3", and the transition probability is 0.68.
[0193] 603. Determine the weight of the arbitrary character string to be selected according to the transition probability corresponding to each target character in the arbitrary character string to be selected;
[0194] The weight of the candidate character string is obtained by calculating the transition probability corresponding to each target character in the same candidate character string.
[0195] In a possible implementation, the transition probabilities corresponding to the target characters in the same candidate character string may be accumulated, and the obtained sum may be used as the weight of the candidate character string.
[0196] As an example, the target characters determined in the candidate character string "PF374850" contain a transition probability of 0.50 for the fifth digit "4" and a transition probability of 0.49 for the sixth digit "8", and the weight of the candidate character string is 0.99.
[0197] In a possible implementation, the transition probabilities corresponding to the target characters in the same candidate character string may be multiplied, and the obtained sum is used as the weight of the candidate character string.
[0198] As an example, the target characters determined in the candidate character string "PF374850" contain a transition probability of 0.50 for the fifth digit "4" and a transition probability of 0.49 for the sixth digit "8", and the weight of the candidate character string is 0.245.
[0199] In a possible implementation, the transition probabilities corresponding to the target characters in the same candidate character string may be accumulated and the average value may be taken, and the obtained average value may be used as the weight of the candidate character string.
[0200] As an example, the target characters determined in the candidate character string "PF374850" contain a transition probability of 0.50 for the fifth digit "4" and a transition probability of 0.49 for the sixth digit "8", and the weight of the candidate character string is 0.495.
[0201] 604. Select a second target character string from among the candidate character strings, where a weight of the second target character string is higher than a weight of non-second target character strings from among the candidate character strings.
[0202] According to the weights of the candidate character strings, a character string with the highest weight is selected from the plurality of candidate character strings as a second target character string, and the second target character string is a character string output as a recognition result.
[0203] In a possible implementation, all possible character strings are stored in a preset database, and it is necessary to determine whether the finally selected second target character string exists in the preset database.
[0204] In this embodiment, each target character is first combined with the corresponding second character in the first character string to obtain at least two candidate character strings; then the preset transition probability matrix is queried to obtain the transition probability corresponding to each target character in any candidate character string; the weight of the arbitrary candidate character string is determined according to the transition probability corresponding to each target character in the arbitrary candidate character string; a second target character string is selected from each candidate character string, and the weight of the second target character string is higher than the weight of the non-second target character string in each candidate character string, so as to achieve the transition probability of the second character recorded in the preset transition probability matrix to transfer the target character, and determine the weight of each candidate character string, so as to select one as the second target character string from multiple candidate character strings by using the weight, so as to ensure that the target character string finally selected has a high degree of matching with the voice information to be recognized, thereby optimizing the recognition result and improving the accuracy of voice information recognition.
[0205] Figure 7 This is a schematic diagram of a flow chart of selecting a second target character string from among candidate character strings provided by an embodiment of the present application, which may include steps 701 to 703, and these steps are described in detail below.
[0206] 701. Sort the candidate character strings according to their weights;
[0207] Before querying whether there are any character strings to be selected from the preset database, the character strings to be selected are sorted in advance.
[0208] In a possible implementation, the candidate character strings may be sorted in descending order of weight.
[0209] In a possible implementation, the candidate character strings may be sorted in order of weight from small to large.
[0210] 702. According to the ranking of each candidate character string, query whether the candidate character string exists from a preset database;
[0211] Among them, each character string to be selected is queried from the setting database according to the order to see whether it exists.
[0212] As an example, 10 character strings to be selected are obtained, and a preset database is queried to determine whether each of the 10 character strings to be selected exists.
[0213] In a possible implementation, some of the character strings to be selected may exist in the preset database, while others may not; all of the character strings to be selected may exist in the preset database; or none of the character strings to be selected may exist in the preset database.
[0214] 703. Query a preset database for an existing target character string to be selected as a second target character string, wherein the weight of the target character string to be selected is higher than the weight of the non-target character string to be selected.
[0215] In order to reduce the number of queries, each candidate character string may be queried in the preset database in order from largest to smallest to see whether it exists.
[0216] The target candidate character string is a candidate character string found in a preset database, and its weight is smaller than the target candidate character string.
[0217] When any candidate character string is found in the preset database, the search is stopped and the candidate character string obtained by the search is used as the second target character string.
[0218] Among them, the candidate character string that is not found in the preset database is not considered as the second target character string.
[0219] In a possible implementation, the overall confidence of each candidate character string may also be determined by accumulating the matching score of each character for calculation. The matching score is affected by whether the character exists in a preset database and the transition probability of the character.
[0220] When determining the confidence of the candidate character string, the following formula can be used for calculation:
[0221] (2)
[0222] Among them, S represents the overall confidence of the selected string, and W i Indicates the weighted value of the preset database comparison result. A successful match (found in the preset database) is valued at 1, and a failed match (not found in the preset database) is valued at 0. i) represents the transition probability of the character, i represents the i-th character in the candidate string, and n represents that the candidate string contains n characters.
[0223] The above formula (2) indicates that the candidate character string is found in the preset database, and the confidence of the candidate character string is obtained by accumulating the transition probability of each target character in the candidate character string.
[0224] In a possible implementation, when none of the candidate character strings can be found in the preset database through the above recognition process, it is determined that the recognition of the voice information to be recognized has failed, and a prompt message can be generated to prompt the user to input the voice again and continue to use the above recognition process for recognition.
[0225] In this embodiment, the candidate strings are sorted according to their weights; based on the sorting of the candidate strings, a query is made from a preset database as to whether the candidate string exists; the target candidate string that exists in the preset database is queried as the second target string, and the weight of the target candidate string is higher than the weight of the non-target candidate string. This realizes the combination with the preset database, the screening of the candidate strings determined by using the preset transition probability matrix, the determination of the second target string from multiple dimensions, and the improvement of the accuracy of speech information recognition.
[0226] In a possible implementation, before obtaining the voice information to be recognized, the method further includes:
[0227] Obtain device information provided by the user terminal.
[0228] The voice information to be recognized may be audio information acquired by the terminal after the user utters the voice at the terminal.
[0229] In an application scenario, the user also uploads device information through the terminal, and the device information may be information such as the model and manufacturer of the device.
[0230] The device information may be input by voice or manually input through an interface provided by the terminal.
[0231] Figure 8 801 and a server 802. The dashed line in the figure indicates the data connection between the terminal and the server. A speech recognition method provided by the present application is applied to the server. The user inputs speech information and device information through the terminal, and the terminal sends the collected speech information and device information to the server. The server performs recognition processing on the speech information to obtain a target string, and performs subsequent processing in combination with the target string.
[0232] Fig. 9It is a flowchart of outputting the target character string as a recognition result provided by an embodiment of the present application, which may include steps 901 to 902. These steps are described in detail below.
[0233] 901. Predicting a service process corresponding to the voice information to be recognized based on matching the target character string with the device information;
[0234] After the target character string is determined, the relevant information of the terminal device can be determined.
[0235] The relevant information may be various service-related information of the device corresponding to the target character string obtained by querying the service database using the target character string.
[0236] In order to improve information security, the device information is also compared with the target string to determine whether the two match. If they match, the service process corresponding to the voice information to be recognized is predicted.
[0237] Among them, determining whether the device information matches the target string may include: querying the device information corresponding to the target string in the device-related database, judging whether the device information corresponding to the target string is consistent with the received device information, and if they are consistent, determining that the two match; otherwise, the two do not match.
[0238] As an example, the device corresponding to the target string queried in the device-related database is a tablet computer, and the received device information is a mobile phone. The target string and the received device information are inconsistent.
[0239] As an example, the device corresponding to the target string queried in the device-related database is a mobile phone, the received device information is a mobile phone, and the target string is consistent with the received device information.
[0240] When the target character string matches the device information, it can be determined that the user who uses the terminal for voice input is a legitimate user, and for the device corresponding to the target character string, the service process that the user may use is predicted.
[0241] As an example, the target string can be used to determine the corresponding device information, and the corresponding device information can be used to determine the service-related information of the device, such as whether the device is under warranty (also known as within the warranty period). If so, jump to the corresponding service process; if not, provide another service process.
[0242] 902. Output the recognition result to a service terminal using the service process, and the service terminal is used to provide services to each user terminal.
[0243] The predicted service process is used to control the service terminal to enter or prepare to enter the corresponding service. The service process is used to output the target character string as a recognition result to the service terminal.
[0244] Among them, the service terminal can also be called a seat, which can encrypt the target string and the corresponding device information. After the service corresponding to the seat is connected, the target string and the corresponding device information are transmitted to the service terminal through an encrypted channel between the electronic device and the seat to ensure the security of the information.
[0245] Moreover, the target character string obtained by the recognition is not fed back to the user terminal. Even if an illegal user enters accurate device information, it matches the target character string obtained by the voice information to be recognized, and the security of the information can be guaranteed.
[0246] As an example, if the device is within the warranty period and it is predicted that it may be renewed, the renewal process is used to output the recognition result to the service terminal. The renewal process option pops up in the service terminal, and the device SN code part of the renewal process is filled in with the target string in the recognition result, so that the staff of the service terminal can quickly respond to the user's renewal request later.
[0247] Moreover, since the target string is directly output to the service process in this application, even if someone uses a false SN code to try to correct it in the system, sensitive information cannot be obtained. Even if the false SN code is corrected, it will only jump in the voice process of the hotline, and the user information will only be displayed to the agent of the terminal call system for further customer service, and no sensitive information will be fed back to the user through the system.
[0248] In one possible implementation, in order to enhance the traceability of the system, real-time monitoring and logging functions are also added, and all failed attempts are recorded and auditable. Logs can help technology respond quickly and take corresponding measures when security incidents occur.
[0249] In this embodiment, the device information provided by the user terminal is obtained; accordingly, the target character string is output as the recognition result, including: based on the matching of the target character string with the device information, the service process corresponding to the voice information to be recognized is predicted; and the recognition result is output to the service terminal using the service process, and the service terminal is used to provide services for each user terminal. The device information provided by the user terminal is compared with the target character string to determine whether the user who inputs the voice is legal, and if it is legal, the service process corresponding to the voice information to be recognized is predicted, and the recognition result is directly output to the service terminal using the service process, so that the staff of the service terminal can quickly respond to the service requests of subsequent users and improve service efficiency.
[0250] In a possible implementation, obtaining the voice information to be recognized includes: obtaining the voice information to be recognized in real time based on responding to a call request.
[0251] The voice information to be recognized is input by the user during a voice call at the terminal. Accordingly, the electronic device as a server responds to the call request and obtains the voice during the call in real time as the voice information to be recognized.
[0252] During a call, the user can input multiple voice messages to be recognized for the device.
[0253] Correspondingly, each time the input voice information to be recognized is processed according to the voice recognition method in the above embodiment to obtain the target character string.
[0254] Fig.10 This is another flowchart of outputting the target character string as a recognition result provided by an embodiment of the present application, which may include steps 1001 to 1003, and these steps are described in detail below.
[0255] 1001. Count the first times of obtaining the target string in this call;
[0256] This is to prevent illegal users from trying to obtain the target string through multiple attempts.
[0257] In this embodiment, during the same call, the voice information to be recognized is recognized, and the number of times the target character string is obtained is counted.
[0258] The first number of times the target character string is obtained in this call is counted, and the first number is the number of times the user tries to input the character string in this call.
[0259] The first number of times the target character string is obtained in a call not only includes the processing of the voice information to be recognized for a certain device, but also includes all the times the target character string is obtained in this call.
[0260] Under normal circumstances, when a user operates the service of a device (such as a phone) through the phone, there will not be a process of operating multiple devices at the same time. Moreover, when inputting a string of characters to a device by voice, if an input error occurs, the user will re-enter the character string more carefully, and generally can enter the correct character string in fewer attempts. Therefore, a first count threshold is set to prevent illegal users from conducting trial and error attacks and making multiple attempts to steal the service operation process of the device.
[0261] 1002. Based on the first number being less than the first number threshold, output the target character string as a recognition result;
[0262] 1003. Based on the fact that the first number is not less than the first number threshold, it is prohibited to output the target character string as a recognition result.
[0263] If the first number is less than the first number threshold, the user can be considered to be a legitimate user, and the target character string is output as a recognition result for subsequent services.
[0264] If the first number is not less than the first number threshold, the user may be considered an illegal user who steals the service operation of the device through multiple inputs, and it is prohibited to output the target character string as a recognition result.
[0265] The first number threshold may be a smaller value, such as 2 or 3, to prevent trial-and-error attacks.
[0266] In a possible implementation, through the above-mentioned recognition process, when the target character string cannot be queried, it can be determined that the recognition of the voice information to be recognized has failed, and the number of failures can be counted. When the number of failures is greater than the set value, an alarm is triggered and protective measures are taken to prevent large-scale attacks.
[0267] Among them, the first number threshold is used to distinguish whether there is a user using a false SN code to obtain more user information through the identification method in this application.
[0268] In this embodiment, based on responding to a call request, voice information to be recognized is obtained in real time; accordingly, the target character string is output as a recognition result, including: counting the first number of times the target character string is obtained in this call; based on the first number being less than a first number threshold, the target character string is output as a recognition result; otherwise, it is prohibited to output the target character string as a recognition result to prevent users from conducting trial and error attacks and improve information security.
[0269] Fig.11 It is a flow chart of the speech recognition method provided in an embodiment of the present application, which may include steps 1101 to 1103, and these steps are described in detail below.
[0270] 1101. Obtain relevant information of a user terminal;
[0271] The relevant information of the user terminal may include the device number, address, telephone number, etc. of the terminal used by the user.
[0272] In a possible implementation, the voice information to be recognized is obtained by making a wireless call via a telephone, and the telephone number of the user terminal can be obtained together.
[0273] In a possible implementation, the voice information to be recognized is obtained by making a call over the network, and information such as the IP address (Internet Protocol Address) of the user terminal can be obtained together.
[0274] 1102. Based on the relevant information of the user terminal, count the second number of times that the user terminal obtains the target character string during the call within a preset time period;
[0275] The preset duration may be a longer period of time, such as 1 month, 1 week, or 1 day, etc., and may be set based on actual conditions.
[0276] In normal usage scenarios, a user will not frequently input the device SN code through the terminal. If frequent input occurs, it can be determined that it is a trial-and-error attack.
[0277] Based on the relevant information of the user terminal, a second number of times the terminal obtains the target character string through voice input during a call within the preset time period is counted.
[0278] In a possible implementation, through the above recognition process, when the target character string cannot be queried, it is determined that the recognition of the voice information to be recognized has failed, and the number of failures is counted. When the number of failures is greater than the set value, an alarm is triggered and protective measures are taken to prevent large-scale attacks.
[0279] 1103. Based on the fact that the second number is not less than the second number threshold, and according to the relevant information of the target terminal, prohibiting responding to the voice information to be recognized provided by the target terminal.
[0280] If the second number is not less than the second number threshold, it can be determined that a trial and error attack process is being carried out.
[0281] In one possible implementation, it is prohibited to respond to the voice information to be recognized provided by the target terminal to protect data security, reduce the data processing burden of the electronic device that executes the voice recognition method provided in this application, and ensure the safe operation of the electronic device.
[0282] In one possible implementation, starting from the current moment, for a relatively long period of time, it is prohibited to respond to the voice information to be recognized provided by the target terminal to protect data security and ensure the operational safety of the electronic device that executes the voice recognition method provided in this application.
[0283] In one possible implementation, the Fig.11 Each step in Figure 1 The steps in are executed in parallel.
[0284] In this embodiment, relevant information of the user terminal is obtained; based on the relevant information of the user terminal, the second number of times the user terminal is involved in obtaining the target character string in the call within the preset time length is counted; based on the second number being not less than the second number threshold, according to the relevant information of the target terminal, it is prohibited to respond to the voice information to be recognized provided by the target terminal. Using the relevant information of the user terminal, if the number of times the user terminal is involved in obtaining the target character string in the call within the preset time length is large and does not belong to normal use, it is prohibited to respond to the voice information to be recognized provided by the user terminal, so as to ensure the data security of the user terminal and the operation security of the electronic device when a trial and error attack occurs.
[0285] A speech recognition method provided in an embodiment of the present application is introduced above, and a device for executing the above speech recognition method will be introduced below.
[0286] See also Fig.12 , Fig.12 Schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application. Fig.12 As shown, the speech recognition device 1200 includes:
[0287] Interface 1201, used to obtain voice information to be recognized;
[0288] The processor 1202 is configured to recognize the speech information to be recognized to obtain a first character string; replace characters in the first character string whose confidences meet a replacement condition to obtain a target character string; and output the target character string as a recognition result.
[0289] In one possible implementation, the processor includes:
[0290] A recognition module, used for recognizing the voice information to be recognized to obtain a first character string;
[0291] A replacement module, used for replacing characters in the first character string whose confidences meet the replacement condition, to obtain a target character string;
[0292] The output module is used to output the target character string as a recognition result.
[0293] Wherein, in the case where the speech recognition device has an audio acquisition device, the interface is used to connect the processor and the audio acquisition device, and send the audio information acquired by the audio acquisition device to the processor.
[0294] Wherein, when the speech recognition device is connected to a user terminal, the user terminal collects speech information to be recognized, the user terminal is connected to the speech recognition device through the interface, and the interface sends the acquired audio information to the processor.
[0295] In a possible implementation, the replacement module includes:
[0296] A first determining unit, configured to determine, based on the confidences of the characters in the first character string, that a first character in the first character string satisfies a replacement condition, and that the confidence of the first character is lower than a set threshold;
[0297] The replacement unit is used to replace the first character according to the second character in the first character string to obtain a target character string, wherein the matching degree between the target character string and the voice information to be recognized is greater than the matching degree between the first character string and the voice information to be recognized.
[0298] In a possible implementation, the replacement unit is specifically configured to:
[0299] According to the second character in each first character string and the position of the second character, at least one candidate character string is obtained by searching in a preset database, wherein the characters in the candidate character string are consistent with the second character in the corresponding first character string, and the number of the first character string is at least one;
[0300] Determining the matching degree between each candidate character string and each first character string in turn;
[0301] According to the matching degree, one of the at least one candidate character string is selected as a first target character string, and the matching degree between the first target character string and each first character string is higher than the matching degree between non-first target character strings in the candidate character string and each first character string.
[0302] In a possible implementation, the replacement unit is specifically configured to:
[0303] According to the second character in each first character string and the position of the second character, no candidate character string is found in the preset database, and the number of the first character string is at least one;
[0304] According to the second character in each first character string, a preset transition probability matrix is searched to obtain at least one target character to which any second character can be transferred, the target character corresponds to the first character in each first character string, the target character is a character sorted after the second character, the preset transition probability matrix includes a number of characters, a transfer character of each character, and a transition probability corresponding to each character and the transfer character;
[0305] A second target character string is determined according to the second characters in each first character string and the corresponding target characters.
[0306] In a possible implementation, the replacement unit determines the second target character string according to the second characters in each first character string and the corresponding target characters, specifically including:
[0307] Combining the target character with the corresponding second character in the first character string to obtain at least two character strings to be selected;
[0308] Query the preset transition probability matrix to obtain the transition probability corresponding to each target character in any candidate character string;
[0309] Determine the weight of the arbitrary character string to be selected according to the transition probability corresponding to each target character in the arbitrary character string to be selected;
[0310] A second target character string is selected from each candidate character string, and a weight of the second target character string is higher than a weight of a non-second target character string from each candidate character string.
[0311] In a possible implementation, the replacement unit selects the second target character string from each candidate character string, specifically including:
[0312] Sort the candidate strings according to their weights;
[0313] According to the ranking of each candidate character string, query whether the candidate character string exists in a preset database;
[0314] An existing target character string to be selected is queried from a preset database as a second target character string, and a weight of the target character string to be selected is higher than a weight of a non-target character string to be selected.
[0315] In a possible implementation, the method further includes:
[0316] An acquisition module, used to obtain device information provided by a user terminal before obtaining the voice information to be recognized;
[0317] Accordingly, the output module includes:
[0318] A prediction unit, configured to predict a service flow corresponding to the voice information to be recognized based on matching the target character string with the device information;
[0319] The first output unit is used to output the recognition result to the service terminal using the service process, and the service terminal is used to provide services for each user terminal.
[0320] In a possible implementation, the interface is specifically used to:
[0321] Based on the response to the call request, the voice information to be recognized is obtained in real time;
[0322] Accordingly, the output module includes:
[0323] A counting unit, used to count the first number of times the target character string is obtained in this call;
[0324] A second output unit, configured to output the target character string as a recognition result based on the first number being less than a first number threshold;
[0325] The prohibition unit is used to prohibit the target character string from being output as a recognition result based on the first number being not less than a first number threshold.
[0326] In a possible implementation, the method further includes:
[0327] Obtain relevant information of the user terminal;
[0328] Based on the relevant information of the user terminal, counting the second number of times the user terminal obtains the target character string during the call within a preset time period;
[0329] Based on the fact that the second number is not less than the second number threshold, and according to the relevant information of the target terminal, it is prohibited to respond to the voice information to be recognized provided by the target terminal.
[0330] It should be noted that for the functional explanation of each component structure in a speech recognition device provided in an embodiment of the present application, please refer to the explanation in the aforementioned method embodiment, and no further explanation will be given here.
[0331] In this embodiment, after the interface obtains the voice information to be recognized, the processor recognizes the voice information to be recognized and obtains a first character string; replaces the characters in the first character string whose confidence meets the replacement condition to obtain a target character string; and outputs the target character string as a recognition result. In this process, the recognition result is optimized by replacing the characters in the first character string whose confidence meets the replacement condition, thereby obtaining a target character string with higher accuracy and improving the accuracy of voice information recognition.
[0332] The present application also provides an electronic device in an embodiment. Fig.13 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the speech recognition method in the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Fig.13 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0333] like Fig.13As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage device 1308 to a random access memory (RAM) 1303. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 1303. The processing device 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0334] Typically, the following devices may be connected to the I / O interface 1305: an input device 1306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1308 including, for example, a memory card, a hard disk, etc.; and a communication device 1309. The communication device 1309 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Fig.13 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0335] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any of the speech recognition methods provided in the embodiments of the present application.
[0336] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any speech recognition method provided in the embodiment of the present application.
[0337] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0338] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0339] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0340] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. A speech recognition method, comprising: Obtaining voice information to be recognized; Recognize the voice information to be recognized to obtain a first character string; Replacing characters in the first character string whose confidences meet a replacement condition to obtain a target character string; The target character string is output as a recognition result.
2. The speech recognition method according to claim 1, wherein replacing characters in the first character string whose confidences meet a replacement condition to obtain a target character string comprises: Determining, based on the confidences of the characters in the first character string, that the first character in the first character string meets a replacement condition, and that the confidence of the first character is lower than a set threshold; The first character is replaced according to the second character in the first character string to obtain a target character string, wherein the matching degree between the target character string and the voice information to be recognized is greater than the matching degree between the first character string and the voice information to be recognized.
3. The speech recognition method according to claim 2, wherein the step of replacing the first character with the second character in the first character string to obtain the target character string comprises: According to the second character in each first character string and the position of the second character, at least one candidate character string is obtained by searching in a preset database, wherein the characters in the candidate character string are consistent with the second character in the corresponding first character string, and the number of the first character string is at least one; Determining the matching degree between each candidate character string and each first character string in turn; According to the matching degree, one of the at least one candidate character string is selected as a first target character string, and the matching degree between the first target character string and each first character string is higher than the matching degree between non-first target character strings in the candidate character string and each first character string.
4. The speech recognition method according to claim 2, wherein the step of replacing the first character with the second character in the first character string to obtain a target character string comprises: According to the second character in each first character string and the position of the second character, no candidate character string is found in the preset database, and the number of the first character string is at least one; According to the second character in each first character string, a preset transition probability matrix is searched to obtain at least one target character to which any second character can be transferred, wherein the target character corresponds to the first character in each first character string, and the target character is a character sorted after the second character, and the preset transition probability matrix includes a number of characters, a transfer character of each character, and a transition probability corresponding to each character and the transfer character; A second target character string is determined according to the second characters in each first character string and the corresponding target characters.
5. The speech recognition method according to claim 4, wherein determining the second target character string according to the second character in each first character string and the corresponding target character comprises: Combining the target character with a corresponding second character in the first character string to obtain at least two character strings to be selected; Query the preset transition probability matrix to obtain the transition probability corresponding to each target character in any candidate character string; Determining the weight of the arbitrary character string to be selected according to the transition probability corresponding to each target character in the arbitrary character string to be selected; A second target character string is selected from each candidate character string, wherein the weight of the second target character string is higher than the weight of non-second target character strings from each candidate character string.
6. The speech recognition method according to claim 5, wherein the step of selecting the second target character string from the candidate character strings comprises: Sort the candidate strings according to their weights; According to the ranking of each candidate character string, querying from a preset database whether the candidate character string exists; An existing target character string to be selected is queried from a preset database as a second target character string, and the weight of the target character string to be selected is higher than the weight of the non-target character string to be selected.
7. The speech recognition method according to any one of claims 1 to 6, before obtaining the speech information to be recognized, further comprising: Obtaining device information provided by a user terminal; Accordingly, outputting the target character string as a recognition result includes: Based on the matching of the target character string and the device information, predicting the service process corresponding to the voice information to be recognized; The service process is used to output the recognition result to a service terminal, and the service terminal is used to provide services to each user terminal.
8. The speech recognition method according to any one of claims 1 to 6, wherein obtaining the speech information to be recognized comprises: Based on the response to the call request, the voice information to be recognized is obtained in real time; Accordingly, outputting the target character string as a recognition result includes: Count the first times the target string is obtained in this call; Based on the first number of times being less than a first number threshold, outputting the target character string as a recognition result; Based on the fact that the first number is not less than a first number threshold, outputting the target character string as a recognition result is prohibited.
9. The speech recognition method according to claim 8, further comprising: Obtain relevant information of the user terminal; Based on the relevant information of the user terminal, counting a second number of times the user terminal obtains the target character string during the call within a preset time period; Based on the second number being not less than a second number threshold, and according to the relevant information of the target terminal, it is prohibited to respond to the voice information to be recognized provided by the target terminal.
10. A sound recognition device, comprising: Interface, used to obtain the voice information to be recognized; A processor, configured to identify the voice information to be recognized and obtain a first character string; Replacing characters in the first character string whose confidences meet a replacement condition to obtain a target character string; The target character string is output as a recognition result.
Citation Information
Patent Citations
Method, device, server and terminal for correcting license plate character string
CN108257602A
Building service facility control method and system based on intelligent semantic instruction recognition
CN109949803A
Server supporting device to perform speech recognition and method of operating server
CN114223029A
Speech recognition method and display equipment
CN115588431A
Data correction method, system and equipment and storage medium
CN119629407A