Speech recognition method, electronic equipment and computer readable storage medium

By first filtering the acoustic probability score during the speech recognition process, and then combining the language probability score to reduce the number of candidate characters, the problem of slow speech recognition speed of electronic devices is solved and the user experience is improved.

CN120340499AActive Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202410045802.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-18
Estimated Expiration
2044-01-10

AI Technical Summary

Technical Problem

Electronic devices are too slow to recognize voice content, which affects the user experience.

Method used

In the speech recognition process, the candidate characters are first filtered based on the acoustic probability score, and then filtered multiple times with the language probability score, reducing the number of candidate characters and reducing the amount of calculation.

Benefits of technology

It improves the speed of voice recognition, reduces user waiting time, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340499A_ABST
    Figure CN120340499A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method, electronic equipment and a computer readable storage medium. In the method, when the electronic equipment identifies a character corresponding to each pronunciation, a plurality of candidate characters can be determined on the basis of a prediction result of an acoustic model, and then multiple times of screening are carried out on the basis of the pronunciation of the plurality of candidate characters and an acoustic probability score. Finally, the electronic device can further screen the finally determined candidate characters based on a language model to determine a final recognition result. According to the method, the calculation amount of the electronic equipment can be reduced while the speech recognition precision is ensured, so that the speech recognition speed of the electronic equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of terminals and communication technologies, and in particular, to a voice recognition method, an electronic device, and a computer-readable storage medium. Background Art

[0002] An electronic device usually needs to recognize the voice content included in the audio while recording the audio, and then provide one or more functions for the user based on the recognized voice content. For example, the electronic device can display the voice content in a visual form such as text or image for the user to view and edit after recognizing the voice content, or the electronic device can also have a conversation with the user based on the voice content, etc. Sometimes, if the speed of the electronic device for recognizing the voice content is too slow, it will reduce the speed of providing subsequent functions, thereby affecting the user experience. Summary of the Invention

[0003] This application provides a voice recognition method, an electronic device, and a computer-readable storage medium. Among them, the electronic device can first determine multiple candidate characters corresponding to the frame of audio based on a single-frame audio and the recognized character sequence, and then the electronic device can perform multiple screenings on the candidate characters based on the acoustic probability scores and language probability scores of the multiple candidate characters, and finally determine the character corresponding to the frame of audio from the screened candidate characters. After multiple screenings, the electronic device reduces the number of candidate characters and the computational complexity during voice recognition, thus reducing the user's waiting time and improving the user experience when using voice recognition-related functions.

[0004] In a first aspect, the present application provides a speech recognition method, which is applied to an electronic device. The method includes: the electronic device receives a first audio segment, the first audio segment includes the K-th frame audio and the N frames of audio before the K-th frame audio, where K is a positive integer and N is a positive integer; the electronic device determines multiple character sequences based on the first prediction result of the K-th frame audio and the first character sequence corresponding to the previously recognized N frames of audio. Any one of the multiple character sequences contains a candidate character corresponding to the first character sequence and the K-th frame audio. The first prediction result includes multiple candidate characters of the K-th frame audio; the electronic device screens out L1 character sequences from the multiple character sequences. The acoustic probability scores of the L1 character sequences are the highest among the multiple character sequences. The acoustic probability score of a character sequence represents the probability that the speech content of the K-th frame audio and the previous N frames of audio is a character sequence, and L1 is a positive integer; the electronic device screens out L2 character sequences from the multiple character sequences according to the L1 character sequences. The pronunciation of the candidate character of the K-th frame audio in the L2 character sequences satisfies the first condition with the pronunciation of the candidate character of the K-th frame audio in any one of the L1 character sequences, and L2 is a positive integer; the electronic device screens out L3 character sequences from the L2 character sequences. The acoustic probability scores of the L3 character sequences satisfy the second condition, and L3 is a positive integer; the electronic device selects a second character sequence from the L3 character sequences based on the second prediction result of the K-th frame audio and determines the second character sequence as the speech recognition result of the K-th frame audio and the previous N frames of audio. The second prediction result is obtained by the electronic device based on the semantic prediction of the speech content of the K-th frame audio.

[0005] In combination with the first aspect, in some embodiments, the first condition includes any one of the following: the finals of the two pronunciations are the same, the two pronunciations only differ in tone, and the two pronunciations are front nasal and back nasal.

[0006] Among them, the first audio segment may be a valid audio segment, which corresponds to a complete sentence. The N frames of audio before the K-th frame audio in the first audio segment are audio containing pronunciations, and each pronunciation may correspond to a character. The electronic device may determine the first prediction result based on the first character sequence, where the first prediction result may include a probability distribution matrix of the characters corresponding to the K-th frame audio output by the acoustic model. Among them, the character sequence A may include multiple character sequences. Therefore, for each character sequence, an electronic device can obtain a probability distribution matrix. A probability distribution matrix represents the acoustic probability scores of multiple candidate characters under a character sequence. Figure 4Taking the illustrated embodiment as an example, the characters in the (K-1)-th frame include "chew" and "pay". The character sequence A may include the character sequence corresponding to the audio before and including the (K-1)-th frame. That is to say, the character sequence A includes both the character sequence composed of the recognized characters under the branch of "chew" and the character sequence composed of the recognized characters under the branch of "pay". The multiple character sequences determined by the electronic device based on the first prediction result include the recognized character sequence and the candidate characters of the audio in the K-th frame under this character sequence, so as to Figure 4Taking the illustrated embodiment as an example, the recognized character sequence includes "chew" and "pay", wherein the candidate characters of the K-th frame audio under the character sequence where "chew" is located include: "mouth", "pull", "kou", "kou", etc., and the candidate characters of the K-th frame audio under the character sequence where "pay" is located include: "kou", "gain", "mouth", "knock", etc. In this way, the multiple character sequences determined by the electronic device include: "chew mouth", "chew pull", "chew kou", "chew kou", "chew kou", "pay kou", "seize", "pay mouth", "pay kou", etc. Then the electronic device can filter out L1 character sequences with the highest acoustic probability scores from multiple character sequences, that is, filter out the characters with the highest acoustic probability scores of the characters corresponding to the K-th frame audio under a certain character sequence A. For example, the acoustic probability scores of "chew mouth", "chew pull", "pay kou", "seize" are filtered out from multiple character sequences. Then the electronic device can determine the pronunciation of the characters corresponding to the K-th frame audio in the filtered character sequence, and then filter out characters with the same or similar pronunciation as the above-mentioned characters corresponding to the K-th frame audio from multiple character sequences to form L2 new character sequences. For example, the electronic device selects the pronunciations that are the same or similar to "口" and "抠" from all the candidate characters of the Kth frame to form a new character sequence, so that the candidate character sequence is expanded to "咬口", "咬抠", "咬蔻", "咬扣" and so on. Here, the L2 character sequences include both the above-mentioned L1 character sequences and the character sequences expanded based on the L1 character sequences. Then the electronic device can select L3 character sequences with the highest acoustic probability scores from the L2 character sequences, and finally the electronic device can select the second character sequence from the L3 character sequences based on the second prediction result of the language model. Among them, the second preset condition can be that the character sequence belongs to one of the character sequences with the highest acoustic probability score in the L2 character sequences, and the highest acoustic probability score here can refer to the first Q1 acoustic probability scores, Q1 is a positive integer, and Q1 can be 1, 2, 3, etc. The second prediction result can include the probability distribution matrix output by the language model, and the probability distribution matrix contains the language probability scores of the candidate characters corresponding to the Kth frame audio under each character sequence A. Among them, the electronic device will search the language probability score of the character corresponding to the Kth frame audio in the L3 character sequence from the second prediction result, and then calculate the total score based on the acoustic probability score and language probability score of the L3 character sequence. Finally, the electronic device can determine the second character sequence based on the total score, and use the second character sequence as the speech recognition result of the Kth frame audio and the N frames of audio before the Kth frame audio. The second character sequence can be the character sequence with the highest total score among the L3 character sequences. The value of L1 can be smaller, and the value of L3 can be larger than L1. For example, L1 can be 3, 4 or 5, etc., and L3 can be 8, 9 or 10, etc.

[0007] That is to say, the electronic device can first determine the L1 character sequences with the highest acoustic probability scores from all character sequences, then filter out L2 character sequences from all character sequences based on the pronunciation of the Kth frame of audio in the L1 character sequences, and finally filter out the character sequence with the highest acoustic probability score from the L2 character sequences. Compared with determining the final recognition result from a large number of character sequences, the electronic device can reduce the amount of calculation and improve the speed of speech recognition by using the above method.

[0008] In combination with the first aspect, in some embodiments, the electronic device selects L3 character sequences from L2 character sequences, specifically including: the electronic device selects L4 character sequences from the L2 character sequences, the pronunciation of the character corresponding to the Kth frame audio in the L4 character sequences has the highest frequency of occurrence, and L4 is a positive integer; the electronic device selects L3 character sequences from the L4 character sequences, and the L4 character sequence has the highest acoustic probability score among the L3 character sequences.

[0009] refer to Figure 4 In the embodiment shown, after the electronic device selects characters with the same or similar pronunciations as the selected candidate characters as candidate characters, it can first select the characters with the highest pronunciation frequency from the candidate characters. In the branch with the K-1 frame being "交", the candidate characters such as "扣", "获", "口", and "敲" are screened, and "获" is removed, among which "获" is a character with a lower pronunciation frequency. The highest pronunciation frequency may refer to the number of times the pronunciation appears in the top Q2 of all candidate characters. Then the electronic device selects the candidate characters with the highest acoustic probability scores from the characters with the highest pronunciation frequency, and obtains L3 character sequences, including: "交扣", "交口", and "交敲".

[0010] That is, after the electronic device expands the candidate characters of the Kth frame based on the pronunciation of the existing candidate characters, it can filter out some characters that are pronounced less frequently. In this way, the electronic device can remove some candidate characters introduced by the prediction network that are irrelevant to the pronunciation in the Kth frame audio (such as Figure 4 In the embodiment shown, the "obtain" can ensure that the candidate characters of the K-th frame audio to be recognized determined by the acoustic model are more dependent on the audio features of the K-th frame audio, which not only reduces the amount of calculation but also ensures the final recognition accuracy.

[0011] In combination with the first aspect, in some embodiments, the first character sequence includes a third character sequence and a fourth character sequence, and any one of the multiple character sequences contains a candidate character corresponding to the Kth frame of audio from the third character sequence, or contains a candidate character corresponding to the Kth frame of audio from the fourth character sequence.

[0012] That is to say, the first character sequence can include multiple character sequences. Figure 4Taking the illustrated embodiment as an example, according to the differences in the content recognized from N frames before the Kth frame, the first character sequence may include the character sequence where "chew" is located and the character sequence where "mouth" is located. Here, the character sequence where "chew" is located can be referred to as the third character sequence, and the character sequence where "mouth" is located can be referred to as the fourth character sequence. Then, among the multiple character sequences determined by the electronic device, there are both the third character sequence and the candidate characters of the Kth frame audio recognized based on the third character sequence (including character sequences such as "chew mouth", "chew pick", "chew kou", "chew buckle", etc.), and the fourth character sequence and the candidate characters of the Kth frame audio recognized based on the fourth character sequence (including character sequences such as "seize buckle", "seize", "seize mouth", "seize kowtow", etc.). When the electronic device screens L1 character sequences from the multiple character sequences, it can be screened from the third character sequence or from the fourth character sequence.

[0013] In combination with the first aspect, in some embodiments, the method further includes: the electronic device makes predictions based on the Kth frame audio and the first character sequence to obtain multiple first candidate characters of the Kth frame audio, and determines the acoustic probability scores of the multiple first candidate characters. The first prediction result includes the acoustic probability scores; the electronic device makes semantic predictions on the Kth frame audio according to the first character sequence to obtain multiple second candidate characters of the Kth frame audio, and determines the language probability scores of the multiple second candidate characters. The language probability scores include the probabilities that the characters corresponding to the Kth frame audio under the first character sequence are each candidate character among the second candidate characters, where the second prediction result includes the language probability scores.

[0014] In combination with the first aspect, in some embodiments, the electronic device selects the second character sequence from L3 character sequences based on the second prediction result of the characters corresponding to the Kth frame audio, including: the electronic device calculates the total score of each character sequence among the L3 character sequences, and the total score is equal to the product of the language probability score and the first weight plus the acoustic probability score; the electronic device selects the second character sequence based on the total scores of the L3 character sequences.

[0015] Among them, the electronic device can use an acoustic model to make predictions based on the Kth frame audio and the first character sequence to obtain the first candidate characters, and use a language model to determine the second candidate characters according to the semantics of the first character sequence. Furthermore, the electronic device can determine the total score based on the acoustic probability score and the language probability score of each character sequence. Referring to Figure 5 the illustrated embodiment, the electronic device calculates the total score of each character sequence respectively, and then obtains the total scores of "chew mouth", "chew pick", "chew kou", "seize buckle", "seize mouth", "seize kowtow". Finally, the electronic device can select the character sequence with the highest total score as the final recognition result. For example, Figure 5In the illustrated embodiment, the "chew mouth" has the highest total score, and the "chew mouth" is the second character sequence finally determined by the electronic device. Optionally, the electronic device may input multiple character sequences with relatively high total scores into the acoustic model and the language model again, so that the acoustic model and the language model can identify the characters in the (K + 1)-th frame based on the currently recognized character sequences. For example, the electronic device may input "chew mouth" and "deduct" into the acoustic model and the language model.

[0016] Combined with the first aspect, in some embodiments, the first weight gradually increases as the N value increases.

[0017] That is to say, the electronic device will increase the magnitude of the first weight based on the recognized character sequences. When the recognized character sequences are fewer, there is less previous text that the language model can refer to when predicting characters, and its accuracy is relatively low. As the recognized character sequences increase, the language model can predict the next character based on more previous text, and its prediction accuracy gradually increases as the recognized character sequences increase. The electronic device can gradually increase the first weight as the length of the recognized character sequences increases, that is to say, gradually increase the influence of the language probability score on the total score. In this way, when there are fewer recognized characters, the electronic device can determine the characters corresponding to the audio more based on the acoustic model, and when there are more recognized characters, the language model can play a greater role in semantically correcting the prediction results of the acoustic model, thereby improving the accuracy of speech recognition.

[0018] Combined with the first aspect, in some embodiments, when the K value is 1, the first weight is 0.

[0019] Since when the electronic device recognizes the first character of a sentence, the language model has no recognized character sequences to refer to, at this time, the language model determines the language probability scores of candidate characters more based on the training data, and the accuracy is low. Therefore, when recognizing the first character, the electronic device can set the first weight to 1, that is, ignore the language probability scores obtained by the language model for predicting the first character. In this way, the total score of the candidate characters of the first character only depends on its acoustic probability score and will not be affected by the language model, improving the accuracy of the first character recognition.

[0020] In a second aspect, the present application provides an electronic device, which includes a display screen, a memory, and a processor coupled to the memory; the display screen is used to display an interface, the memory stores a computer program, and when the processor executes the above computer program, the electronic device implements the method described in any item of the first aspect above.

[0021] In a third aspect, the present application provides a computer-readable storage medium, which stores a computer program or computer instructions, and the foregoing computer program or computer instructions are executed by a processor to implement the method described in any item of the first aspect above.

[0022] Fourthly, an embodiment of the present application provides a computer program product. When the computer program product is executed by a processor, the method described in any one of the above first aspects will be implemented.

[0023] Fifthly, an embodiment of the present application provides a chip, which includes a processor and a memory. The memory is used to store a computer program or computer instructions, and the processor is used to execute the computer program or computer instructions stored in the memory, so that the chip executes the method described in any one of the above first aspects.

[0024] The solutions provided in the above second to fifth aspects are used to implement or cooperate with the implementation of the corresponding methods provided in the above first aspect. Therefore, the same or corresponding beneficial effects can be achieved as those of the corresponding methods in the first aspect, and details are not described herein again. Description of the Drawings

[0025] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0026] Figure 2 is a schematic architecture diagram of a speech recognition system 200 provided by an embodiment of the present application;

[0027] Figure 3 is a flowchart of a speech recognition method provided by an embodiment of the present application;

[0028] Figures 4-5 is a specific example diagram when the electronic device provided by an embodiment of the present application recognizes speech. Detailed Embodiments

[0029] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above", "said", "this" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It should also be understood that the term " / and" as used in the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0030] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as implying or indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0031] An electronic device usually identifies the speech content included in the audio while recording the audio, and then provides one or more functions for the user based on the speech content. The above functions may include, but are not limited to: voice notes, voice translation, human-computer conversation, and so on. Among them, voice notes may refer to the electronic device displaying the speech content recognized from the audio in a visual form such as text or image for the user to view and edit. Voice translation may refer to the electronic device converting the recognized speech content from one language to another, and playing the speech content in the converted language, or displaying the text of the speech content in the converted language. Human-computer conversation may refer to the electronic device determining the meaning of the speech content after recognizing the speech content in the audio, and then the electronic device can have a conversation with the user or execute the instructions included in the speech content, and so on. Sometimes the electronic device recognizes the speech content from the audio too slowly, resulting in a long time for it to provide the corresponding function for the user. For example, in the voice note function, the electronic device may display the text of the speech content corresponding to the audio only after receiving the audio for a long time. In the human-computer conversation function, the electronic device takes a long time to execute an action after receiving the user's voice command, which causes the user to wait too long and reduces the user experience.

[0032] In order to improve the speed of speech recognition of the electronic device, the embodiments of the present application provide a speech recognition method, an electronic device, and a computer-readable storage medium. In this method, for a single pronunciation, the electronic device first filters out a small group of candidate characters based on the acoustic probability scores of a large number of candidate characters, and then the electronic device determines the character corresponding to the pronunciation from this group of candidate characters based on the acoustic probability scores and language probability scores of the filtered candidate characters. Among them, the acoustic probability score is the probability of each candidate character corresponding to the pronunciation predicted by the acoustic model based on the acoustic features of the pronunciation, and the language probability score is the probability of each candidate character corresponding to the pronunciation predicted by the language model based on the language features of the context. The electronic device can reduce its computational amount during speech recognition and improve the speed of speech recognition by using the above method, which reduces the waiting time of the user and improves the user experience when using speech recognition-related functions.

[0033] For the convenience of description and better understanding, the subsequent embodiments of the present application introduce the speech recognition method by taking the electronic device recognizing Chinese as an example. Without being limited to recognizing Chinese, the electronic device can also recognize the speech of other languages through this speech recognition method. The method for the electronic device to recognize the speech of other languages can refer to the introduction of the method for the electronic device to recognize Chinese in the subsequent embodiments, and the embodiments of the present application will not elaborate on this.

[0034] Next, an exemplary electronic device 100 provided by the embodiments of the present application will be introduced first.

[0035] Figure 1It is a schematic structural diagram of the electronic device 100 provided by an embodiment of the present application.

[0036] The following takes the electronic device 100 as an example to specifically illustrate the embodiment. It should be understood that the electronic device 100 may have more or fewer components than those Figure 1 shown in, may combine two or more components, or may have different component configurations. Figure 1 The various components shown in may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.

[0037] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a display screen 190, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a touch sensor 180B, etc.

[0038] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or may be integrated in one or more processors.

[0039] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0040] A memory can also be set in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0041] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger.

[0042] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 190, etc.

[0043] The modem processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 190. In some embodiments, the modem processor can be an independent device.

[0044] The electronic device 100 realizes the display function through the GPU, the display screen 190, and the application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 190 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.

[0045] The display screen 190 is used to display images, videos, etc. The display screen 190 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 190, where N is a positive integer greater than 1.

[0046] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0047] The NPU is a neural-network (NN) computing processor. By learning from the biological neural network structure, such as learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn.

[0048] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0049] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as face recognition function, fingerprint recognition function, mobile payment function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as face information template data, fingerprint information template, etc.). In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0050] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0051] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals.

[0052] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.

[0053] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by bringing the receiver 170B close to the human ear.

[0054] The microphone 170C, also known as the "microphone", "transmitter", is used to convert a sound signal into an electrical signal.

[0055] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0056] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 190. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 190, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A.

[0057] The touch sensor 180B, also known as the "touch panel". The touch sensor 180B can be disposed on the display screen 190. The touch sensor 180B and the display screen 190 form a touch screen, also known as the "touch screen". The touch sensor 180B is used to detect touch operations acting on or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 190. In some other embodiments, the touch sensor 180B can also be disposed on the surface of the electronic device 100, at a different position from the display screen 190.

[0058] In some embodiments, the electronic device can record an audio stream through the microphone 170C. The audio module 170 can convert the audio data in the form of an analog signal output by the microphone 170C into a digital signal. After the DSP resamples and denoises the audio in the form of a digital signal output by the audio module 170, the audio data can be sent to the NPU. Furthermore, the NPU can identify the speech content in the audio. The method for the NPU to identify the speech content in the audio can refer to the introduction of subsequent embodiments and will not be elaborated here.

[0059] A speech recognition system can be deployed in the electronic device 100, and the speech recognition system is used to identify the speech content in the audio received by the electronic device.

[0060] The speech recognition system provided by the embodiments of the present application will be introduced below.

[0061] Figure 2 It is a schematic diagram of the architecture of the speech recognition system 200 provided by the embodiments of the present application. As Figure 2 shown, the speech recognition system 200 can include an acoustic model 201, a language model 202, and a decoder 203.

[0062] Among them, the acoustic model 201 can receive audio input and then determine the probability distribution matrix corresponding to each pronunciation based on audio features. The probability distribution matrix output by the acoustic model 201 can be a four-dimensional tensor, and the probability distribution matrix can include the acoustic probability scores of multiple candidate characters corresponding to the pronunciation. The acoustic probability score is used to represent the probability that the character corresponding to the pronunciation to be recognized is each candidate character.

[0063] Exemplarily, the acoustic model 201 may include an encoder 204, a prediction network 205, and a fusion network 206. Among them, the encoder 204 can encode the audio data to obtain an acoustic encoding sequence (or referred to as acoustic encoding features). The prediction network 205 can receive the text that has been recognized, and then encode the text to obtain a text encoding sequence (or referred to as text encoding features). Furthermore, the prediction network 205 can predict the subsequent text based on the text encoding sequence, and then generate a new text encoding sequence. The fusion network 206 can splice the text encoding sequence and the acoustic encoding sequence to form a new encoding sequence, and then output the probability distribution matrix to the decoder 203 based on the spliced encoding sequence.

[0064] Figure 2 The shown acoustic model 201 can also be referred to as an acoustic model based on the Transducer architecture. For the convenience of description and better understanding, the subsequent embodiments of this application introduce the speech recognition method with the acoustic model based on the Transducer architecture. It is not limited to the acoustic model based on the Transducer architecture. The acoustic model 201 can also adopt other architectures. And the acoustic model 201 can include more or fewer modules than Figure 2 shown. Or combine certain modules, or add or subtract certain modules. The embodiments of this application do not limit the architecture of the acoustic model 201.

[0065] The language model 202 can receive text inputs and predict subsequent characters based on semantic features of the recognized text, etc. For example, the language model 202 can determine the character corresponding to the 5th pronunciation in the current sentence based on the characters corresponding to the 1st - 4th recognized pronunciations. The more characters that have been recognized in the current sentence, the more text the language model 202 can use for prediction, and the stronger the semantic correction function it can exert. The language model 202 can output a probability distribution matrix to the decoder 203. The probability distribution matrix output by the language model 202 can include language probability scores of multiple candidate characters, and this language probability score is used to represent the probability that the character corresponding to the pronunciation to be recognized is each candidate character. The language model 202 can be trained based on the text in the corresponding field when recognizing speech content in different fields. For example, when the electronic device uses the speech recognition system 200 to recognize speech in daily life, the language model 202 can be trained based on common phrases in daily life; when the electronic device uses the speech recognition system 200 to recognize speech in the technology field, the language model 202 can be trained based on the professional terms in the corresponding technology field. In this way, the speech recognition system 200 in the electronic device can achieve relatively good speech recognition effects in the face of different application scenarios.

[0066] The language model 202 can be a neural network model based on a convolutional neural networks (CNN) model, a recurrent neural network (RNN) model, a long short - term memory (LSTM) model, a deep neural network (DNN) model, a generative pre - training transformer (GPT) model, and so on.

[0067] The decoder 203 is used to fuse the prediction results of the acoustic model 201 and the language model 202 to generate the text of the recognized speech content. The method by which the decoder 203 fuses the acoustic model 201 and the language model 202 can refer to the introduction in the subsequent embodiments and will not be elaborated here.

[0068] After the electronic device 100 records an audio stream through the microphone 170C, it can cut the audio stream to obtain valid audio segments. The above-mentioned valid audio segments may refer to audio segments containing human voices. Among them, the electronic device may cut the audio stream to obtain valid audio segments when there is a long pause in the audio stream. Among them, the audio stream is continuous audio data transmitted using streaming technology. Among them, the electronic device 100 can screen out the valid audio in the audio stream by calculating the short-time energy value or short-time zero-crossing rate of the audio segments in the audio stream. For example, the electronic device 100 can calculate the short-time energy value within each time window in the audio stream by sliding a time window. When the short-time energy value within a certain time window in the audio stream is lower than a preset threshold, the electronic device can select the starting time point of the time window to cut the audio stream. Without being limited to this method, the electronic device 100 can also cut the valid audio segments by other methods, and the embodiments of the present application do not limit this. The electronic device can cut the valid audio segments according to a preset duration to obtain multiple frames of audio. The above-mentioned preset time length may be 500 milliseconds, that is to say, the duration of each frame of audio is 500 milliseconds. Then the electronic device 100 can input the audio frame by frame into the speech recognition system to recognize the speech content.

[0069] Exemplarily, after the electronic device inputs the first frame of audio into the acoustic model 201, the encoder 204 in the acoustic model 201 will generate an acoustic coding sequence corresponding to the first frame of audio. When the electronic device recognizes the first frame of audio, there is no recognized character. At this time, the fusion network 206 mainly determines the character corresponding to the pronunciation contained in the first frame based on the acoustic coding sequence output by the encoder 204. The fusion network 206 can input the probability distribution matrix corresponding to the first frame of audio into the decoder 203. Similarly, since there is no recognized character at present, the decoder 203 can determine one or more candidate characters with higher acoustic probability scores based on the probability distribution matrix output by the fusion network 206, and input the one or more candidate characters into the language model 202 and the prediction network model 205. For example, the first frame of audio contains the pronunciation of "wo". After receiving the first frame of audio, the acoustic model 201 can input the probability distribution matrix into the decoder 203. The decoder 203 can determine multiple candidate characters with higher acoustic probability scores from the probability distribution matrix, such as "wo", "wo", and "wo". Since there are no recognized characters at present, the language model 202 cannot give an accurate prediction based on the context, so the decoder 203 can use the acoustic probability score of each candidate character output by the acoustic model 201 as the total score of the candidate character. Finally, the decoder 203 can input several determined candidate characters with high total scores into the language model 202 and the prediction network 205. Optionally, in some scenarios, the electronic device can display the candidate character with the highest total score on the screen. For example, in the voice note function, when the total score of "I", "Wo", and "Oh" is the highest, the electronic device can display the character "I" first.

[0070] Then the electronic device can input the second frame of audio into the acoustic model 201, and the encoder 204 can generate an acoustic coding sequence corresponding to the second frame of audio. The prediction network 205 can predict the candidate characters of the second frame based on the previous text (at this time, only the candidate characters of the first frame are included). For example, when the candidate characters of the first frame of audio input by the decoder 203 into the prediction network 205 include "我", "窝", "哦", the prediction network can predict the characters of the second frame based on each candidate character respectively, and then generate a text coding sequence, which is used to strengthen some characters in the acoustic coding sequence based on the semantics of the previous text. The second frame of audio can contain the pronunciation of "shi". Taking the candidate character "我" of the first frame as an example, the prediction network 205 can determine that "我" can be followed by the character "是", and then the prediction network 205 can strengthen "是" through the text coding sequence. Although "shi" can be recognized as "是", "试", "四", the acoustic probability score of "是" in the subsequent probability distribution matrix will be improved due to the strengthening of the text coding sequence. Finally, the fusion network 206 will concatenate the text encoding sequence and the acoustic encoding sequence to obtain the probability distribution matrix corresponding to the acoustic model. After receiving the language text "I" corresponding to the first frame of audio, the language model 202 will also predict the characters of the second frame to obtain the probability distribution matrix corresponding to the language model. Finally, the decoder 203 can fuse the probability distribution matrix corresponding to the acoustic model and the probability distribution matrix corresponding to the language model to determine one or more candidate characters for the second frame. The method by which the decoder 203 fuses the probability distribution matrix corresponding to the acoustic model and the probability distribution matrix corresponding to the language model can be referred to the introduction of subsequent embodiments, which will not be expanded here.

[0071] The method for the speech recognition system to recognize the speech content in subsequent frames of audio can refer to the above-mentioned method for recognizing the speech content in the first and second frames of audio, which will not be repeated here.

[0072] It should be noted that the speech recognition system 200 will determine multiple candidate characters for each frame of audio, and then recognize subsequent characters based on each candidate character, so that the speech recognition system 200 will obtain multiple groups of speech content texts. Exemplarily, when the speech recognition system 200 determines two candidate characters for each frame, it can obtain four groups of speech texts after recognizing the first frame of audio and the second frame of audio. For example, the speech recognition system 200 determines two candidate characters "I" and "Nest" based on the first frame of audio, and then determines "Yes" and "Try" based on "I", and determines "Yes" and "City" based on "Nest", so that four groups of speech texts are obtained: "I am", "I try", "Nest is", "Nest City", and then the speech recognition system will determine the candidate characters corresponding to the third frame of audio based on the above four groups of speech texts. Finally, after the speech recognition system recognizes a valid audio segment, it can select a group of speech texts with the highest total score as the final recognition result of the valid audio segment.

[0073] Combine the following Figure 2 The illustrated speech recognition system 200 specifically introduces the speech recognition method provided in the embodiment of the present application.

[0074] Figure 3 It is a flow chart of the speech recognition method provided in an embodiment of the present application. Since the speech recognition method includes two screenings of candidate characters, the speech recognition method can also be called a double-topk-based speech recognition method.

[0075] like Figure 3 As shown, the method may include but is not limited to the following steps:

[0076] S301. An electronic device receives an audio segment A, where the audio segment A includes a Kth audio frame and N audio frames before the Kth audio frame.

[0077] The audio segment A may be a frame of audio in a valid audio segment, and the audio segment A may contain pronunciation A. In Chinese recognition, each pronunciation may correspond to a Chinese character. In English recognition, a pronunciation may correspond to a phoneme. In the embodiment of the present application, the audio segment A may also be referred to as the first audio segment.

[0078] The electronic device can determine the speech content corresponding to a valid audio segment as a sentence. The valid audio segment is an audio segment in which the pause time of the speech content is less than a preset time. The electronic device can segment the valid audio segment from the audio stream according to the short-time energy value or short-time zero-crossing rate of the audio in the sliding window.

[0079] S302. The electronic device determines multiple character sequences based on the character sequence A corresponding to the recognized N frames of audio and the prediction result A of the character corresponding to the Kth frame of audio, wherein any one of the multiple character sequences contains a candidate character corresponding to the character sequence A and the Kth frame of audio, wherein the prediction result A is determined based on the Kth frame of audio and the character sequence A.

[0080] The electronic device can input the audio segment A into the acoustic model 201, and the acoustic model 201 can determine the prediction result A based on the Kth frame of audio and the character sequence A. The character sequence A may include the characters corresponding to the N frames of audio that the electronic device has recognized before the Kth frame of audio. The prediction result A may include the probability distribution matrix of the acoustic model 201 for the characters corresponding to the Kth frame of audio, and the probability distribution matrix may include the acoustic probability score of the character corresponding to the Kth frame of audio being a candidate character under a character sequence. The character sequence may be a sequence containing one or more characters, and each character sequence corresponds to a recognition result of the Kth frame of audio and the N frames of audio before the Kth frame of audio. In an embodiment of the present application, the prediction result A may also be referred to as the first prediction result, and the character sequence A may also be referred to as the first character sequence.

[0081] In the embodiment of the present application, since the speech recognition system 200 determines the candidate characters corresponding to the current frame audio based on the character sequence that has been recognized before the frame audio being recognized, the probability of the candidate characters of the current frame is related to the recognized character sequence. And because the speech recognition system 200 can respectively identify the candidate characters corresponding to the current frame audio based on multiple character sequences, in order to facilitate distinction, the acoustic probability score of the candidate characters corresponding to the frame audio under different character sequences can be called the acoustic probability score of the entire character sequence. For example, "I am", "I try", "My home is", and "My home city" are four character sequences, among which the acoustic probability score of the character sequence "I am" is also the acoustic probability score of the character corresponding to the next frame audio being "yes" when "I" has been recognized, which is used to indicate the probability of the character corresponding to the next frame audio being "yes" when "I" has been recognized.

[0082] In some embodiments, the character result A may include multiple characters. This is because the decoder 203 may include multiple recognition results for each frame of audio, so that after the decoder 203 recognizes each frame of audio, multiple character sequences will be input into the language model 202 and the prediction network 205. The acoustic model 201 will determine the candidate character corresponding to the Kth frame of audio based on the character sequence A corresponding to the N frames of audio that have been recognized in the audio segment. Finally, the acoustic model 201 can determine multiple character sequences, each of which consists of the character sequence A and a candidate character corresponding to the Kth frame of audio.

[0083] S303. The electronic device filters out M1 character sequences from multiple character sequences. The acoustic probability scores of the M1 character sequences are the highest among the multiple character sequences. The acoustic probability score of a character sequence represents the probability that the speech content of the K-th frame of audio and the previous N frames of audio is this character sequence.

[0084] The electronic device can filter out the M1 character sequences with the highest acoustic probability scores from multiple character sequences based on the prediction result A. Among them, the electronic device can rank the acoustic probability scores of each character sequence from high to low, and then determine the top M1 character sequences with the highest acoustic probability scores. Since this step selects the character sequences with the top several acoustic probability scores, it can also be called a topk of this speech recognition method. M1 can take a relatively small value to reduce the calculation amount. For example, M1 can be 5.

[0085] S304. The electronic device filters out M2 character sequences from multiple character sequences according to the M1 character sequences. The pronunciation of the candidate characters of the K-th frame of audio in the M2 character sequences is the same as or similar to the pronunciation of the candidate characters of the K-th frame of audio in any one of the M1 character sequences.

[0086] In step S303, only the top M1 character sequences with higher acoustic probability scores are determined. This may result in the situation that the correct character sequence has a relatively low acoustic probability score and is not included in the M1 character sequences, which will lead to the exclusion of the correct character sequence. In step S304, the electronic device can expand the candidate character sequences based on the pronunciation of the characters in the M1 character sequences.

[0087] The electronic device can determine one or more pronunciations of the K-th frame of audio in the above M1 character sequences, and then filter out the character sequences from multiple character sequences (i.e., the total character sequences) whose candidate characters corresponding to the K-th frame of audio are the same as or similar to the above one or more pronunciations. Finally, M2 character sequences are filtered out from multiple character sequences. The above M2 character sequences include both the above M1 character sequences and one or more newly added character sequences. The pronunciation of the candidate characters corresponding to the K-th frame of audio in the one or more newly added character sequences is similar to the one or more pronunciations corresponding to the K-th frame of audio in the above M1 character sequences.

[0088] Among them, similar pronunciations can include any of the following situations: two pronunciations with front and back nasal sounds (such as "yin" and "ying"), two pronunciations with the same final sound (such as "qing" and "xing"), and two pronunciations with different tones (such as the pronunciations corresponding to "light" and "celebrate"). In some embodiments, when two pronunciations simultaneously include multiple above situations, they can also be called similar, for example, the final sounds of the pronunciations of "light" and "walk" are the same and the tones are different, and the pronunciations of "light" and "walk" can be called similar pronunciations. The pronunciations of "sound" and "win" are front and back nasal sounds, and their tones are different, and the pronunciations of "sound" and "win" can also be called similar pronunciations.

[0089] The pronunciation of the character corresponding to the K-th frame audio in the expanded character sequence is more diverse, so as to avoid the situation that due to non-standard pronunciations in the audio, such as incorrect front and back nasal sounds, incorrect tones, etc., the acoustic probability score of the correct character sequence is relatively low, and then the correct result is excluded in step S303.

[0090] S305. The electronic device screens out M3 character sequences from M2 character sequences, and the pronunciation of the character corresponding to the K-th frame audio in the M3 character sequences appears the most frequently.

[0091] After the expansion in step S304, the number of character sequences increases, that is to say, there are a relatively large number of recognition results for the K-th frame audio and the first N frames of audio before the K-th frame audio currently. Since the prediction network 205 in the speech recognition system 200 will strengthen the acoustic probability scores of some candidate characters based on the semantics of the first N frames of audio before the K-th frame audio when determining the character corresponding to the K-th frame audio, this may increase the acoustic probability scores of the character sequences where some candidate characters with pronunciations not close to the actual pronunciation in the K-th frame audio are located. To exclude the character sequences where such candidate characters are located, the electronic device can determine the frequency of occurrence of the pronunciation of the character corresponding to the K-th frame audio in the M3 character sequences, and then determine the one or more pronunciations with the highest frequency of occurrence and the corresponding characters. Then the electronic device can screen out M3 character sequences from M2 character sequences where the pronunciation of the character corresponding to the K-th frame audio appears the most frequently.

[0092] Not limited to the M3 character sequences where the pronunciation of the character corresponding to the K-th frame audio appears the most frequently, in some embodiments, the electronic device can also screen out the character sequences from M2 character sequences where the pronunciation of the character corresponding to the K-th frame audio appears more frequently than a certain preset threshold.

[0093] For example, when the electronic device determines that the character corresponding to the K-th frame audio in the M2 character sequences has the third tone "ni" the most, it can screen out all the character sequences from M2 character sequences where the pronunciation of the character corresponding to the K-th frame audio is the third tone "ni".

[0094] S306. The electronic device screens out M4 character sequences from M3 character sequences, and the acoustic probability scores of the M4 character sequences are the highest among the M3 character sequences.

[0095] The electronic device can rank the acoustic probability scores of the M3 character sequences from high to low, and then determine the top M4 character sequences with the highest acoustic probability scores. This can be called the second topk of the speech recognition method. Among them, M4 can be relatively large compared to M1. For example, when M1 is 5, M4 can be 10. This can improve the diversity of candidate character sequences without increasing too much computational complexity.

[0096] S307. The electronic device selects one character sequence from the M4 character sequences based on the prediction result B of the character corresponding to the Kth frame of audio, and determines this one character sequence as the speech recognition result of the Kth frame of audio and the first N frames of audio before the Kth frame of audio, where the prediction result B is determined based on the character sequence A.

[0097] Both the acoustic model 201 and the language model 202 will output probability distribution matrices. Among them, the prediction result A includes the probability distribution matrix output by the acoustic model 201, and the prediction result B includes the probability distribution matrix output by the language model 202. Among them, the prediction result B is determined based on the semantic features of the recognized character sequence A, etc. In the embodiments of the present application, the prediction result B can also be referred to as the second prediction result.

[0098] After the electronic device 100 determines the M4 character sequences, it can look up the language probability scores of the M4 character sequences in the language model. Here, the language probability score is denoted as Score lm , and the acoustic probability score is denoted as Score am , then the total score Score final of each character sequence = Score am + α · Score lm , where α is the weight of the language probability score. Among them, the value of α can gradually increase as the number of recognized character sequences (i.e., the value of N) increases. That is to say, the weight α of the language probability score is proportional to the length of the recognized character sequence. For example, when the speech recognition system 200 recognizes the fifth character in the same character sequence, the weight α can be 0.2, and when it recognizes the sixth character, the weight α can be 0.22. In the embodiments of the present application, the weight α can also be referred to as the first weight.

[0099] When calculating the total score, the weight α of the electronic device increases as the length of the recognized text sequence increases. In this way, when there are fewer recognized characters, the context (i.e., the recognized characters) that the language model 202 can refer to is less, so the accuracy may also be lower. At this time, the total score of the candidate characters determined by the speech recognition system 200 depends more on the acoustic probability score of the candidate character. When the recognized characters gradually increase, since the language model 202 can make predictions based on more context (i.e., the recognized characters), the accuracy of its prediction will relatively increase. At this time, increasing the weight α by the speech recognition system can play a semantic correction function on the prediction result of the acoustic model 201, and improve the accuracy of the speech recognition system 200 in recognizing speech content. Taking the electronic device's recognition of the audio corresponding to the sentence "How many years old does the ancient term 'huajia' refer to" as an example. When the electronic device always uses the same weight to recognize the audio, it may always be mainly affected by the prediction result of the acoustic model, which may lead to its recognition of "is referred to as" as "market value", "is worth", "set", etc. When the electronic device adjusts and increases the weight of the language probability score as the recognized character sequence grows, the language probability score of the language model can play a semantic correction role on the prediction result of the acoustic model. It can be seen that when the electronic device dynamically adjusts the weight of the language probability score, the correct result "How many years old does the ancient term 'huajia' refer to" will have a higher total score.

[0100] When N is 0, that is to say, there are no recognized characters in the current sentence. At this time, the weight α can be 0. That is to say, when the speech recognition system 200 recognizes the first character in a sentence, the weight α can be 0. At this time, the total score of each character only depends on its acoustic probability score. Taking the electronic device's recognition of the audio corresponding to the sentence "Fire has no shadow" as an example, when the weight of the language model is non-zero when the electronic device recognizes the first character, the language model may determine from a semantic perspective that the probability of "I" as the first character in a sentence is relatively large, which may lead to the sequence with the highest total score in the final recognition result being "I have no shadow", which makes the recognition result of the first character deviate from its original pronunciation "huo". When the weight of the language model is 0 when the electronic device recognizes the first character, the speech recognition system determines the total score of each character only based on the recognition result of the acoustic model, and the total score of "Fire has no shadow" in the final recognition result will be higher. It can be seen that ignoring the language probability score of the first character in each sentence when the electronic device calculates the total score can improve the accuracy of the first character and avoid the language model only recognizing the first character based on the statistical probability in the training data, which affects the final recognition result.

[0101] Finally, the electronic device can use the character sequence with the highest total score as the speech recognition result of the Kth frame audio and the N frames of audio before the Kth frame.

[0102] In some embodiments, the electronic device may input multiple character sequences with higher total scores into the language model 202 and the prediction network 205, so that when the speech recognition system 200 recognizes the character corresponding to the (K + 1)-th frame of audio, it can make a prediction based on the first N + 1 audio of the (K + 1)-th frame of audio, and then obtain multiple character sequences.

[0103] When the speech recognition system 200 determines that a sentence ends, the speech recognition system 200 may select a character sequence with the highest total score as the final recognition result of the sentence. The above-mentioned sign that a sentence ends may be that the speech recognition system 200 has recognized all the frame audio in a valid audio segment (such as audio segment A).

[0104] The following uses a specific example to introduce the speech recognition method provided by the embodiments of the present application.

[0105] Figures 4-5 It is a specific example diagram when the electronic device provided by the embodiments of the present application recognizes speech.

[0106] Here, an example is given in which the speech recognition system 200 selects two characters for input into the language model 202 and the prediction network 205 when recognizing each frame of audio. This example introduces the method of the electronic device when recognizing part of the speech content "chew gum" in a sentence. Here, the pronunciations to be recognized are: "jiao", "kou", "xiang", "tang". For the convenience of description and better understanding, Figure 4 、 Figure 5 only two characters corresponding to the K-th frame of audio are shown. It can be understood that there may be more frames of audio before the K-th frame of this sentence, and each frame of audio may include two recognition results (characters). The multiple characters recognized from multiple frames of audio together form multiple character sequences.

[0107] As Figure 4 shown, the electronic device determines that the characters corresponding to the pronunciations in the (K - 1)-th frame of audio are "chew" and "pay". Among them, the character sequences where "chew" and "pay" are located may be the two character sequences with the highest total scores ranked in descending order when the electronic device recognizes the (K - 1)-th frame of audio. In this example, the character sequences where "chew" and "pay" are located can be called two character sequences in character sequence A.

[0108] After the electronic device 100 obtains the K-th frame of audio from the valid audio clip, the K-th frame of audio can be input into the acoustic model 201. Then the acoustic model 201 will output the probability distribution matrix corresponding to the K-th frame of audio, and the probability distribution matrix contains the probability that the character corresponding to the pronunciation in the K-th frame of audio is each character in the total character set. The acoustic model can determine the probability that the character corresponding to the audio is each character in the total character set each time the audio is recognized. The electronic device inputs a total of two character sequences: the character sequence in which the characters corresponding to the K-th frame of audio are "食用" and "付". The acoustic model 201 will obtain two probability distribution matrices for "食用" and "付" respectively. The probability distribution matrix can indicate the probability that the K-th frame of audio and the N frames of audio before the K-th frame of audio are a certain character sequence. Among them, when the character corresponding to the pronunciation in the K-1 frame audio is "chew", the candidate characters with acoustic probability scores from high to low in the K frame audio are "mouth", "knocking", "bandit", "button", etc., and the acoustic probability scores of the character sequence from high to low can also be called "chew mouth", "chew knocking", "chew bandit", "chew button", etc. When the character corresponding to the pronunciation in the K-1 frame audio is "pay", the candidate characters with acoustic probability scores from high to low in the K frame audio can be "button", "win", "mouth", "knocking", etc., and the acoustic probability scores of the character sequence from high to low can also be called "pay button", "win", "pay mouth", "pay knock". Then the electronic device can execute step S303 to determine M1 characters based on the probability distribution matrix corresponding to the pronunciation in the K-1 frame audio. Figure 4 In the embodiment shown, M1 can be 2. Since the acoustic probability scores of the character sequences "咬口" and "咬抠" are relatively high, the electronic device 100 can use "口" and "抠" as candidate characters for the Kth frame of audio when the K-1th frame of audio is recognized as "咬". Similarly, the electronic device can use "扣" and "获" as candidate characters for the Kth frame of audio when the K-1th frame of audio is recognized as "交".

[0109] Then the electronic device executes step S304 to expand the candidate character set based on the pronunciation of the candidate characters, that is, the electronic device uses characters with the same or similar pronunciation as the selected candidate characters as candidate characters for the Kth frame of audio. For example, the electronic device can also use characters with the same or similar pronunciation as the candidate characters excluded in step S303 as candidate characters. Figure 4 As shown, in the branch (hereinafter referred to as branch 1) where the pronunciation corresponding character is "食用" in the K-1 frame audio, the electronic device can determine characters such as "蔻" and "扣" that are similar in pronunciation to "口" and "抠" as candidate characters corresponding to the K-th frame audio in the branch. Similarly, the electronic device can also expand branch 2, for example, "口" and "敲" that are similar in pronunciation to "扣" are also used as candidate characters. There are few characters with the same pronunciation as "获" among the candidate characters, so "获" is excluded from the candidate characters.

[0110] Then the electronic device can perform step S305 to filter the characters corresponding to one or more pronunciations with the highest frequency of occurrence in each branch. In branch 1, three characters corresponding to pronunciations are retained, namely, "抠" corresponding to the first tone "kou", "口" corresponding to the third tone "kou", and "蔻" and "扣" corresponding to the fourth tone "kou". In branch 2, the frequency of occurrence of the corresponding pronunciation of "获" is relatively low, so "获" is excluded, leaving only "扣", "口" and "敲" as candidate characters.

[0111] Finally, the electronic device can select M4 characters with the highest acoustic probability scores from the candidate characters. In this way, under branch 1, the electronic device can determine the three characters with the highest acoustic probability scores from the candidate characters "口", "抠", "蔻", "扣" and so on. Under branch 2, the electronic device can determine the three characters with the highest acoustic probability scores under this branch: "扣", "口", "敲".

[0112] like Figure 5 As shown, after the decoder determines the candidate characters in branch 1 and branch 2, it can determine the total score of each character sequence based on the candidate characters in branch 1 and branch 2 respectively. Taking branch 1 as an example, after the decoder 203 determines that the candidate characters are "口", "抠", and "蔻", it will search for the language probability scores of "口", "抠", and "蔻" in the language model 202. The language probability scores of "口", "抠", and "蔻" in the language model 202 are determined by the language model 202 based on the prediction result "食用" of the Kth frame. Then the decoder 203 can calculate the total score of each character sequence. Here, taking the weight α as 0.2 as an example, the decoder 203 can calculate the total score of the character sequence "咬口" = acoustic probability score (0.15) + weight (0.2) × language probability score (0.1) = 0.17. Similarly, the decoder 203 can calculate the total scores of the character sequences "咬抠", "咬蔻", and the total scores of the character sequences "交扣", "交口", and "交敲" under branch 2 in turn. Finally, the decoder 203 may compare the total scores of all character sequences to determine the recognition results of the K-1th frame and the Kth frame audio.

[0113] In some embodiments, the electronic device needs to display the recognized speech content while performing speech recognition. Since the total score of the character sequence "嚼口" exceeds the total scores of the other five, the electronic device can use "嚼口" as the current recognition result and display it on the display screen. It should be noted that the speech recognition content displayed by the electronic device here does not represent the final recognition result of the speech. The electronic device will continue to recognize subsequent frames and select the character sequence with the highest total score in the subsequent frames as the new recognition result.

[0114] In this example, when the speech recognition system 200 recognizes each frame of audio, it selects two characters and inputs them into the language model 202 and the prediction network 205. The decoder 203 determines the two character sequences with the highest total scores, including the "bit" sequence with a total score of 0.17 in branch 1 and the "payment" sequence with a total score of 0.158 in branch 2, and then inputs these two character sequences into the language model 202 and the prediction network 205 to predict the characters corresponding to the (K + 1)-th frame of audio. That is to say, after calculating the total scores of all character sequences, the decoder 203 determines the final recognition result based on the total scores of all branch character sequences.

[0115] When the speech recognition system recognizes the (K + 1)-th frame of audio, the prediction network 205 in the acoustic model 201 will increase the acoustic probability scores of some candidate characters based on the recognized character sequences (including the "bit" character sequence and the "payment" character sequence), and the language model 202 will also determine the language probability scores of the characters corresponding to the (K + 1)-th frame of audio based on the recognized character sequences. Finally, the decoder 203 can determine the character sequence containing the characters corresponding to the (K + 1)-th frame of audio based on the prediction results of the acoustic model 201 and the language model 202. The method by which the decoder 203 determines the (K + 1)-th frame of audio based on the prediction results of the acoustic model 201 and the language model 202 can refer to the introduction in the above embodiment and will not be elaborated here.

[0116] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

[0117] As used in the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0118] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.

[0119] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media that can store program codes such as ROM or random access memory RAM, magnetic disks, or optical discs.

Claims

1. A speech recognition method, characterized in that, The method is applied to an electronic device, and the method includes: The electronic device receives a first audio segment, where the first audio segment includes the Kth frame of audio and the first N frames of audio before the Kth frame of audio, K is a positive integer, and N is a positive integer; The electronic device determines multiple character sequences based on a first prediction result of the Kth frame of audio and a first character sequence corresponding to the identified first N frames of audio. Any one of the multiple character sequences includes a candidate character corresponding to the first character sequence and the Kth frame of audio. The first prediction result includes multiple candidate characters of the Kth frame of audio; The electronic device filters out L1 character sequences from the multiple character sequences. The acoustic probability scores of the L1 character sequences are the highest among the multiple character sequences. The acoustic probability score of a character sequence represents the probability that the speech content of the Kth frame of audio and the first N frames of audio is the character sequence. L1 is a positive integer; The electronic device filters out L2 character sequences from the multiple character sequences according to the L1 character sequences. The pronunciation of the candidate character of the Kth frame of audio in the L2 character sequences satisfies a first condition with the pronunciation of the candidate character of the Kth frame of audio in any one of the L1 character sequences. L2 is a positive integer; The electronic device filters out L3 character sequences from the L2 character sequences. The acoustic probability scores of the L3 character sequences satisfy a second condition. L3 is a positive integer; The electronic device selects a second character sequence from the L3 character sequences based on a second prediction result of the Kth frame of audio and determines the second character sequence as the speech recognition result of the Kth frame of audio and the first N frames of audio. The second prediction result is obtained by the electronic device through semantic prediction based on the speech content of the Kth frame of audio.

2. The method according to claim 1, wherein The electronic device filters out the L3 character sequences from the L2 character sequences, specifically including: The electronic device filters out L4 character sequences from the L2 character sequences. The pronunciation of the character corresponding to the Kth frame of audio in the L4 character sequences appears most frequently. L4 is a positive integer; The electronic device filters out the L3 character sequences from the L4 character sequences. The L4 character sequences are those with the highest acoustic probability scores among the L3 character sequences.

3. The method according to claim 1 or 2, characterized in that, The first condition includes any one of the following: the finals of the two pronunciations are the same, the two pronunciations only differ in tone, and the two pronunciations are front nasal and back nasal.

4. The method according to any one of claims 1 to 3, characterized in that, The first character sequence includes a third character sequence and a fourth character sequence. Any one of the multiple character sequences includes a candidate character corresponding to the third character sequence and the Kth frame of audio, or includes a candidate character corresponding to the fourth character sequence and the Kth frame of audio.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: The electronic device makes predictions based on the K-th frame audio and the first character sequence to obtain multiple first candidate characters for the K-th frame audio, and determines the acoustic probability scores of the multiple first candidate characters. The first prediction result includes the acoustic probability scores. The electronic device performs semantic prediction on the K-th frame audio according to the first character sequence to obtain multiple second candidate characters for the K-th frame audio, and determines the language probability scores of the multiple second candidate characters. The language probability scores include the probabilities that the characters corresponding to the K-th frame audio under the first character sequence are each of the candidate characters among the second candidate characters. The second prediction result includes the language probability scores.

6. The method according to claim 5, wherein The electronic device selects a second character sequence from the L3 character sequences based on the second prediction result of the character corresponding to the K-th frame audio, including: The electronic device calculates the total score of each character sequence among the L3 character sequences. The total score is equal to the product of the language probability score and the first weight value plus the acoustic probability score. The electronic device selects the second character sequence based on the total scores of the L3 character sequences.

7. The method according to claim 6, wherein The first weight value gradually increases as the N value increases.

8. The method according to claim 6 or 7, characterized in that, When the K value is 1, the first weight value is 0.

9. An electronic device, characterized in that, The electronic device includes: a display screen, a memory, and a processor coupled to the memory. The display screen is used to display a user interface. The memory stores a computer program. When the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instruction runs on the electronic device, the electronic device executes the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speech recognition method and device

    CN110942763A

  • Speech recognition method and device, electronic equipment and readable storage medium

    CN112509565A

  • Speech recognition method and device, equipment and storage medium

    CN112599128A

  • Speech recognition method and device, storage medium and electronic equipment

    CN113990293A

  • Speech recognition method, interaction method, storage medium and program product

    CN114141232A