Keyword recognition method and apparatus, and electronic device

By decoding the identification information in electronic devices using keyword sets, full sets or non-keyword sets, and determining the output results based on similarity, the problem of degradation in speech keyword recognition performance in complex background noise environments is solved, and the balance between recognition accuracy and false alarm rate is achieved.

WO2025112446A1PCT designated stage expired Publication Date: 2025-06-05HUAWEI TECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099051
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-06-13
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In complex background noise environments, the performance of speech keyword recognition has dropped sharply, resulting in an increase in false alarm rate or a decrease in recognition accuracy, and even causing the system to fail to work normally.

Method used

By obtaining the information to be identified in the electronic device and decoding it through the keyword set and the full set or the non-keyword set respectively, the first sequence and the second sequence are obtained. Then, the output of the keyword recognition result is determined based on whether the similarity between the two satisfies the set threshold condition.

Benefits of technology

The balance of voice keyword recognition performance in different environments is achieved, the balance of recognition accuracy and false alarm rate is ensured, thereby ensuring the stable availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099051_05062025_PF_FP_ABST
    Figure CN2024099051_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electronic devices. Disclosed are a keyword recognition method and apparatus, and an electronic device, which are used for ensuring the balance between recognition accuracy and a false alarm rate. The method comprises: acquiring information to be subjected to recognition, and decoding said information by means of a keyword set, so as to obtain a first sequence; decoding said information by means of a full set or a non-keyword set, so as to obtain a second sequence; determining the similarity between the first sequence and the second sequence; and when the similarity meets a set threshold condition, outputting the first sequence as a keyword recognition result. In the present application, the first sequence is output as the keyword recognition result only when the similarity between the first sequence and the second sequence meets the set threshold condition, so that the balance between the keyword recognition accuracy and the false alarm rate can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Keyword recognition method, device and electronic device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 29, 2023, with application number 202311626722.3 and application name “A keyword recognition method, device and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The embodiments of the present application relate to the field of speech recognition technology, and in particular to a keyword recognition method, device, and electronic device. Background Art

[0004] With the emergence of smart speakers, voice assistants, and other applications, users are using voice to communicate with machines. Voice keyword recognition is a key technology for enabling human-machine voice interaction and is widely used in various smart devices and voice search systems to wake up and control devices.

[0005] Currently, speech keyword recognition achieves excellent results in noise-free or lightly noisy environments. However, in real-world environments with complex background noise (e.g., subway stations, karaoke bars, conference rooms, etc.), performance can decline dramatically (false alarm rates increase dramatically or recognition accuracy decreases sharply), and can even cause the system to malfunction. Therefore, a method that can guarantee speech keyword recognition performance in diverse environments is urgently needed.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a keyword recognition method, device, and electronic device to ensure a balance between the accuracy and false alarm rate of keyword recognition.

[0008] In a first aspect, embodiments of the present application provide a keyword recognition method for an electronic device. In this method, the electronic device obtains information to be recognized; decodes the information to be recognized using a keyword set to obtain a first sequence; decodes the information to be recognized using a full set or a non-keyword set to obtain a second sequence; determines the similarity between the first sequence and the second sequence, and when the similarity meets a set threshold, outputs the first sequence as a keyword recognition result. Optionally, the keyword set may include multiple keywords or a single keyword.

[0009] In this method, after obtaining the information to be identified, the electronic device can decode the information to be identified using a keyword set, a full set, or a non-keyword set to obtain a first sequence and a second sequence, so that the electronic device can determine the final output of the keyword recognition result by determining whether the similarity between the first sequence and the second sequence meets the set threshold condition, thereby achieving a balance between the recognition accuracy and the false alarm rate.

[0010] In one possible design, the electronic device can perform feature extraction on the information to be identified to obtain feature information of the information to be identified; the electronic device can input the feature information into a keyword model, perform keyword recognition on the feature information based on the keyword model, and obtain at least one first candidate sequence; the electronic device can output the first candidate sequence corresponding to the maximum likelihood probability as the first sequence.

[0011] With this design, the electronic device can output the first candidate sequence corresponding to the maximum likelihood probability identified by the keyword model as the first sequence, thereby ensuring the accuracy of keyword recognition.

[0012] In one possible design, the keyword model is constructed based on a first modeling unit, and the first modeling unit includes at least one of a pronunciation, a phoneme, and a phrase corresponding to the keyword set.

[0013] Through this design, the keyword module can be constructed in multiple ways and can be applied to terminal devices of various sizes.

[0014] In one possible design, the electronic device can extract features from the information to be identified to obtain feature information of the information to be identified; the electronic device inputs the feature information into a background model, identifies the feature information based on the background model, and obtains at least one second candidate sequence; the electronic device can output the second candidate sequence corresponding to the maximum likelihood probability as a second sequence.

[0015] With this design, the electronic device can output the second candidate sequence corresponding to the maximum likelihood probability identified by the background model as the second sequence, thereby ensuring recognition accuracy.

[0016] In one possible design, the background model is constructed based on a second modeling unit, and the second modeling unit includes at least one of a pronunciation, a phoneme, and a phrase. Optionally, the second modeling unit includes at least one of a pronunciation, a phoneme, and a phrase corresponding to the full set or the non-keyword set.

[0017] With this design, electronic devices can build background models based on a balancing strategy between storage and computing resources. This allows the background model to cover a wide range of electronic devices, allowing them to choose different background model building methods based on their resource availability.

[0018] In one possible design, the electronic device may further determine similarity based on the probability corresponding to the first sequence and the probability corresponding to the second sequence. The probability corresponding to the first sequence represents the probability of training the first sequence using elements in the keyword set, and the probability corresponding to the second sequence represents the probability of training the second sequence using elements in the full set or non-keyword set.

[0019] With this design, the electronic device can determine the similarity between the first sequence and the second sequence by using the probability of the keyword model constructing the first sequence and the probability of the background model constructing the second sequence, and can intuitively judge whether the keyword recognition is accurate or not.

[0020] In a second aspect, the present application provides a keyword recognition device, the keyword recognition device comprising:

[0021] An identification module is configured to obtain information to be identified; decode the information to be identified using a keyword set to obtain a first sequence; and decode the information to be identified using a full set or a non-keyword set to obtain a second sequence;

[0022] The determination module is configured to determine the similarity between the first sequence and the second sequence; when the similarity satisfies a set threshold condition, output the first sequence as a keyword recognition result.

[0023] In a third aspect, the present application provides an electronic device comprising one or more processors and one or more memories. The one or more memories are configured to store one or more computer programs and data information, and the one or more processors are configured to execute the computer programs stored in the one or more memories, so that the electronic device performs the method described in any of the first aspects above. The electronic device may be the electronic device described in the first aspect above.

[0024] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed by a computing device, the computing device executes the method of any aspect of the above-mentioned first aspect and any possible implementation of any aspect.

[0025] In a fifth aspect, an embodiment of the present application provides a chip, which is coupled to a memory in an electronic device so that the chip calls a computer program stored in the memory during operation to implement a method that may be designed in any aspect of the first aspect of the embodiment of the present application.

[0026] In a sixth aspect, the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a computing device, the computing device executes the method of any aspect of the above-mentioned first aspect and any possible implementation of any aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic diagram of a keyword recognition scenario;

[0028] FIG2 is a schematic diagram of the hardware structure of a possible electronic device exemplarily provided in this application;

[0029] FIG3 is a flow chart of a keyword recognition method provided by the present application;

[0030] FIG4a is a schematic diagram of a decoding search graph of a keyword space provided by the present application;

[0031] FIG4 b is a schematic diagram of another decoding search graph of a keyword space provided by the present application;

[0032] FIG4c is a schematic diagram of a decoding search graph of another keyword space provided by the present application;

[0033] FIG4 d is a schematic diagram of a decoding search graph of another keyword space provided by the present application;

[0034] FIG5 is a schematic diagram of a decoding search graph of a full decoding search space provided by the present application;

[0035] FIG6 is a schematic diagram of an application architecture of an electronic device provided by this application;

[0036] FIG7 is a schematic diagram of the structure of a decoding image module provided by the present application;

[0037] FIG8 is a schematic diagram of a training process of a decoding graph module provided by the present application;

[0038] FIG9 is a schematic diagram of a complete flow chart of a keyword recognition method provided by this application;

[0039] FIG10 is a schematic structural diagram of a keyword recognition device provided in this application. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be described in detail below in conjunction with the drawings in the following embodiments of the present application.

[0041] First, the concepts related to the embodiments of the present application are explained.

[0042] (1) Keyword spotting (KWS), a subfield of speech recognition, aims to detect all occurrences of a specified word in a speech signal. Currently, it is primarily used for keyword spotting and isolated word recognition in unconstrained speech. This application primarily addresses isolated word recognition.

[0043] (2) Modeling unit refers to the basic unit for constructing a model.

[0044] (3) False alarm rate refers to the probability of detecting a target when it is not present, even though it is not present, due to the prevalence and fluctuation of noise when using threshold detection methods during radar detection or other detection processes. In this application, false alarm rate refers to the probability that an electronic device determines that a keyword is present when it is not present.

[0045] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.

[0046] Furthermore, unless otherwise indicated, ordinal numbers such as "first" and "second" in the embodiments of this application are used to distinguish between multiple objects and are not used to limit the size, content, order, timing, priority, or importance of the multiple objects. For example, "second file" and "second file" are only used to distinguish different files and do not indicate a difference in size, content, priority, or importance between the two files.

[0047] Currently, speech keyword recognition achieves excellent results in noise-free or lightly noisy environments. However, in real-world environments with complex background noise (e.g., subway stations, karaoke bars, conference rooms, etc.), performance can decline dramatically (false alarm rates increase dramatically or recognition accuracy decreases sharply), and can even cause the system to malfunction. Therefore, a method that can guarantee speech keyword recognition performance in diverse environments is urgently needed.

[0048] In order to ensure the accuracy of keyword recognition, an embodiment of the present application provides a keyword recognition method, which is applied to an electronic device. In this method, after obtaining the information to be recognized, the electronic device can decode the information to be recognized through a keyword set to obtain a first sequence, and decode the information to be recognized through a full set or a non-keyword set to obtain a second sequence. The electronic device can also determine the similarity between the first sequence and the second sequence, and when the similarity meets a set threshold condition, output the first sequence as a keyword recognition result. In this way, the electronic device can decide the output of the keyword recognition result based on the obtained similarity, so that the false alarm rate of keyword recognition can meet actual needs while ensuring the accuracy of keyword recognition, thereby ensuring a balance between the accuracy rate and the false alarm rate.

[0049] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0050] The embodiments of the present application can be applied to various scenarios in which related functions are executed through keyword recognition. For example, as shown in FIG1 , a user can communicate with an electronic device through voice. The electronic device acquires voice information to be recognized by collecting voice information from the surrounding environment. The electronic device can perform keyword recognition on the information to be recognized. For another example, the electronic device can also connect to an information collection device or database to obtain the information to be recognized for keyword recognition.

[0051] In some examples, the electronic device can decode the voice information to be recognized using a keyword set to obtain a first sequence. At the same time, the electronic device can also decode the voice information to be recognized using a full set or a non-keyword set to obtain a second sequence. The electronic device can determine the output of the keyword recognition result based on the similarity between the first sequence and the second sequence. Exemplarily, when the similarity between the first sequence and the second sequence meets the set threshold condition, the electronic device can output the first sequence as the keyword recognition result. When the electronic device determines that the keyword recognition result is the first sequence, it can execute related functions, such as waking up the phone.

[0052] FIG2 shows a schematic diagram of the hardware structure of a possible electronic device. The electronic device 100 may be the electronic device in FIG1 . As shown in FIG2 , the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0053] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a microcontroller unit (MCU), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. The controller may serve as the nerve center and command center of the electronic device 100. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The processor 110 may also include memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a high-speed cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. Repeated access is avoided, the waiting time of the processor 110 is reduced, and the efficiency of the system is improved.

[0054] In an embodiment of the present application, the processor 110 may store one or more applications, obtain information to be identified, decode the information to be identified using a keyword set to obtain a first sequence, and decode the information to be identified using a full set or a non-keyword set to obtain a second sequence. Furthermore, the processor 110 may determine whether the similarity between the first sequence and the second sequence meets a set threshold condition to output a keyword recognition result.

[0055] The USB interface 130 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices. The charging management module 140 is used to receive charging input from the charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.

[0056] The wireless communication functionality of electronic device 100 can be implemented using antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0057] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0058] The wireless communication module 160 can provide wireless communication solutions including wireless local area network (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signal, and sends the processed signal to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0059] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-CDMA), long term evolution (LTE), the fifth generation (5G) mobile communication system, future communication systems such as the sixth generation (6G) system, BT, GNSS, WLAN, NFC, FM and / or IR technology, etc. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0060] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode or an active-matrix organic light emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1. In an embodiment of the present application, the display screen 194 can be used to display the main interface, application interface, granularity adjustment controls, etc.

[0061] The camera 193 is used to capture still images or videos. The camera 193 may include a front camera and a rear camera.

[0062] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, and software code of at least one application (such as Huawei Video, Changlian, etc.). The data storage area can store data (such as images, videos, etc.) generated during the use of the electronic device 100. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0063] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as pictures and videos can be stored on the external memory card.

[0064] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0065] It is understood that the components shown in FIG2 do not constitute a specific limitation on the electronic device. The electronic device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. In the following embodiments, the electronic device 100 shown in FIG2 is used as an example for description.

[0066] It should be noted that the electronic device involved in the keyword recognition method provided in the embodiment of the present application can be a terminal device or a server. For example, the terminal device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) device, a virtual reality (VR) device, a laptop computer, a drone, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart speaker, a smart TV and other smart devices. In the embodiment of the present application, there is no special restriction on the specific form of the electronic device. For example, the electronic device can be a mobile phone or a smart speaker.

[0067] The present application is described below with reference to specific embodiments.

[0068] FIG3 is a schematic diagram of a keyword recognition method provided by an embodiment of the present application. As shown in FIG3 , the method includes:

[0069] S301: The electronic device obtains information to be identified.

[0070] The information to be identified can be any information that satisfies the search range of a finite set. For example, the information to be identified can be voice information, text information, blood pressure, or heart rate. Blood pressure and heart rate are waveform signals.

[0071] As an example, the electronic device can collect information about its surroundings to obtain information to be identified. For example, if the electronic device is a mobile phone and the information to be identified is voice information, the mobile phone can automatically collect information about its surroundings to obtain the voice information.

[0072] As another example, the electronic device may also obtain the information to be identified from an information collection device or database that collects the information to be identified. For example, if the information to be identified is blood pressure, the user may use a blood pressure monitor to measure blood pressure. The electronic device may connect to the blood pressure monitor and obtain the blood pressure data measured by the blood pressure monitor.

[0073] S302: The electronic device decodes the information to be identified using a keyword set to obtain a first sequence, wherein the keyword set is composed of one or more keywords.

[0074] In some examples, after obtaining the information to be identified, the electronic device may perform feature extraction on the information to be identified to obtain feature information of the information to be identified. For example, when the information to be identified is voice information, the feature information may be mel-scale frequency cepstral coefficients (MFCC) and the like. For another example, when the information to be identified is blood pressure data or heart rate, the electronic device may sample and quantize the blood pressure data or heart rate signal, and convert the time domain continuous signal into discrete data (array, matrix), which is the feature information.

[0075] The electronic device may input feature information of the information to be identified into a keyword model, and perform keyword recognition on the feature information based on the keyword model to obtain at least one first candidate sequence. The keyword model is constructed based on a keyword set. The electronic device may output the first candidate sequence corresponding to the maximum likelihood probability as the first sequence. Exemplarily, when the keyword model performs keyword recognition on the feature information, it outputs the first candidate sequence corresponding to the maximum likelihood probability as the keyword model's recognition result, thereby ensuring the accuracy of the keyword model's keyword recognition.

[0076] Exemplarily, when the information to be recognized is voice information or text information, the keyword may be a text or phrase, that is, the electronic device performs keyword recognition on the feature information based on the keyword model in order to determine the keyword sequence in the voice information or text information.

[0077] For another example, when the information to be identified is blood pressure data, the keyword can be a waveform signal with a mutation point. That is, the electronic device performs keyword recognition on the characteristic information of the blood pressure data based on the keyword model in order to determine the mutation value in the blood pressure data. The mutation value in the blood pressure data is the mutation point in the blood pressure data waveform signal.

[0078] For another example, when the information to be identified is heart rate data, the keyword may be a waveform signal with a mutation point, that is, the electronic device performs keyword recognition on the characteristic information of the heart rate data through a keyword model in order to identify the mutation value in the heart rate data.

[0079] The keyword model is constructed based on the first modeling unit, and the first modeling unit can be the basic constituent unit of the keyword set. In some embodiments, when the keyword is speech or text, the first modeling unit includes at least one of the pronunciation, phoneme, and word (phrase) corresponding to the keyword set. Pronunciation, phoneme, and word are the basic constituent units of speech. For example, when the speech is Chinese speech, the pronunciation is the pinyin of a single Chinese character, the phoneme is the initial and final consonants, and the word is a Chinese character phrase. Exemplarily, the electronic device can construct a decoding search graph of the keyword space based on the pronunciation, phoneme, or word, and use the constructed decoding search graph of the keyword space as the keyword model. Additionally, multiple groups of keywords can be constructed in the keyword model.

[0080] As an example, the keyword model can be constructed based on the pronunciation corresponding to the keyword set. For example, as shown in Figure 4a, when the keyword set includes the keyword "Hello Xiaoka", the pronunciation corresponding to the keyword set is "ni", "hao", "xiao", "ka". The electronic device performs keyword recognition on the feature information based on the keyword model, obtains one or two first candidate sequences, and calculates the likelihood probability corresponding to the first candidate sequence. Among them, the first candidate sequence can be "ni", "hao", "xiao", "ka", or can also be the empty sequence "non". Among them, the keyword model regards all sequences different from the keyword sequence "ni", "hao", "xiao", "ka" as garbage edges, that is, the empty sequence. When the first candidate sequence corresponding to the maximum likelihood probability is "ni", "hao", "ni", "hao", the electronic device outputs "ni", "hao", "ni", "hao" as the first sequence.

[0081] Another example is that, as shown in Figure 4b, when the keyword set includes two groups of keywords, "Hello Xiaoka" and "Hello Hello", the keyword model has two sequences of pronunciations corresponding to the keywords, "ni", "hao", "xiao", "ka" and "ni", "hao", "ni", "hao", as well as a garbage edge. The electronic device performs keyword recognition on the feature information based on the keyword model, obtains at least one first candidate sequence, and calculates the likelihood probability corresponding to at least one first candidate sequence respectively. Among them, the first candidate sequence can be "ni", "hao", "xiao", "ka", can also be "ni", "hao", "ni", "hao", or can also be the empty sequence. When the first candidate sequence corresponding to the maximum likelihood probability is "ni", "hao", "ni", "hao", the electronic device outputs "ni", "hao", "ni", "hao" as the first sequence.

[0082] As another example, the keyword model can also be constructed based on the phonemes corresponding to the keyword set. For example, as shown in FIG. 4c, when the keyword set includes one keyword "Hello", the phonemes corresponding to the keyword set are "n", "i", "h", "ao", with a total of 4 phoneme elements. The present application can construct a decoding search graph for the keyword space based on these 4 phoneme elements.

[0083] As another example, the keyword model can also be constructed based on the phrases corresponding to the keyword set. For example, when the keyword set includes the keyword "Hello Xiaoka", the phrase corresponding to the keyword set consists of "Ni", "Hao", "Xiao", "Ka". As shown in FIG. 4d, the present application can construct a decoding search graph for the keyword space based on the keyword phonetic sounds, where the decoding search graph has a total of 4 phonetic sound elements.

[0084] In some other embodiments, when the information to be recognized is blood pressure data or heart rate, the first modeling unit may include information such as the waveform amplitude, peak, and trough corresponding to the mutation point. Exemplarily, when the information to be recognized is blood pressure data, the electronic device can construct a keyword recognition model based on the waveform amplitude corresponding to the mutation point. After the electronic device inputs the feature information of the blood pressure data into the keyword model, it performs keyword recognition based on the keyword model to determine at least one first candidate sequence. The first candidate sequence can be an array with a mutation point or an empty sequence. Among them, the keyword model sets all arrays without mutation points as empty sequences. The electronic device can respectively determine the likelihood probabilities corresponding to at least one first candidate sequence, and output the first candidate sequence corresponding to the maximum likelihood probability as the first sequence.

[0085] S303: The electronic device decodes the information to be recognized through the full set or the non-keyword set to obtain a second sequence. The full set includes the keyword set and other information, and the non-keyword set includes the information in the full set except the keyword set. For example, when the keyword set consists of "Hao", the full set consists of all Chinese characters, and the non-keyword set consists of Chinese characters other than "Hao".

[0086] In some embodiments, before the electronic device decodes the information to be recognized through the full set or the non-keyword set, it can also extract the features of the information to be recognized to obtain the feature information of the information to be recognized. The process by which the electronic device obtains the feature information is the same as the process of determining the feature information in step S302.

[0087] The electronic device may also input the feature information into a background model (BGM), decode the feature information based on the background model, and obtain at least one second candidate sequence. The background model is constructed based on a full set or a non-keyword set. The electronic device may output the second candidate sequence with the maximum likelihood probability as the second sequence. The background model in the embodiment of the present application may absorb non-keyword information or noise input to reduce environmental interference.

[0088] The background model is constructed based on the second modeling unit, which can be a basic constituent unit of the full set or the non-keyword set. In some embodiments, when the keyword is voice or text, the second modeling unit includes at least one of the pronunciations, phonemes, and phrases corresponding to the full set or the non-keyword set. Exemplarily, the second modeling unit may include at least one of the full pronunciations, full phonemes, and full phrases, and may also include at least one of the pronunciations, phonemes, and phrases corresponding to non-keywords.

[0089] The present application can construct a decoding search graph of a full decoding search space or a non-keyword space based on pronunciation, phonemes or phrases, and use the constructed decoding search graph of the full decoding search space or the non-keyword space as a background model. Exemplarily, the present application can select different second modeling units to construct a background model based on the resource situation of the electronic device. Furthermore, the present application can select a suitable modeling unit to construct a background model in combination with the storage resource and computing resource balancing strategy of the electronic device.

[0090] As an example, when computing resources are sufficient and storage resources are tight, the present application can adopt a phoneme-based model construction method. Exemplarily, the present application can construct a decoding search graph of the background space based on the full amount of phonemes or non-keyword phonemes to obtain a background model. For example, when the information to be recognized is Chinese speech, the present application can construct a decoding search graph of the background space based on the full amount of phonemes. Among them, the full amount of phonemes includes Chinese initials and finals, initials: b, p, m, f, d, t, n, l, g, k, h, j, q, x, zh, ch, sh, z, c, s, y, w, r; finals: a, o, e, i, u, v, ai, ei, ui, ao, ou, iu, ie, ve, er, an, en, in, un, vn, ang, eng, ing, ong.

[0091] As another example, when both computing resources and storage resources are abundant, the present application may adopt a model construction method based on phonetic sounds. Exemplarily, the present application may construct a decoding search graph of the background space based on all phonetic sounds or non-keyword phonetic sounds to obtain a background model. For example, when the information to be recognized is Chinese and the keyword is "Hello, Xiaoka", the present application may construct a decoding search graph of the background space based on all phonetic sounds. Among them, the decoding search graph of the background space contains 4 phonetic sound elements of the keyword "Hello, Xiaoka", namely "ni", "hao", "xiao", and "ka". Another example is that when the present application constructs a decoding search graph of the background space based on non-keyword phonetic sounds, the decoding search graph does not contain these 4 phonetic sound elements of "ni", "hao", "xiao", and "ka".

[0092] As another example, when computing resources are scarce and storage resources are sufficient, the present application may adopt a model construction method based on phrases. Exemplarily, the present application may construct a decoding search graph of the background space based on all phonetic sounds or non-keyword phonetic sounds to obtain a background model. For example, as shown in FIG. 5, when the information to be recognized is Chinese and the keyword is "Hello, Xiaoka", the present application may construct a decoding search graph of the background space based on all phrases. Among them, the decoding search graph of the background space shown in FIG. 5 contains 4 phrase elements of the keyword "Hello, Xiaoka", namely "ni", "hao", "xiao", and "ka". Another example is that when the information to be recognized is Chinese and the keyword is "Hello, Xiaoka", the present application may also construct a decoding search graph of the background space based on non-keyword phrases. Among them, the decoding search graph of the background space does not contain 4 phrase elements of the keyword "Hello, Xiaoka", namely "ni", "hao", "xiao", and "ka".

[0093] In some other embodiments, when the information to be recognized is blood pressure data or heart rate, the second modeling unit may include information such as waveform amplitude, wave peak, and wave trough, which is not limited herein. Exemplarily, when the information to be recognized is blood pressure data, the electronic device may construct a background model based on the waveform amplitude. After the electronic device inputs the characteristic information of the blood pressure data into the background model, it performs recognition based on the background model to determine at least one second candidate sequence. Among them, the second candidate sequence is a waveform array. The electronic device may respectively determine the likelihood probabilities corresponding to at least one second candidate sequence, and output the second candidate sequence corresponding to the maximum likelihood probability as the second sequence.

[0094] S304: The electronic device determines the similarity between the first sequence and the second sequence.

[0095] In some embodiments, after the electronic device outputs the first sequence through the keyword model and the second sequence through the background model, it can also determine the similarity between the first sequence and the second sequence based on the probability corresponding to the first sequence and the probability corresponding to the second sequence. The probability corresponding to the first sequence represents the probability of training the first sequence using elements in the keyword set, and the probability corresponding to the second sequence represents the probability of training the second sequence using elements in the full set or non-keyword set. For example, the closer the probability corresponding to the first sequence is to the probability corresponding to the second sequence, the higher the similarity between the first sequence and the second sequence.

[0096] As an example, the probability corresponding to the first sequence and the probability corresponding to the second sequence can be likelihood probability or posterior probability. In addition, the probability corresponding to the first sequence and the second sequence can also be other probabilities, which are not limited here.

[0097] As an example, the electronic device may use the ratio of the probability corresponding to the first sequence to the probability corresponding to the second sequence as the similarity between the first and second sequences. The probability corresponding to the first sequence is the probability that the elements in the keyword model constitute the first sequence, and the probability corresponding to the second sequence is the probability that the elements in the background model constitute the second sequence. Alternatively, the electronic device may determine the similarity between the first and second sequences using other calculation methods, which are not limited here.

[0098] For example, the embodiment of the present application may determine the similarity between the first sequence and the second sequence by the following formula: R likelihood =L kws / L bg ; Among them, R likelihood is the similarity, L kws represents the probability corresponding to the first sequence output by the keyword model, L bg Represents the probability corresponding to the second sequence output by the background model. The electronic device can likelihood As the similarity between the first sequence and the second sequence. likelihood The closer it is to 1, the higher the similarity between the first sequence and the second sequence. kws Indicates the likelihood probability corresponding to the first sequence, L bg When the likelihood probability of the second sequence is expressed, the similarity R likelihood is the likelihood probability ratio.

[0099] For example, as shown in Figure 4d, each arrow in the decoding search graph of the keyword model has a corresponding weight value. The weight value represents the probability of the keyword model selecting this element. The probability corresponding to the first sequence is the probability L that the four elements "you", "good", "small", and "card" in the keyword model form the path "you"-"good"-"small"-"card". kws=m1*m2*m3*m4*m5. As shown in Figure 5, each arrow in the decoding search graph of the background model has a corresponding weight value. The weight value represents the probability of the background model selecting this element. The probability corresponding to the second sequence is the probability L that the four elements "you", "good", "small", and "card" in the background model form the path "you"-"good"-"small"-"card". bg =n1*n2*n3*n4*n5. Electronic devices can be kws / L bg As the similarity between the first sequence and the second sequence.

[0100] In other embodiments, when the information to be identified is blood pressure data or heart rate data, the electronic device may determine the similarity between the first sequence and the second sequence. For example, the electronic device may compare elements in the first sequence with elements in the second sequence to determine the similarity between the first sequence and the second sequence.

[0101] S305: When the similarity meets the set threshold condition, the electronic device outputs the first sequence as a keyword recognition result.

[0102] In some embodiments, after obtaining the similarity, the electronic device can control the output of the keyword recognition result by determining whether the similarity meets the set threshold condition. Among them, the set threshold condition can be manually set according to actual needs. The threshold in the set threshold condition can be related to the false alarm rate, that is, the user can set the relevant threshold condition according to the actual needs of the false alarm rate. For example, when the information to be identified contains keywords, the similarity between the first sequence and the second sequence is infinitely close to 1. Taking into account situations such as loss of calculation accuracy, the similarity between the first sequence and the second sequence is usually not equal to 1 in actual situations. For another example, when the information to be identified does not contain keywords, the similarity between the first sequence and the second sequence must be less than 1. The user can set the set threshold condition according to this situation to maintain a balance between recognition accuracy and false alarm rate.

[0103] In an embodiment of the present application, a user can directly set a threshold condition according to actual usage requirements. For example, the user can set the threshold condition in a configuration user interface (UI) corresponding to the threshold condition. In addition, the user can also configure the threshold condition through an application program interface (API) for configuring the threshold condition. Among them, the false alarm rate and the threshold value in the threshold condition are linearly negatively correlated. For example, when the false alarm rate is less than 0.5%, the user can configure the threshold condition to have a similarity greater than 0.995.

[0104] As an example, when the false alarm rate is X, the threshold conditions include but are not limited to the following:

[0105] Case 1: Similarity is greater than or equal to 1-X;

[0106] Case 2: The similarity is less than or equal to 1 and greater than or equal to 1-X;

[0107] Case 3: When X is between M and N, the similarity is greater than or equal to 1-N and less than or equal to 1-M, where N is greater than M.

[0108] When the electronic device determines that the similarity meets a set threshold, it determines that the information to be identified is a keyword and outputs the first sequence as the keyword identification result. As another example, when the similarity does not meet the set threshold, it is determined that the information to be identified does not contain a keyword, and the keyword identification result is determined to be null. In this way, the electronic device will only output the first sequence as the keyword identification result when the first and second sequences meet the set threshold, allowing the electronic device to perform the relevant function.

[0109] As an example, after the user configures the threshold conditions, the embodiment of the present application can determine the keyword recognition result through binary decision making. The decision can be expressed by the following formula:

[0110] Among them, R likelihood is the similarity, L kws represents the probability corresponding to the first sequence output by the keyword model, L bg Represents the probability corresponding to the second sequence output by the background model. When the decision output is 1, it indicates that the first sequence is the same as the second sequence, that is, the information to be identified includes keywords. In this case, the electronic device will output the first sequence as the keyword recognition result. When the decision output is 0, it indicates that the information to be identified does not include keywords. In this case, the electronic device will not output a keyword recognition result, or the output keyword recognition result will be empty. In this way, a balance between keyword recognition accuracy and false alarm rate can be maintained in different scenarios and application environments.

[0111] Based on the content shown in the above embodiment, the electronic device decides the output of the keyword recognition result by determining whether the similarity between the first sequence and the second sequence meets the set threshold condition. This allows users to directly and simply limit the false alarm rate of keyword recognition by adjusting the set threshold condition in different scenarios, different application environments, different user needs, etc., so that the false alarm rate does not change with environmental changes, thereby ensuring a balance between the keyword recognition accuracy and the false alarm rate, and further ensuring the stable availability of the keyword recognition system.

[0112] The content executed by the electronic device shown in the above embodiment can be executed by an application in the electronic device. Figure 6 is an application architecture diagram of an electronic device provided in an embodiment of the present application. The electronic device includes a feature extraction module and a decoding graph module.

[0113] The feature extraction module is used to extract features from the information to be identified to obtain feature information of the information to be identified.

[0114] The decoding graph module is used to perform keyword recognition on feature information to obtain keyword recognition results.

[0115] As shown in FIG7 , the decoding graph module may include a keyword model module, a background model module, a likelihood ratio module, and a decision module.

[0116] The keyword model module is used to build a keyword model, perform keyword recognition on the feature information of the information to be recognized through the keyword model, and output the sequence corresponding to the maximum likelihood probability. The sequence corresponding to the maximum likelihood probability is the sequence with the highest score in the keyword model.

[0117] The background model module is used to build a background model, identify the feature information of the information to be identified through the background model, and output the sequence corresponding to the maximum likelihood probability.

[0118] The likelihood ratio module is used to determine the similarity between the sequence output by the keyword model module and the sequence output by the background model module.

[0119] The decision module is used to decide the output of keyword recognition results based on similarity.

[0120] In some embodiments, the embodiments of the present application may train the decoding graph module in the following manner.

[0121] As shown in Figure 8, in the embodiment of the present application, sample features can be input into the decoding graph module to train the decoding graph module. The sample features include sample keyword features and sample non-keyword features.

[0122] The electronic device can construct a keyword model based on the first modeling unit corresponding to the sample keyword. Furthermore, the electronic device can also construct a background model for the entire search space or non-keyword space based on the second modeling unit. The keyword model construction process is the same as the keyword model construction process in step S302 above, and the background model construction process is the same as the background model construction process in step S303 above, and will not be repeated here.

[0123] After constructing the keyword model and background model, the electronic device places them into the keyword model module and background model module, respectively, within the decoding graph module. After inputting sample features into the keyword model module, the electronic device uses the keyword module to perform keyword recognition on the sample features, obtaining at least one sample sequence and outputting the sample sequence with the maximum likelihood probability as the first sequence. Similarly, after inputting sample features into the background model module, the electronic device uses the background model to perform keyword recognition on the sample features, obtaining at least one sample sequence and outputting the sample sequence with the maximum likelihood probability as the second sequence.

[0124] The keyword model module outputs the first sequence to the likelihood ratio module, and the background model module outputs the second sequence to the likelihood ratio module. The likelihood ratio module can calculate the ratio of the probability corresponding to the first sequence and the probability corresponding to the second sequence to obtain the likelihood probability ratio, and output the likelihood probability ratio as the similarity between the first sequence and the second sequence to the decision module. Among them, the probability corresponding to the first sequence is the probability that the elements in the keyword model constitute the first sequence, and the probability corresponding to the second sequence is the probability that the elements in the background model constitute the second sequence. For example, the probability corresponding to the first sequence can be the likelihood probability of the first sequence, and the probability corresponding to the second sequence can be the likelihood probability of the second sequence.

[0125] The decision module determines the output of keyword recognition results by determining whether the similarity meets a set threshold. Users can customize the threshold in the decision module through the UI or corresponding API. The electronic device can adjust the decoding graph module based on the keyword recognition results, completing the decoding graph module training process.

[0126] Based on the contents shown in the above embodiments, users can customize the set threshold conditions of the decision module in the decoding graph module according to actual conditions, so that the electronic device can output keyword recognition results according to user needs, and thus can achieve a balance between the accuracy and false alarm rate of the keyword recognition system in any environment, and the performance is stable and user-configurable in accordance with the actual needs of users.

[0127] FIG9 is a schematic diagram of a complete flow chart of a keyword identification method provided by the present application. As shown in FIG9 , the method includes the following steps:

[0128] S901: The electronic device obtains information to be identified.

[0129] The information to be recognized may be voice information.

[0130] S902: The electronic device extracts features from the information to be identified to obtain feature information of the information to be identified.

[0131] S903: The electronic device inputs the feature information into a keyword model, performs keyword recognition on the feature information based on the keyword model, and obtains at least one first candidate sequence.

[0132] The keyword model is constructed based on a first modeling unit, and the first modeling unit includes at least one of a pronunciation, a phoneme, and a phrase corresponding to the keyword set.

[0133] S904: The electronic device uses the first candidate sequence corresponding to the maximum likelihood probability as the first sequence.

[0134] S905: The electronic device inputs the feature information into the background model, identifies the feature information based on the background model, and obtains at least one second candidate sequence.

[0135] The background model is constructed based on the second modeling unit, and the second modeling unit includes at least one of a pronunciation, a phoneme, or a phrase corresponding to the full set or the non-keyword set. The number of pronunciation elements, phoneme elements, or phrase elements in the background model is greater than the number of pronunciation elements, phoneme elements, or phrase elements in the keyword model.

[0136] S906: The electronic device uses the second candidate sequence corresponding to the maximum likelihood probability as the second sequence.

[0137] S907: The electronic device uses the ratio of the probability corresponding to the first sequence to the probability corresponding to the second sequence as the similarity between the first sequence and the second sequence. The probability corresponding to the first sequence represents the probability that the elements in the keyword model constitute the first sequence, and the probability corresponding to the second sequence represents the probability that the elements in the background model constitute the second sequence. The closer the ratio of the probability corresponding to the first sequence to the probability corresponding to the second sequence is to 1, the higher the similarity between the first sequence and the second sequence.

[0138] S908: The electronic device determines whether the similarity meets a set threshold condition, and if so, executes step S909; if not, executes step S910.

[0139] S909: The electronic device outputs the first sequence as a keyword recognition result.

[0140] When the electronic device outputs the first sequence as a keyword result, it determines that the information to be identified is a keyword and performs a related function. For example, when the keyword is a wake-up keyword for the electronic device, the electronic device wakes up the electronic device when it determines that the information to be identified is a keyword.

[0141] S910: The electronic device determines that the keyword recognition result is empty.

[0142] Based on the content shown in Figure 9, the electronic device can output the recognition result according to the user's needs by determining whether the similarity between the first sequence and the second sequence meets the set threshold condition. The user can also freely modify the set threshold condition according to their own needs, ensuring the balance between the keyword recognition accuracy and the false alarm rate in different scenarios and different application environments, thereby ensuring the stable availability of the keyword recognition system.

[0143] Based on the above content and the same concept, FIG10 is a structural diagram of a possible keyword recognition device 1000 provided by the present application, and the keyword recognition device 1000 is applied to an electronic device.

[0144] In some embodiments, the electronic device includes the keyword recognition device 1000, or the electronic device is the keyword recognition device 1000. The keyword recognition device 1000 can be used to implement the functions of the electronic device in the above method embodiment, and thus can also achieve the beneficial effects of the above method embodiment.

[0145] Identification module 1001 is used to obtain information to be identified; decode the information to be identified using a keyword set to obtain a first sequence; and decode the information to be identified using a full set or a non-keyword set to obtain a second sequence;

[0146] The determination module 1002 is configured to determine the similarity between the first sequence and the second sequence; when the similarity satisfies a set threshold condition, the first sequence is output as a keyword recognition result.

[0147] Optionally, when the similarity does not meet the set threshold condition, the determination module 1002 may not output the keyword recognition result.

[0148] In an optional implementation manner, the identification module 1001 is specifically configured to:

[0149] Extracting features of the information to be identified to obtain feature information of the information to be identified;

[0150] Inputting the feature information into a keyword model, performing keyword recognition on the feature information based on the keyword model, and obtaining at least one first candidate sequence;

[0151] The first candidate sequence corresponding to the maximum likelihood probability is used as the first sequence.

[0152] An optional implementation is that the keyword model is constructed based on a first modeling unit, and the first modeling unit includes at least one of a pronunciation, a phoneme, and a phrase corresponding to the keyword set.

[0153] In an optional implementation manner, the identification module 1001 is specifically configured to:

[0154] Extracting features of the information to be identified to obtain feature information of the information to be identified;

[0155] Inputting the feature information into a background model, and identifying the feature information based on the background model to obtain at least one second candidate sequence;

[0156] The second candidate sequence corresponding to the maximum likelihood probability is used as the second sequence.

[0157] An optional implementation manner is that the background model is constructed based on a second modeling unit, and the second modeling unit includes at least one of a pronunciation, a phoneme, and a phrase.

[0158] In an optional implementation manner, the determination module 1002 is specifically configured to:

[0159] The similarity is determined based on the probability corresponding to the first sequence and the probability corresponding to the second sequence; the probability corresponding to the first sequence represents the probability of training the first sequence through the elements in the keyword set, and the probability corresponding to the second sequence represents the probability of training the second sequence through the elements in the full set or the non-keyword set.

[0160] Based on the above content and the same concept, the present application provides an electronic device, including a processor and a memory, the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so that the electronic device performs the steps performed by the electronic device in the above method embodiment.

[0161] Based on the above content and the same concept, the present application provides a computer-readable storage medium, which stores a computer program or instruction. When the computer program or instruction is executed by a computing device, the computing device executes the steps performed by the electronic device in the above method embodiment.

[0162] Based on the above content and the same concept, the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a computing device, the computing device executes the steps performed by the electronic device in the above method embodiment.

[0163] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0164] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flow charts and / or one or more boxes in the block diagram.

[0165] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0167] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A keyword recognition method, characterized in that: The method comprises: Obtaining information to be identified; Decoding the information to be identified by using a keyword set to obtain a first sequence; Decoding the information to be identified through a full set or a non-keyword set to obtain a second sequence; Determine the similarity between the first sequence and the second sequence; when the similarity meets a set threshold condition, output the first sequence as a keyword recognition result.

2. The method according to claim 1, characterized in that The decoding of the information to be identified by using a keyword set to obtain a first sequence includes: Extracting features of the information to be identified to obtain feature information of the information to be identified; Inputting the feature information into a keyword model, performing keyword recognition on the feature information based on the keyword model to obtain at least one first candidate sequence; the keyword model is constructed based on the keyword set; The first candidate sequence corresponding to the maximum likelihood probability is used as the first sequence.

3. The method according to claim 2, characterized in that The keyword model is constructed based on a first modeling unit, and the first modeling unit includes at least one of a pronunciation, a phoneme, and a phrase corresponding to the keyword set.

4. The method according to any one of claims 1 to 3, characterized in that: The decoding of the information to be identified by using the full set or the non-keyword set to obtain a second sequence includes: Extracting features of the information to be identified to obtain feature information of the information to be identified; Inputting the feature information into a background model, identifying the feature information based on the background model, and obtaining at least one second candidate sequence; the background model is constructed based on the full set or the non-keyword set; The second candidate sequence corresponding to the maximum likelihood probability is used as the second sequence.

5. The method according to claim 4, characterized in that The background model is constructed based on a second modeling unit, and the second modeling unit includes at least one of a character sound, a phoneme, and a phrase.

6. The method according to any one of claims 1 to 5, characterized in that: The determining the similarity between the first sequence and the second sequence includes: The similarity is determined based on the probability corresponding to the first sequence and the probability corresponding to the second sequence; the probability corresponding to the first sequence represents the probability of training the first sequence through the elements in the keyword set, and the probability corresponding to the second sequence represents the probability of training the second sequence through the elements in the full set or the non-keyword set.

7. A keyword recognition device, characterized in that: The device comprises: An identification module is used to obtain information to be identified; decode the information to be identified through a keyword set to obtain a first sequence; decode the information to be identified through a full set or a non-keyword set to obtain a second sequence; The determination module is used to determine the similarity between the first sequence and the second sequence; when the similarity meets a set threshold condition, the first sequence is output as a keyword recognition result.

8. The device according to claim 7, characterized in that The identification module is specifically used for: Extracting features of the information to be identified to obtain feature information of the information to be identified; Inputting the feature information into a keyword model, and performing keyword recognition on the feature information based on the keyword model to obtain at least one first candidate sequence; The first candidate sequence corresponding to the maximum likelihood probability is used as the first sequence.

9. The device according to claim 8, characterized in that The keyword model is constructed based on a first modeling unit, and the first modeling unit includes at least one of a pronunciation, a phoneme, and a phrase corresponding to the keyword set.

10. The device according to any one of claims 7 to 9, characterized in that: The identification module is specifically used for: Extracting features of the information to be identified to obtain feature information of the information to be identified; Inputting the feature information into a background model, identifying the feature information based on the background model, and obtaining at least one second candidate sequence; The second candidate sequence corresponding to the maximum likelihood probability is used as the second sequence.

11. The device according to claim 10, characterized in that The background model is constructed based on a second modeling unit, and the second modeling unit includes at least one of a character sound, a phoneme, and a phrase.

12. The device according to any one of claims 7 to 11, characterized in that: The determination module is specifically used for: The similarity is determined based on the probability corresponding to the first sequence and the probability corresponding to the second sequence; the probability corresponding to the first sequence represents the probability of training the first sequence through the elements in the keyword set, and the probability corresponding to the second sequence represents the probability of training the second sequence through the elements in the full set or the non-keyword set.

13. An electronic device, characterized in that: include: one or more processors; one or more memories; The one or more memories store one or more computer programs, and the one or more computer programs include instructions, which, when executed by the one or more processors, enable the electronic device to perform the method as claimed in any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.

15. A computer program product, characterized in that The invention comprises a computer program, which, when being executed on a computer, enables the computer to execute the method according to any one of claims 1 to 6.

16. A chip, characterized in that: The chip is coupled to a memory, and the chip reads a computer program stored in the memory to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice recognition method and system

    CN110223678A

  • Voice keyword recognition method and device

    CN111798840A

  • Audio processing method and device, language model training method and device and computer equipment

    CN111933129A

  • Voice keyword recognition method and device, computer equipment and storage medium

    CN112259101A

  • Voice keyword recognition method and system

    CN114937449A