Voice keyword recognition method, system and device and medium
By obtaining contextual information of candidate keywords in the ASR system and using a preset keyword library for error correction, the problem of misidentification in noisy and accented environments in the ASR system is solved, and high-precision speech keyword recognition is achieved.
Patent Information
- Application Number
- CN202511616614.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-13
AI Technical Summary
Existing ASR systems are prone to misidentifying speech keywords under noise, accents, and environmental interference, leading to errors in the execution of critical commands, and lack real-time error correction capabilities.
By obtaining contextual information of candidate keywords, semantic analysis and error correction are performed using a pre-set keyword library, and misidentified candidate keywords are replaced with standard keywords, thereby improving recognition accuracy.
It effectively reduced the misrecognition rate of voice keywords, improved the recognition accuracy of the ASR system in multiple scenarios, and ensured the accurate execution of keyword commands.
Smart Images

Figure CN121528202A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and in particular to a speech keyword recognition method, system, device and medium. BACKGROUND
[0002] With the wide application of speech recognition (ASR) technology in intelligent cockpit, smart home, vehicle-mounted system and cross-device interaction and other various speech interaction scene fields, accurately recognizing keywords is crucial to improving product experience.
[0003] In related technologies, ASR systems often misrecognize similar words due to noise interference, ambiguous pronunciation, accent differences and environmental interference in actual application, resulting in incorrect execution of key commands. SUMMARY
[0004] The present application aims to solve at least one of the technical problems existing in the prior art and proposes a speech keyword recognition method, system, device and medium.
[0005] In a first aspect, the present application embodiment provides a speech keyword recognition method, which comprises:
[0006] obtaining a candidate keyword in the to-be-processed speech;
[0007] determining whether the candidate keyword is a misrecognized keyword according to context information of the candidate keyword in the to-be-processed speech;
[0008] in response to determining that the candidate keyword is a misrecognized keyword, obtaining a standard keyword corresponding to the candidate keyword from a preset keyword library, and generating a speech recognition result according to the standard keyword.
[0009] In a second aspect, the present application embodiment provides a speech keyword recognition system, which comprises:
[0010] a keyword obtaining unit configured to obtain a candidate keyword in the to-be-processed speech;
[0011] a misrecognition determining unit configured to determine whether the candidate keyword is a misrecognized keyword according to context information of the candidate keyword in the to-be-processed speech;
[0012] a recognition result generating unit configured to, in response to determining that the candidate keyword is a misrecognized keyword, obtain a standard keyword corresponding to the candidate keyword from a preset keyword library, and generate a speech recognition result according to the standard keyword.
[0013] In a third aspect, an electronic device is provided, comprising:
[0014] one or more processors;
[0015] a memory for storing one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the voice keyword recognition method of the embodiments of the present application.
[0017] In a fourth aspect, a computer readable medium is provided, and the computer readable medium stores a computer program, and the computer program is executed by a processor to implement the voice keyword recognition method of the embodiments of the present application.
[0018] The voice keyword recognition method provided by the embodiments of the present application can effectively reduce the misrecognition rate of the voice keyword, improve the recognition accuracy of the ASR system for the voice keyword, and realize high-precision multi-scene voice keyword enhanced recognition by analyzing and judging whether there is misrecognition of the candidate keyword of the recognized voice, and correcting the misrecognized candidate keyword through the preset keyword library. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 A flowchart of a voice keyword recognition method provided by the embodiments of the present application is provided.
[0020] Figure 2 A flowchart of a candidate keyword acquisition method in the embodiments of the present application is provided.
[0021] Figure 3 A flowchart of another voice keyword recognition method provided by the embodiments of the present application is provided.
[0022] Figure 4 A structural block diagram of a voice keyword recognition system provided by the embodiments of the present application is provided.
[0023] Figure 5 A structural block diagram of an electronic device provided by the embodiments of the present application is provided. DETAILED DESCRIPTION
[0024] In order for those skilled in the art to better understand the technical solutions of the present application, the following describes exemplary embodiments of the present application in conjunction with the accompanying drawings, which include various details of the embodiments of the present application to help understanding, and should be considered only as exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Also, in order to be clear and concise, the description in the following description omits the description of well-known functions and structures.
[0025] In the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0026] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0027] The terms used herein are only used to describe specific embodiments and are not intended to limit the present application. As used herein, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprise" and / or "consist of", when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The terms "connected" or "coupled" and / or similar terms are not limited to a physical or mechanical connection, but can include an electrical connection, whether direct or indirect.
[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.
[0029] In the related art, the ASR system often misrecognizes similar words due to noise, accent, and environmental interference in actual application, leading to incorrect execution of key commands, and the ASR system relies on a static keyword library for key word recognition, lacking the ability to correct keywords in real time for context changes.
[0030] To solve at least one of the technical problems existing in the above related art, the present application provides a voice keyword recognition method. Figure 1 A flowchart of a voice keyword recognition method provided by an embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 1 As shown in FIG. 1, the voice keyword recognition method includes the following steps:
[0031] Step S11, candidate keywords in the to-be-processed voice are acquired.
[0032] In the embodiment of the present application, the to-be-processed voice can be recognized by a preset voice recognition system, and the candidate keywords in the to-be-processed voice can be recognized by a preset keyword library.
[0033] Step S12, whether the candidate keywords are misrecognized keywords is determined according to context information of the candidate keywords in the to-be-processed voice.
[0034] In the embodiment of the present application, the semantic understanding and analysis of the context information of the candidate keyword is performed to determine whether the candidate keyword is misrecognized, i.e., whether the keyword is correctly recognized, so as to correct the keyword of the candidate keyword.
[0035] In step S13, in response to determining that the candidate keyword is a misrecognized keyword, a standard keyword corresponding to the candidate keyword is obtained from the preset keyword library, and a speech recognition result is generated according to the standard keyword.
[0036] In the embodiment of the present application, the preset keyword library includes a standard keyword and a similar keyword corresponding to the standard keyword, and the similar keyword refers to a word having similarity in pronunciation, spelling or semantic characteristics with the standard keyword.
[0037] When it is determined that the candidate keyword is a misrecognized keyword, it indicates that the currently recognized candidate keyword is incorrect, and the candidate keyword needs to be corrected. The candidate keyword is replaced by the corresponding standard keyword in the preset keyword library as the current speech recognition result.
[0038] According to the speech keyword recognition method in the embodiment of the present application, the context information of the recognized speech is analyzed to determine whether there is misrecognition, and the misrecognized candidate keyword is corrected by the preset keyword library, thereby effectively reducing the misrecognition rate of the speech keyword, improving the recognition accuracy of the ASR system for the speech keyword, and realizing high-precision speech keyword enhanced recognition in a multi-speech interaction scene.
[0039] In some embodiments, after determining whether the candidate keyword is a misrecognized keyword, the speech keyword recognition method further includes: in response to determining that the candidate keyword is a correctly recognized keyword, generating a speech recognition result according to the candidate keyword.
[0040] When it is determined that the candidate keyword is a correctly recognized keyword, it indicates that the currently recognized candidate keyword is accurate and there is no misrecognition, so no keyword correction is needed, and therefore the candidate keyword can be directly used as the current speech recognition result.
[0041] Figure 2 A flowchart of a candidate keyword acquisition method in the embodiment of the present application is shown in some embodiments, as shown in FIG. 1, the step of obtaining the candidate keyword in the to-be-processed speech, i.e., the step S11 described above, can further include: Figure 2
[0042] In step S21, the to-be-processed speech is input into a preset speech recognition system for speech recognition to obtain a to-be-processed text corresponding to the to-be-processed speech.
[0043] The preset speech recognition system is a speech recognition system based on automatic speech recognition (ASR) technology, which is a technology for converting speech signals into text.
[0044] After obtaining the user input speech to be processed, the speech to be processed is input into the preset speech recognition system for speech recognition to obtain the text to be processed corresponding to the speech to be processed.
[0045] The speech recognition system can use an end-to-end ASR model (such as DeepSpeech, Whisper, Wav2Vec2.0, etc.) or a HMM-DNN (Hidden Markov Model-Deep Neural Network) based method for speech recognition. Among them, the end-to-end ASR model can use a multi-channel speech input fusion speech recognition method, a streaming speech recognition method or a non-streaming speech recognition method for speech recognition. For vehicle or smart home application scenarios, a multi-channel speech input fusion strategy can be used to improve recognition robustness, or a streaming ASR method or a non-streaming ASR method can be selected according to real-time response and high-precision batch processing application requirements.
[0046] In some embodiments, inputting the speech to be processed into the preset speech recognition system for speech recognition to obtain the text to be processed corresponding to the speech to be processed includes: obtaining speech features of the speech to be processed; inputting the speech features of the speech to be processed into the preset speech recognition system for speech recognition to obtain the text to be processed.
[0047] In some embodiments, obtaining the speech features of the speech to be processed includes: performing speech feature extraction on the speech to be processed to obtain the speech features corresponding to the speech to be processed.
[0048] Among them, the speech features are acoustic features of the speech to be processed, such as MFCC (Mel Frequency Ceptral Coefficient, Mel Frequency Cepstrum Coefficient) features, FBank (Filter Bank, Filter Bank) features, etc.
[0049] In some embodiments, before the step of inputting the speech to be processed into the preset speech recognition system for speech recognition to obtain the text to be processed corresponding to the speech to be processed, the speech keyword recognition method can further include: obtaining the speech to be processed; and pre-processing the speech to be processed, the pre-processing including filtering processing and / or speech endpoint detection processing.
[0050] In some embodiments, a microphone array or a single microphone can be used to collect user speech input to obtain the speech to be processed of the user speech input and to perform signal enhancement on the speech to be processed.
[0051] In some embodiments, the pre-processing of the to-be-processed speech includes: filtering the to-be-processed speech; and performing speech endpoint detection on the filtered to-be-processed speech to obtain valid to-be-processed speech.
[0052] In some embodiments, an adaptive filter (such as a Wiener filter, spectral subtraction, etc.) can be used to filter the to-be-processed speech, remove noise in the to-be-processed speech, improve the quality of the speech signal, and achieve noise suppression and echo cancellation of the to-be-processed speech.
[0053] In some embodiments, an endpoint detection algorithm (such as energy detection, deep learning, voice activity detection VAD) can be used to perform speech endpoint detection on the to-be-processed speech, determine valid speech segments in the to-be-processed speech, and obtain valid to-be-processed speech, thereby avoiding interference from irrelevant noise in the to-be-processed speech.
[0054] In step S22, the to-be-processed text is input into a preset keyword library for keyword matching to obtain candidate keywords. The preset keyword library includes standard keywords and similar keywords corresponding to the standard keywords.
[0055] In some embodiments, a preset keyword library for different application scenarios can be constructed in advance for products in different application scenarios (such as vehicle-mounted, smart home, conference system, etc.), and a mapping relationship table between standard keywords and similar keywords in the preset keyword library can be maintained. In the preset keyword library for each application scenario, each preset standard keyword and similar keywords corresponding to the standard keyword are included.
[0056] In some embodiments, for each product in an application scenario, a certain number of input speech or text samples under the product in the application scenario can be collected, keywords in the samples can be extracted using the TF-IDF (Term Frequency-Inverse Document Frequency) method, and each standard keyword and similar keyword in the application scenario can be determined through manual review. The preset keyword library for the application scenario can be constructed according to the determined standard keywords and similar keywords in the application scenario, and a mapping relationship table between the standard keywords and the similar keywords in the preset keyword library can be maintained.
[0057] The mapping relationship between the standard keywords and the similar keywords can be determined using at least one of various matching methods, such as semantic similarity based on word vectors, spelling similarity based on edit distance, and pronunciation similarity based on phoneme comparison.
[0058] The semantic similarity based on the word vector refers to that the semantic similarity between the standard keyword and the similar keyword is represented by calculating the vector distance between the word vector of the standard keyword and the word vector of the similar keyword; when the semantic similarity between the word vector of the standard keyword and the word vector of the similar keyword is greater than or equal to a preset semantic similarity, the similar keyword is determined as the keyword similar to the standard keyword, and a mapping relationship between the similar keyword and the standard keyword is established.
[0059] The spelling similarity based on the edit distance refers to that the spelling similarity between the standard keyword and the similar keyword is represented by calculating the edit distance between the standard keyword and the similar keyword; when the spelling similarity between the standard keyword and the similar keyword is greater than or equal to a preset spelling similarity, the similar keyword is determined as the keyword similar to the standard keyword, and a mapping relationship between the similar keyword and the standard keyword is established.
[0060] The pronunciation similarity based on the phoneme contrast refers to that the pronunciation similarity between the pronunciation of the standard keyword and the pronunciation of the similar keyword is calculated by a phoneme contrast algorithm; when the pronunciation similarity between the standard keyword and the similar keyword is greater than or equal to a preset pronunciation similarity, the similar keyword is determined as the keyword similar to the standard keyword, and a mapping relationship between the similar keyword and the standard keyword is established.
[0061] In some embodiments, the similarity between the standard keyword and the similar keyword can also be judged by comprehensively matching the semantic similarity based on the word vector, the spelling similarity based on the edit distance, the pronunciation similarity based on the phoneme contrast and the like, for example, the similarities corresponding to the above-mentioned matching manners are weighted and summed, and the result of the weighted sum is taken as the comprehensive similarity between the standard keyword and the similar keyword; when the comprehensive similarity is greater than or equal to a preset similarity threshold, the similar keyword is determined as the keyword similar to the standard keyword, and a mapping relationship between the similar keyword and the standard keyword is established.
[0062] Figure 3 Another flowchart of the voice keyword recognition method provided by the embodiments of the present application is shown in FIG. 6. In some embodiments, after the step of obtaining the candidate keyword in the voice to be processed, i.e., after the step S11, the voice keyword recognition method can further include: Figure 3
[0063] The step S31, in response to the candidate keyword being the standard keyword in the preset keyword library, generates a voice recognition result according to the standard keyword.
[0064] When the candidate keyword is a standard keyword in the preset keyword library, it means that the identified candidate keyword is correct and there is no misidentification. Therefore, there is no need to judge the misidentified keyword. Thus, the candidate keyword, i.e. the standard keyword, can be used as the current speech recognition result.
[0065] Step S32: In response to the candidate keyword being a similar keyword in the preset keyword library, determine whether the candidate keyword is a misidentified keyword based on the context information of the candidate keyword in the speech to be processed.
[0066] When the candidate keyword is a similar keyword in the preset keyword library, it indicates that the identified candidate keyword may be misidentified. In order to determine whether there is misidentification, the candidate keyword is determined to be a misidentified keyword based on the context information of the candidate keyword in the speech to be processed.
[0067] Then, when it is determined that the candidate keyword is a misidentified keyword, step S13 is executed. When it is determined that the candidate keyword is a correctly identified keyword, a speech recognition result is generated based on the candidate keyword.
[0068] When a candidate keyword is determined to be a correctly identified keyword, it means that the identified candidate keyword is accurate and there is no misidentification. Therefore, no keyword correction is needed, and the candidate keyword can be directly used as the current speech recognition result.
[0069] In some embodiments, the step of determining whether a candidate keyword is a misidentified keyword based on the context information of the candidate keyword in the speech to be processed may further include: inputting the candidate keyword and the context information of the candidate keyword in the speech to be processed into a large language model for keyword correction, and determining whether the candidate keyword is a misidentified keyword.
[0070] Among them, Large Language Model (LLM) is based on large-scale pre-trained language models and is used for semantic understanding and contextual reasoning, such as GPT model, BERT model or ERNIE model.
[0071] In some embodiments, a large language model is used to perform contextual semantic understanding and reasoning on the contextual information of candidate keywords, thereby determining whether the candidate keywords are accurately identified and whether they are misidentified keywords.
[0072] When the large language model outputs information indicating that the candidate keyword is a misidentified keyword, the above step S13 is executed; when the large language model outputs information indicating that the candidate keyword is a correctly identified keyword, a speech recognition result is generated based on the candidate keyword; when the large language model outputs information indicating that it is impossible to determine whether the candidate keyword is misidentified, a multi-round confirmation mechanism can be triggered.
[0073] For example, in a multi-round confirmation mechanism, the current speech recognition result containing candidate keywords can be provided to the user, and user feedback can be received. The accuracy of the candidate keywords can then be determined based on the user feedback. The user feedback may include keywords corrected by the user, which can replace the candidate keywords as the speech recognition result.
[0074] For example, in a multi-round confirmation mechanism, the ASR recognition process can be restarted to re-recognize keywords and correct misidentifications in the speech to be processed.
[0075] In some embodiments, after generating the current speech recognition result, the method may further include: providing feedback on the current speech recognition result to the user; receiving user feedback information; and determining whether the speech recognition result is accurate based on the user feedback information. The user feedback information may include keywords corrected by the user, indicating that the current speech recognition result is inaccurate, and the corrected keywords can replace the candidate keywords as the new speech recognition result. Furthermore, the preset keyword library and keyword correction rules can be adjusted and optimized based on the determination result.
[0076] In some embodiments, if the candidate keyword is a similar keyword and the misidentification result of the large language model is a correctly identified keyword, the preset keyword library can be optimized based on the candidate keyword, and the preset keyword library can also support personalized custom expansion keywords.
[0077] In some embodiments, the ASR model, the keyword matching model of the preset keyword library, and the large language model can be fine-tuned based on the error between the currently identified candidate keywords and the corrected keywords, so as to improve the performance of the model and enhance its adaptability to specific user groups and specific scenarios.
[0078] In practical applications, API interfaces can be provided to systems such as smart cockpits, smart homes, and cross-device interaction, allowing these systems to call and execute the voice keyword recognition method of this invention to achieve accurate voice recognition control. Furthermore, it can be combined with an NLU (Natural Language Understanding) module to perform intent parsing on the finally recognized text and execute corresponding operations, such as navigation and voice assistant interaction.
[0079] In practical applications, the speech keyword recognition method of this invention significantly reduces the misrecognition rate of speech keywords and improves the accuracy of keyword recognition by analyzing the contextual information of candidate keywords identified by ASR and automatically correcting errors. This ensures that keyword commands are accurately triggered. The method can dynamically adjust the keyword library and error correction strategy for different scenarios and product characteristics, exhibiting strong cross-scenario applicability. Leveraging the deep semantic understanding capabilities of a large-scale pre-trained language model, it can more accurately grasp the contextual semantics and achieve intelligent error correction. Through dynamic error correction and online learning optimization mechanisms, it can effectively reduce the operational error rate, improve the user interaction experience, and bring greater market competitiveness to the product.
[0080] Figure 4 This invention provides a structural block diagram of a speech keyword recognition system according to an embodiment of the present invention. The present invention also provides a speech keyword recognition system, such as... Figure 4 As shown, the speech keyword recognition system includes:
[0081] Keyword acquisition unit 401 is used to acquire candidate keywords in the speech to be processed.
[0082] The misidentification determination unit 402 is used to determine whether a candidate keyword is a misidentified keyword based on the context information of the candidate keyword in the speech to be processed.
[0083] The recognition result generation unit 403 is used to, in response to determining that the candidate keyword is a misidentified keyword, obtain the standard keyword corresponding to the candidate keyword from the preset keyword library, and generate the speech recognition result based on the standard keyword.
[0084] The speech keyword recognition system provided in this embodiment of the invention can be used to implement the speech keyword recognition method of the above embodiments. For a detailed description of the functions of each module, please refer to the relevant description in the speech keyword recognition method of the above embodiments, which will not be repeated here.
[0085] Based on the same inventive concept, embodiments of the present invention also provide an electronic device. Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Figure 5 As shown, an embodiment of the present invention provides an electronic device including: one or more processors 501, a memory 502, and one or more I / O interfaces 503. The memory 502 stores one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement any of the voice keyword recognition methods described in the above embodiments; the one or more I / O interfaces 503 are connected between the processor and the memory, configured to enable information interaction between the processor and the memory.
[0086] Among them, processor 501 is a device with data processing capabilities, including but not limited to central processing unit (CPU); memory 502 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); I / O interface (read-write interface) 503 is connected between processor 501 and memory 502, and can realize information interaction between processor 501 and memory 102, including but not limited to data bus (Bus).
[0087] In some embodiments, the processor 101, memory 502, and I / O interface 503 are interconnected via bus 504, and thus connected to other components of the computing device.
[0088] In some embodiments, the one or more processors 501 include a field-programmable gate array.
[0089] This invention also provides a computer-readable medium. The computer-readable medium stores a computer program, which, when executed by a processor, implements the steps of any of the speech keyword recognition methods described in the above embodiments. The computer-readable storage medium may be volatile or non-volatile.
[0090] This invention also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described voice keyword recognition method.
[0091] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0092] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0093] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0094] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0095] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0096] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0097] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0098] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0100] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A method for recognizing speech keywords, characterized in that, include: Obtain candidate keywords from the speech to be processed; Based on the contextual information of the candidate keywords in the speech to be processed, determine whether the candidate keywords are misidentified keywords; In response to determining that the candidate keyword is a misidentified keyword, the standard keyword corresponding to the candidate keyword is obtained from the preset keyword library, and a speech recognition result is generated based on the standard keyword.
2. The method according to claim 1, characterized in that, The process of obtaining candidate keywords in the speech to be processed includes: The speech to be processed is input into a preset speech recognition system for speech recognition to obtain the text to be processed corresponding to the speech; The text to be processed is input into the preset keyword library for keyword matching to obtain the candidate keywords; the preset keyword library contains standard keywords and similar keywords corresponding to the standard keywords.
3. The method according to claim 2, characterized in that, The step of inputting the speech to be processed into a preset speech recognition system for speech recognition to obtain the text to be processed corresponding to the speech includes: Speech features are extracted from the speech to be processed to obtain the speech features corresponding to the speech to be processed; The speech features of the speech to be processed are input into a preset speech recognition system for speech recognition to obtain the text to be processed.
4. The method according to claim 2, characterized in that, Before inputting the speech to be processed into a preset speech recognition system for speech recognition to obtain the text to be processed corresponding to the speech, the method further includes: Obtain the speech to be processed; The speech to be processed is preprocessed, including filtering and / or speech endpoint detection.
5. The method according to claim 2, characterized in that, After obtaining candidate keywords in the speech to be processed, the method further includes: In response to the candidate keyword being a standard keyword in the preset keyword library, a speech recognition result is generated based on the standard keyword. In response to the candidate keyword being a similar keyword in the preset keyword library, the step of determining whether the candidate keyword is a misidentified keyword based on the context information of the candidate keyword in the speech to be processed is performed.
6. The method according to claim 1, characterized in that, The step of determining whether a candidate keyword is a misidentified keyword based on the context information of the candidate keyword in the speech to be processed includes: The candidate keywords and their contextual information in the speech to be processed are input into a large language model for keyword correction to determine whether the candidate keywords are misidentified keywords.
7. The method according to claim 1, characterized in that, After determining whether the candidate keyword is a misidentified keyword, the method further includes: In response to determining that the candidate keyword is a correctly identified keyword, a speech recognition result is generated based on the candidate keyword.
8. A voice keyword recognition system, characterized in that, include: The keyword acquisition unit is used to acquire candidate keywords in the speech to be processed. A misidentification determination unit is used to determine whether the candidate keyword is a misidentified keyword based on the context information of the candidate keyword in the speech to be processed; The recognition result generation unit is used to, in response to determining that the candidate keyword is a misidentified keyword, obtain the standard keyword corresponding to the candidate keyword from a preset keyword library, and generate a speech recognition result based on the standard keyword.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.