A speech recognition method, device, computer equipment and storage medium
By matching the decoding diagram according to the business scenario in the speech recognition system and building a fusion decoding diagram in combination with customer hot word lists, the problems of low speech recognition accuracy and efficiency in the prior art are solved, and more efficient and accurate speech recognition is achieved.
Patent Information
- Application Number
- CN202210983356.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-08-16
AI Technical Summary
In the prior art, speech recognition has low accuracy and low recognition efficiency, especially in highly personalized application scenarios, it is difficult to meet the requirements of real-time in conversations.
By matching the corresponding service decoding diagram and static decoding diagram of speech recognition according to the business scenario, the voice to be recognized and its corresponding customer hot word list are obtained. If the customer hot word list contains customer hot words, a fusion decoding diagram is constructed as the target decoding diagram, otherwise the static decoding diagram is used for decoding.
It improves the accuracy and efficiency of speech recognition, so that the final decoding results are more in line with the customer's speaking habits, adapt to different usage scenarios, and reduces the impact on recognition efficiency.
Smart Images

Figure CN115376496B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition method, apparatus, computer equipment and storage medium. Background Art
[0002] Speech recognition technology is often used in various specific fields in production environments, such as intelligent customer service, intelligent robots and other interactive fields. Each field has its own specific proper nouns, and it is difficult for speech recognition systems in general scenarios to accurately recognize these proper nouns. Hot word enhancement refers to improving the recognition rate of hot words in speech recognition results based on the proper noun hot words provided by users.
[0003] Traditional hot word enhancement methods, faced with highly personalized application scenarios, all use scene hot word lists constructed from basic hot words in the target scenario to enhance speech recognition. However, since the hot words corresponding to different customers are somewhat different, the speech recognition accuracy is low, and the data in the scene hot word list is redundant, making it difficult to meet the real-time requirements in the conversation. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a speech recognition method, apparatus, computer device and storage medium to solve the problems of low speech recognition accuracy and low recognition efficiency in the prior art.
[0005] In order to solve the above technical problems, the present application provides a speech recognition method, which adopts the following technical solution:
[0006] Matching a business decoding graph and a static decoding graph corresponding to speech recognition according to a business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph;
[0007] Acquire a speech to be recognized and a customer hot word list corresponding to the speech to be recognized;
[0008] Decoding the speech to be recognized by using the service decoding graph to obtain a preliminary decoding result;
[0009] If the customer hot word list contains a customer hot word, construct a customer decoding graph according to the customer hot word in the customer hot word list, and construct a fused decoding graph according to the customer decoding graph and the static decoding graph, and use the fused decoding graph as a target decoding graph;
[0010] If the client hot word list does not contain the client hot word, using the static decoding graph as the target decoding graph;
[0011] The preliminary decoding result is decoded by using the target decoding graph to obtain a target decoding result.
[0012] Furthermore, before the step of matching the business decoding graph and the static decoding graph corresponding to the speech recognition according to the business scenario, it also includes:
[0013] Obtaining a fusion type of the service decoding graph and the basic decoding graph;
[0014] A specific expression of the static decoding graph is determined according to the fusion type.
[0015] Furthermore, the step of determining a specific expression of the static decoding graph according to the fusion type includes:
[0016] If the fusion type is a linear fusion type, the specific expression of the static decoding graph is C(s G (w|H), s B (w|H))=α1*s G (w|H)+β1*s B (w|H);
[0017] If the fusion type is an exponential linear fusion type, the specific expression of the static decoding graph is C(s G (w|H), s B (w|H))=-log(α1*exp(-s G (w|H))+β1*exp(s B (w|H)));
[0018] Among them, in the specific expression of the static decoding graph, α1 and β1 are variables, s G (w|H) is the score output by the service decoding graph based on the historical decoding state, s B (w|H) is the score output by the basic decoding graph based on the historical decoding state.
[0019] Furthermore, the step of decoding the preliminary decoding result by using the target decoding graph includes:
[0020] If the target decoding graph is a fusion decoding graph, the first formula of the target decoding graph is s(w|H)=-log(α2*exp(-C(s G (w|H), s B (w|H)))+β2*exp(s s (w|H))) decodes the preliminary decoding result, where α2 and β2 are variables, s s (w|H) is the score output by the client decoding graph based on the historical decoding state;
[0021] If the target decoding graph is a static decoding graph, the second formula of the target decoding graph is Decode the preliminary decoding result, where s G (w|H) is the score output by the service decoding graph based on the historical decoding state;
[0022] In the first formula and the second formula of the target decoding graph, s(w|H) is the target decoding result, C(s G (w|H), s B (w|H)) is the score output by the static decoding graph based on the historical decoding state.
[0023] Furthermore, the step of decoding the speech to be recognized by using the service decoding graph includes:
[0024] Extracting audio features from the speech to be recognized;
[0025] Converting the audio features into a phoneme sequence through an acoustic model;
[0026] The phoneme sequence is decoded by using the service decoding graph.
[0027] Furthermore, after the step of obtaining the target decoding result, the method further includes:
[0028] The new customer hot words are extracted from the target decoding result, and the new customer hot words are updated into the customer hot word table.
[0029] Furthermore, the step of updating the new customer hot word into the customer hot word table includes:
[0030] If the customer hot word list does not match the new customer hot word, adding the new customer hot word to the customer hot word list;
[0031] If the customer hot word corresponding to the new customer hot word is matched in the customer hot word table, the customer hot word corresponding to the new customer hot word in the customer hot word table is not modified.
[0032] In order to solve the above technical problems, the embodiment of the present application further provides a speech recognition device, which adopts the following technical solution:
[0033] A decoding graph matching module, used to match a business decoding graph and a static decoding graph corresponding to speech recognition according to a business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph;
[0034] An acquisition module, used for acquiring a speech to be recognized and a customer hot word list corresponding to the speech to be recognized;
[0035] A preliminary decoding module, used to decode the speech to be recognized through the service decoding graph to obtain a preliminary decoding result;
[0036] A first determination module is used for constructing a customer decoding graph according to the customer hot words in the customer hot word list if the customer hot word list contains the customer hot words, and constructing a fused decoding graph according to the customer decoding graph and the static decoding graph, and using the fused decoding graph as a target decoding graph;
[0037] A second determination module is configured to use the static decoding graph as a target decoding graph if the client hot word list does not contain the client hot word; and
[0038] The target decoding module is used to decode the preliminary decoding result through the target decoding graph to obtain a target decoding result.
[0039] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0040] The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the speech recognition method as described above when executing the computer-readable instructions.
[0041] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0042] The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps of the speech recognition method described above are implemented.
[0043] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: by matching the business decoding graph and the static decoding graph corresponding to the speech recognition according to the business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph; obtaining the speech to be recognized and the customer hot word list corresponding to the speech to be recognized; decoding the speech to be recognized by the business decoding graph to obtain a preliminary decoding result; if the customer hot word list contains customer hot words, constructing a customer decoding graph according to the customer hot words in the customer hot word list, and constructing a fusion decoding graph according to the customer decoding graph and the static decoding graph, and using the fusion decoding graph as the target decoding graph; if the customer hot word list does not contain customer hot words, using the static decoding graph as the target decoding graph; decoding the preliminary decoding result by the target decoding graph to obtain a target decoding result. In the present application, the speech to be recognized is first decoded through the business decoding graph, so that the preliminary decoding result obtained by decoding conforms to the corpus in the customer's current business scenario, thereby improving the accuracy and efficiency of speech recognition. Then, the corresponding target decoding graph is determined according to whether the customer's hot words are included in the customer's hot word list, so that the final target decoding result conforms to the customer's speaking habits, thereby further improving the accuracy of speech recognition. At the same time, due to the existence of the basic decoding graph, the customer's hot word list can be a lightweight character list, and whether the customer's hot word list includes the customer's hot words is determined to determine whether to merge and form a fused decoding graph, so as to flexibly adapt to the corresponding usage scenarios and reduce the impact on speech recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0046] Figure 2 is a flow chart of an embodiment of a speech recognition method according to the present application;
[0047] Figure 3 is a structural schematic diagram of an embodiment of a speech recognition device according to the present application;
[0048] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0050] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0051] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0052] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0053] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0054] Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, etc.
[0055] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .
[0056] It should be noted that the speech recognition method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the speech recognition device is generally arranged in the server / terminal device.
[0057] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0058] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for speech recognition according to the present application. The speech recognition method comprises the following steps:
[0059] Step S201, matching a business decoding graph and a static decoding graph corresponding to speech recognition according to a business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph.
[0060] Specifically, the above-mentioned business decoding graph is obtained by training the business text corpus according to the language model; in practical applications, multiple business decoding graphs can be pre-trained, and one business decoding graph corresponds to a business scenario.
[0061] The above-mentioned basic decoding graph uses the characters and words in the language model as a dictionary, segments each scene basic hot word in the scene basic hot word table, constructs an AC automaton, and then converts it according to a preset weight ratio; in this way, compared with the business decoding graph, the basic decoding graph contains the most characters and words.
[0062] Match the business decoding graph and basic decoding graph of current speech recognition according to the business scenario, so as to improve the accuracy and efficiency of speech recognition.
[0063] Step S202: Acquire the speech to be recognized and a customer hot word list corresponding to the speech to be recognized.
[0064] Specifically, the electronic device (eg, Figure 1The server / terminal device shown in the figure) can receive the voice to be recognized and the client hot word list corresponding to the voice to be recognized sent by the client / service end through a wired connection or a wireless connection. It should be noted that the above wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0065] The above customer hot word table corresponds to the customer of the speech to be recognized, and the customer hot word table includes N customer hot words, wherein N≥0, N is an integer; and the customer hot word representation is the customer's personalized corpus.
[0066] Step S203: decode the speech to be recognized by using the service decoding graph to obtain a preliminary decoding result.
[0067] Specifically, the business decoding graph is an FST graph defined by an operator. Based on HCLG, the speech to be recognized is decoded through the business decoding graph, and the target text corresponding to the speech to be recognized is recalled to form a preliminary decoded text (preliminary decoding result).
[0068] Step S204: if the customer hot word list contains a customer hot word, construct a customer decoding graph according to the customer hot word in the customer hot word list, and construct a fused decoding graph according to the customer decoding graph and the static decoding graph, and use the fused decoding graph as the target decoding graph.
[0069] Specifically, each customer hot word in the customer hot word list is decomposed into characters and words through a language model (such as an N-garm language model), and an AC automaton is constructed based on the decomposed characters and words of each customer hot word, and then the AC automaton is converted into a customer decoding graph according to a preset weight relationship. The preset weight relationship is represented by the weight of the characters and words in the customer hot word. For example, in the preset weight relationship, the weight of "know" is higher than the weight of "daozhi".
[0070] The above-mentioned customer hot word list is not empty, indicating that the customer hot word list includes at least one customer hot word; a fused decoding graph is constructed through the customer decoding graph and the static decoding graph, so that in the actual decoding process, the fused decoding graph can be decoded according to the customer's personalization, thereby improving the accuracy of speech recognition.
[0071] Step S205: if the client hot word list does not contain the client hot word, determining the static decoding graph as a target decoding graph.
[0072] Specifically, the above-mentioned customer hot word list is empty, which means that the customer hot word list does not include customer hot words; at this time, there is no need to construct a customer decoding graph, and to merge the customer decoding graph and the static decoding graph. It is only necessary to decode the preliminary decoding result through the static decoding graph to flexibly adapt to the corresponding usage scenarios and reduce the impact on speech recognition efficiency.
[0073] Step S206, decoding the preliminary decoding result through the target decoding graph to obtain a target decoding result.
[0074] Specifically, in practical applications, based on the weighted finite-state transducer WFST (weighted finite-state transducer), the preliminary decoding result is decoded through the target decoding graph, and the weight proportion of each decoding graph is comprehensively considered to obtain the text with the highest score to form the target decoding result, and then the text information is determined according to the target decoding result.
[0075] In the present application, the speech to be recognized is first decoded through the business decoding graph, so that the preliminary decoding result obtained by decoding conforms to the corpus in the customer's current business scenario, thereby improving the accuracy and efficiency of speech recognition. Then, the corresponding target decoding graph is determined according to whether the customer's hot words are included in the customer's hot word list, so that the final target decoding result conforms to the customer's speaking habits, thereby further improving the accuracy of speech recognition. At the same time, due to the existence of the basic decoding graph, the customer's hot word list can be a lightweight character list, and whether the customer's hot word list includes the customer's hot words is determined to determine whether to merge and form a fused decoding graph, so as to flexibly adapt to the corresponding usage scenarios and reduce the impact on speech recognition efficiency.
[0076] In some optional implementations of this embodiment, before the step S201 of matching the business decoding graph and the static decoding graph corresponding to the speech recognition according to the business scenario, the method further includes:
[0077] Obtaining a fusion type of the service decoding graph and the basic decoding graph;
[0078] A specific expression of the static decoding graph is determined according to the fusion type.
[0079] Specifically, the fusion type includes a linear fusion type (log-linear, LL) and an exponential linear fusion type (linear, LIN), wherein the exponential linear fusion type (LIN) is calculated relative to the linear fusion type (LL) to obtain C(s G (w|H), s B (w|H)) results are more accurate.
[0080] In some optional implementations of this embodiment, the step of determining a specific expression of the static decoding graph according to the fusion type includes:
[0081] If the fusion type is a linear fusion type, the specific expression of the static decoding graph is C(s G (w|H), s B (w|H))=α1*s G (w|H)+β1*s B (w|H);
[0082] If the fusion type is an exponential linear fusion type, the specific expression of the static decoding graph is C(s G (w|H), s B (w|H))=-log(α1*exp(-s G (w|H))+β1*exp(s B (w|H)));
[0083] Among them, in the specific expression of the static decoding graph, α1 and β1 are variables, s G (w|H) is the score output by the service decoding graph based on the historical decoding state, s B (w|H) is the score output by the basic decoding graph based on the historical decoding state.
[0084] Specifically, α1 and β1 are variables, and the sum of α1 and β1 is 1. In this way, s can be adjusted according to the actual situation by adjusting the size of α1 and β1. G (w|H) and s B (w|H) in C(s G (w|H), s B (w|H)) in the specific expression; if α1 is greater than β1, it is characterized by the final passing of C(s G (w|H), s B The result of calculating the specific expression of (w|H) is the speaking habits in the current business scenario.
[0085] In some optional implementations of this embodiment, in the above step S205, the step of decoding the preliminary decoding result by using the target decoding graph includes:
[0086] If the target decoding graph is a fusion decoding graph, the first formula of the target decoding graph is s(w|H)=-log(α2*exp(-C(s G (w|H), s B (w|H)))+β2*exp(s sDecode the preliminary decoding result, where both α2 and β2 are variables, and s s (w|H) is the score output by the customer decoding graph based on the historical decoding state;
[0087] If the target decoding graph is a static decoding graph, use the second formula of the target decoding graph to decode the preliminary decoding result, where s G (w|H) is the score output by the service decoding graph based on the historical decoding state;
[0088] Among them, in the first formula and the second formula of the target decoding graph, s(w|H) is the target decoding result, and C(s G (w|H), s B (w|H)) is the score output by the static decoding graph based on the historical decoding state.
[0089] Specifically, in the first formula of the target decoding graph, both α2 and β2 are variables, and the sum of α2 and β2 is 1. In this way, according to the actual situation, the weights of C(s G (w|H), s B (w|H)) and s s (w|H) in the first formula of the target decoding graph can be adjusted; for example, when α2 is less than β2, it means that the s(w|H) finally calculated by the first formula of the target decoding graph is more in line with the customer's speaking habit.
[0090] In the second formula of the target decoding graph, B represents the dictionary of the language model, and the basic decoding graph is constructed from this dictionary of the language model; if then it means that the phrase (w|H) does not exist in the dictionary of the language model. At this time, s(w|H) = s G (w|H); on the contrary, if if(w|H) ∈ B, it means that the phrase (w|H) is included in the dictionary of the language model. At this time, s(w|H) = C(s G (w|H), s B (w|H)).
[0091] For example, when w in (w|H) is "service" and H is "industry", judge whether the word "service" is included in the dictionary of the language model. If so, s(w|H) = s G (w|H), if not, then if(w|H) ∈ B, s(w|H) = C(s G (w|H), s B (w|H)).
[0092] In some optional implementations of this embodiment, in the above step S203, the step of decoding the speech to be recognized by using the service decoding graph includes:
[0093] Extracting audio features from the speech to be recognized;
[0094] Converting the audio features into a phoneme sequence through an acoustic model;
[0095] The phoneme sequence is decoded by using the service decoding graph.
[0096] Specifically, after obtaining the speech to be recognized, at least one audio feature (Mel-frequency cepstral coefficient (MFCC)) is first extracted from the speech to be recognized, and then each audio feature in the speech to be recognized is converted into a state sequence / phoneme sequence through a pre-trained acoustic model, and then the phoneme sequence is decoded through a business decoding graph.
[0097] In some optional implementations of this embodiment, the above step S206, after the step of obtaining the target decoding result, further includes:
[0098] The new customer hot words are extracted from the target decoding result, and the new customer hot words are updated into the customer hot word table.
[0099] Specifically, each time the target decoding result is obtained, word segmentation processing will be performed on the text information in the target decoding result to extract new customer hot words to replace the customer hot word list and improve the customer hot word list, which effectively improves the accuracy of subsequent speech recognition and enhances customer experience.
[0100] It should be noted that after the text information is segmented, keywords can be determined for the new customer hot words obtained from each segmentation, and the score of each new customer hot word can be determined according to the preset mapping relationship. New customer hot words that are keywords are determined based on the score of each new customer hot word, and the new customer hot words that are keywords are added to the customer hot word list; this can avoid the customer hot word list being too redundant in subsequent voice recognition, thereby ensuring the accuracy of voice recognition and improving the efficiency of voice recognition.
[0101] In some optional implementations of this embodiment, the step of updating the new customer hot word into the customer hot word table includes:
[0102] If the customer hot word list does not match the new customer hot word, adding the new customer hot word to the customer hot word list;
[0103] If the customer hot word corresponding to the new customer hot word is matched in the customer hot word table, the customer hot word corresponding to the new customer hot word in the customer hot word table is not modified.
[0104] Specifically, when there is no matching customer hot word corresponding to a new customer hot word in the customer hot word list, it is characterized as that the new customer hot word is not included in the customer hot word list. At this time, by adding the new customer hot word to the customer hot word list, the customer hot word list is further optimized to improve the adaptability of the customer hot word list, thereby effectively ensuring the accuracy of speech recognition.
[0105] When a customer hot word corresponding to a new customer hot word is matched in the customer hot word table, it is characterized that the customer hot word table contains the new customer hot word, and the customer hot word table is not modified at this time.
[0106] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned static decoding map and customer decoding map, the above-mentioned static decoding map and customer decoding map information can also be stored in a node of a blockchain.
[0107] The blockchain referred to in this application is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0108] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0109] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0110] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0111] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0112] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a speech recognition device, and the embodiment of the device is Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0113] like Figure 3 As shown, the speech recognition device 300 described in this embodiment includes: a decoding graph matching module 301, an acquisition module 302, a preliminary decoding module 303, a first determination module 304, a second determination module 305 and a target decoding module 306. Among them:
[0114] A decoding graph matching module 301 is used to match a business decoding graph and a static decoding graph corresponding to speech recognition according to a business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph;
[0115] An acquisition module 302 is used to acquire a speech to be recognized and a customer hot word list corresponding to the speech to be recognized;
[0116] A preliminary decoding module 303, configured to decode the speech to be recognized through the service decoding graph to obtain a preliminary decoding result;
[0117] A first determination module 304 is configured to construct a customer decoding graph according to the customer hot words in the customer hot word list if the customer hot word list contains the customer hot words, and to construct a fused decoding graph according to the customer decoding graph and the static decoding graph, and use the fused decoding graph as a target decoding graph;
[0118] A second determination module 305 is configured to use the static decoding graph as a target decoding graph if the customer hot word list does not contain the customer hot word;
[0119] The target decoding module 306 is used to decode the preliminary decoding result through the target decoding graph to obtain a target decoding result.
[0120] In the present application, the speech to be recognized is first decoded through the business decoding graph, so that the preliminary decoding result obtained by decoding conforms to the corpus in the customer's current business scenario, thereby improving the accuracy and efficiency of speech recognition. Then, the corresponding target decoding graph is determined according to whether the customer's hot words are included in the customer's hot word list, so that the final target decoding result conforms to the customer's speaking habits, thereby further improving the accuracy of speech recognition. At the same time, due to the existence of the basic decoding graph, the customer's hot word list can be a lightweight character list, and whether the customer's hot word list includes the customer's hot words is determined to determine whether to merge and form a fused decoding graph, so as to flexibly adapt to the corresponding usage scenarios and reduce the impact on speech recognition efficiency.
[0121] In some optional implementations of this embodiment, a type acquisition module and a third determination module are also included.
[0122] A type acquisition module, used to acquire the fusion type of the business decoding graph and the basic decoding graph;
[0123] The third determination module is used to determine a specific expression of the static decoding graph according to the fusion type.
[0124] In some optional implementations of this embodiment, the third determination module includes a first determination submodule and a second determination submodule; wherein:
[0125] The first determination submodule is used for, if the fusion type is a linear fusion type, the specific expression of the static decoding graph is C(s G (w|H), s B (w|H))=α1*s G (w|H)+β1*s B (w|H);
[0126] The second determination submodule is used for, if the fusion type is a logarithmic linear fusion type, the specific expression of the static decoding graph is C(s G (w|H), s B(w|H))=-log(α1*exp(-s G (w|H))+β1*exp(s B (w|H))).
[0127] In the above specific expressions of the static decoding graph, α1 and β1 are variables, s G (w|H) is the score output by the service decoding graph based on the historical decoding state, s B (w|H) is the score output by the basic decoding graph based on the historical decoding state.
[0128] In some optional implementations of this embodiment, the target decoding module 306 includes a first decoding submodule and a second decoding submodule.
[0129] The first decoding submodule is used for, if the target decoding graph is a fusion decoding graph, to decode the target decoding graph by the first formula s(w|H)=-log(α2*exp(-C(s G (w|H), s B (w|H)))+β2*exp(s s (w|H))) decodes the preliminary decoding result, where α2 and β2 are variables, s s (w|H) is the score output by the client decoding graph based on the historical decoding state;
[0130] The second decoding submodule is used for, if the target decoding graph is a static decoding graph, to obtain the target decoding graph by the second formula of the target decoding graph. Decode the preliminary decoding result, where s G (w|H) is the score output by the service decoding graph based on the historical decoding status.
[0131] In the first and second formulas of the target decoding graph, s(w|H) is the target decoding result, C(s G (w|H), s B (w|H)) is the score output by the static decoding graph based on the historical decoding state.
[0132] In some optional implementations of this embodiment, the preliminary decoding module 203 includes a feature extraction submodule, a sequence conversion submodule and a sequence decoding submodule.
[0133] A feature extraction submodule, used to extract audio features from the speech to be recognized;
[0134] A sequence conversion submodule, used for converting the audio features into a phoneme sequence through an acoustic model;
[0135] The sequence decoding submodule is used to decode the phoneme sequence through the service decoding graph.
[0136] In some optional implementations of this embodiment, a hot word update module is also included.
[0137] A hot word updating module is used to extract new customer hot words from the target decoding result and update the new customer hot words into the customer hot word table.
[0138] In some optional implementations of this embodiment, the hot word update module includes a first update submodule and a second update submodule.
[0139] A first updating submodule, configured to add the new customer hot word to the customer hot word list if there is no customer hot word in the customer hot word list that matches the new customer hot word;
[0140] The second updating submodule is configured to not modify the customer hot word in the customer hot word table that corresponds to the new customer hot word if the customer hot word corresponding to the new customer hot word is matched in the customer hot word table.
[0141] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0142] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.
[0143] The computer device may be a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with the user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.
[0144] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the speech recognition method, etc. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0145] The processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for running the speech recognition method.
[0146] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0147] In the present application, the speech to be recognized is first decoded through the business decoding graph, so that the preliminary decoding result obtained by decoding conforms to the corpus in the customer's current business scenario, thereby improving the accuracy and efficiency of speech recognition. Then, the corresponding target decoding graph is determined according to whether the customer's hot words are included in the customer's hot word list, so that the final target decoding result conforms to the customer's speaking habits, thereby further improving the accuracy of speech recognition. At the same time, due to the existence of the basic decoding graph, the customer's hot word list can be a lightweight character list, and whether the customer's hot word list includes the customer's hot words is determined to determine whether to merge and form a fused decoding graph, so as to flexibly adapt to the corresponding usage scenarios and reduce the impact on speech recognition efficiency.
[0148] The present application also provides another implementation, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the speech recognition method as described above.
[0149] In the present application, the speech to be recognized is first decoded through the business decoding graph, so that the preliminary decoding result obtained by decoding conforms to the corpus in the customer's current business scenario, thereby improving the accuracy and efficiency of speech recognition. Then, the corresponding target decoding graph is determined according to whether the customer's hot words are included in the customer's hot word list, so that the final target decoding result conforms to the customer's speaking habits, thereby further improving the accuracy of speech recognition. At the same time, due to the existence of the basic decoding graph, the customer's hot word list can be a lightweight character list, and whether the customer's hot word list includes the customer's hot words is determined to determine whether to merge and form a fused decoding graph, so as to flexibly adapt to the corresponding usage scenarios and reduce the impact on speech recognition efficiency.
[0150] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0151] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.
Claims
1. A speech recognition method, characterized in that: The steps include: Matching a business decoding graph and a static decoding graph corresponding to speech recognition according to a business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph; Acquire a speech to be recognized and a customer hot word list corresponding to the speech to be recognized; Decoding the speech to be recognized by using the service decoding graph to obtain a preliminary decoding result; If the customer hot word list contains a customer hot word, construct a customer decoding graph according to the customer hot word in the customer hot word list, and construct a fused decoding graph according to the customer decoding graph and the static decoding graph, and use the fused decoding graph as a target decoding graph; If the client hot word list does not contain the client hot word, using the static decoding graph as the target decoding graph; The preliminary decoding result is decoded by using the target decoding graph to obtain a target decoding result.
2. The speech recognition method according to claim 1, characterized in that: Before the step of matching the business decoding graph and the static decoding graph corresponding to the speech recognition according to the business scenario, the method further includes: Obtaining a fusion type of the service decoding graph and the basic decoding graph; A specific expression of the static decoding graph is determined according to the fusion type.
3. The speech recognition method according to claim 2, characterized in that: The step of determining the specific expression of the static decoding graph according to the fusion type comprises: If the fusion type is a linear fusion type, the specific expression of the static decoding graph is: ; If the fusion type is an exponential linear fusion type, the specific expression of the static decoding graph is: ; Among them, the specific expression of the static decoding graph is and are variables, is the score outputted by the service decoding graph based on the historical decoding state, It is the score output by the basic decoding graph based on the historical decoding status.
4. The speech recognition method according to claim 3, characterized in that: The step of decoding the preliminary decoding result by using the target decoding graph comprises: If the target decoding graph is a fused decoding graph, the first formula of the target decoding graph is Decoding the preliminary decoding result, wherein and are variables, A score outputted for the client decoding graph based on a historical decoding state; If the target decoding graph is a static decoding graph, the second formula of the target decoding graph is The preliminary decoding result is decoded, wherein: The dictionary represented as a language model, if , then the dictionary represented as the language model does not contain The phrase, at this time On the contrary, if , then the dictionary represented by the language model contains The phrase, at this time ; Among them, in the first formula and the second formula of the target decoding graph, Decoding result for the target, It is the score output by the static decoding graph based on the historical decoding status.
5. The speech recognition method according to any one of claims 1 to 4, characterized in that: The step of decoding the speech to be recognized by using the service decoding graph comprises: Extracting audio features from the speech to be recognized; Converting the audio features into a phoneme sequence through an acoustic model; The phoneme sequence is decoded by using the service decoding graph.
6. The speech recognition method according to any one of claims 1 to 4, characterized in that: After the step of obtaining the target decoding result, the method further includes: The new customer hot words are extracted from the target decoding result, and the new customer hot words are updated into the customer hot word table.
7. The speech recognition method according to claim 6, characterized in that: The step of updating the new customer hot word into the customer hot word table comprises: If the customer hot word list does not match the new customer hot word, adding the new customer hot word to the customer hot word list; If the customer hot word corresponding to the new customer hot word is matched in the customer hot word table, the customer hot word corresponding to the new customer hot word in the customer hot word table is not modified.
8. A speech recognition device, characterized in that: include: A decoding graph matching module, used to match a business decoding graph and a static decoding graph corresponding to speech recognition according to a business scenario, wherein the static decoding graph is constructed by the business decoding graph and the basic decoding graph; An acquisition module, used for acquiring a speech to be recognized and a customer hot word list corresponding to the speech to be recognized; A preliminary decoding module, used to decode the speech to be recognized through the service decoding graph to obtain a preliminary decoding result; A first determination module is used for constructing a customer decoding graph according to the customer hot words in the customer hot word list if the customer hot word list contains the customer hot words, and constructing a fused decoding graph according to the customer decoding graph and the static decoding graph, and using the fused decoding graph as a target decoding graph; A second determination module is configured to use the static decoding graph as a target decoding graph if the client hot word list does not contain the client hot word; and The target decoding module is used to decode the preliminary decoding result through the target decoding graph to obtain a target decoding result.
9. A computer device, comprising a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the speech recognition method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the speech recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Voice recognition method, device, equipment, system and storage medium
CN113436614A
Speech recognition method and device, equipment and storage medium
CN114360499A