Voice dialogue method, device, computer-readable storage medium, and electronic device
The voice interaction method improves efficiency and reduces delays in multi-channel scenarios by using a preset voice recognition model and caching recognition data, addressing inefficiencies in current voice interaction technologies.
Patent Information
- Application Number
- JP2022558093
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-16
- Filing Date
- 2022-02-16
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2042-02-16
AI Technical Summary
Current voice interaction technologies are inefficient in processing multi-channel voice signals, leading to low speech recognition efficiency and large dialogue delays, especially in multi-user scenarios.
A voice interaction method that utilizes a preset voice recognition model to recognize audio signals, extracts stored recognition data from a cache, and generates part of the recognition result using the model, thereby improving processing efficiency and reducing resource consumption.
Enhances processing efficiency and reduces delays in multi-channel voice interaction scenarios by effectively reusing stored recognition data, meeting requirements for high efficiency and individualization.
Smart Images

Figure 0007792912000001 
Figure 0007792912000002 
Figure 0007792912000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to a Chinese patent application bearing application number 202110279812.4 and entitled "Voice dialogue method, device, computer-readable storage medium and electronic device" filed with the State Intellectual Property Office of the People's Republic of China on March 16, 2021, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates to the field of computer technology, and in particular to a voice interaction method, an apparatus, a computer-readable storage medium, and an electronic device. [Background technology]
[0003] With the continuous advancement of artificial intelligence technology, human-computer interaction has also made great progress. Intelligent voice dialogue technology can be applied to a variety of devices, such as automobiles, robots, home appliances, central control systems, access control systems, and ATM machines.
[0004] For example, in an in-vehicle voice dialogue scenario, a voice dialogue system generally receives only one channel of voice signal, processes the voice signal, and then provides feedback to the user. With the development of artificial intelligence technology, voice dialogue systems are developing toward becoming more efficient, intelligent, and personalized. Summary of the Invention [Problem to be solved by the invention]
[0005] SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a voice interaction method, apparatus, computer-readable storage medium, and electronic device. [Means for solving the problem]
[0006] An embodiment of the present disclosure provides a voice interaction method, the method including: acquiring at least one channel of audio signals; recognizing the at least one channel of audio signals using a preset voice recognition model to obtain a first class of recognition results using the voice recognition model; determining stored recognition data for the at least one channel of audio signals from a cache; generating a second class of recognition results based on the stored recognition data; processing the first class of recognition results and the second class of recognition results using the voice recognition model to obtain phrase recognition results respectively corresponding to the at least one channel of audio signals; performing semantic analysis on the phrase recognition results to obtain at least one analysis result; and generating an instruction to control a voice interaction device to perform a corresponding function based on the at least one analysis result.
[0007] According to another aspect of an embodiment of the present disclosure, there is provided a voice interaction device, the device including: an acquisition module for acquiring at least one channel of audio signal; a recognition module for recognizing the at least one channel of audio signal using a preset voice recognition model and obtaining a first class of recognition result using the voice recognition model; a determination module for determining stored recognition data for the at least one channel of audio signal from a cache; a first generation module for generating a second class of recognition result based on the stored recognition data; a processing module for processing the first class recognition result and the second class recognition result using the voice recognition model to obtain phrase recognition results respectively corresponding to the at least one channel of audio signal; an analysis module for performing semantic analysis on each phrase recognition result respectively to obtain at least one analysis result; and a second generation module for generating an instruction to control the voice interaction device to perform a corresponding function based on the at least one analysis result.
[0008] According to another aspect of an embodiment of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program for executing the above-described voice interaction method.
[0009] According to another aspect of an embodiment of the present disclosure, there is provided an electronic device, the electronic device including a processor and a memory for storing instructions executable by the processor, the processor being adapted to read the executable instructions from the memory and execute the instructions to realize the above-described voice interaction method.
[0010] Based on the voice interaction method, device, computer-readable storage medium, and electronic device provided in the above embodiments, the present disclosure recognizes at least one channel of audio signal using a preset voice recognition model, and during recognition, extracts stored recognition data from a cache to generate part of the recognition result, and generates another part of the recognition result using the voice recognition model, thereby effectively reusing the stored recognition data and eliminating the need to process the entire amount of data using the voice recognition model, thereby improving the processing efficiency of at least one channel of audio signal, and helping to still meet the requirements of low resource consumption and low processing delay even in multi-channel voice interaction scenarios.
[0011] The technical solutions of the present disclosure are described in detail below with reference to the drawings and examples. The above and other objects, features, and advantages of the present disclosure will become apparent from a more detailed description of the embodiments of the present disclosure taken in conjunction with the drawings. The drawings are used to provide a further understanding of the embodiments of the present disclosure, constitute a part of this specification, and are used to interpret the present disclosure together with the embodiments of the present disclosure, but are not intended to limit the present disclosure. In the drawings, the same reference numerals generally indicate the same elements or steps. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram of a system applicable to the present disclosure. [Figure 2]1 is a schematic flowchart of a voice interaction method provided in an exemplary embodiment of the present disclosure. [Figure 3] 1 is a schematic flowchart of a voice interaction method provided in another exemplary embodiment of the present disclosure. [Figure 4] 1 is a schematic flowchart of a voice interaction method provided in another exemplary embodiment of the present disclosure. [Figure 5] 1 is a schematic flowchart of a voice interaction method provided in another exemplary embodiment of the present disclosure. [Figure 6] 1 is a schematic flowchart of a voice interaction method provided in another exemplary embodiment of the present disclosure. [Figure 7] 1 is a schematic diagram of an application scene of a voice interaction method according to an embodiment of the present disclosure; [Figure 8] 1 is a schematic structural diagram of a voice dialogue device provided in an exemplary embodiment of the present disclosure; [Figure 9] FIG. 2 is a schematic structural diagram of a voice dialogue device provided in another exemplary embodiment of the present disclosure; [Figure 10] 1 is a structural diagram of an electronic device provided in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It is clear that the described embodiments are only some of the embodiments of the present disclosure, and are not all of the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0014] Unless otherwise specified, the relative arrangement of components and steps, formulas and values described in these examples do not limit the scope of the present disclosure.
[0015] Those skilled in the art will understand that the terms "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, modules, etc., and do not represent any specific technical meaning or a necessary logical order therebetween.
[0016] It should further be understood that in the embodiments of the present disclosure, "plurality" means two or more than two, and "at least one" means one, two, or more than two.
[0017] Furthermore, it should be understood that any single element, data, or structure referred to in the embodiments of the present disclosure can generally be understood to be one or more unless expressly limited or the context suggests to the contrary.
[0018] Furthermore, in this disclosure, the term "and / or" is merely used to describe the relationship between related objects and indicates that three types of relationships exist, for example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, in this disclosure, the symbol " / " generally indicates that the related objects before and after are in an "or" relationship.
[0019] Furthermore, it should be understood that the description of each embodiment of the present disclosure emphasizes the differences between each embodiment, and that the same or similar aspects may be referenced, and detailed descriptions have been omitted for the sake of brevity.
[0020] It should also be understood that the dimensions of each part shown in the drawings are not drawn to scale for the sake of convenience.
[0021] The following description of at least one exemplary embodiment is merely exemplary in nature and is in no way limiting of the present disclosure and its applications or uses.
[0022] Techniques, methods and equipment known to those of ordinary skill in the relevant art will not be described in detail, but where appropriate, said techniques, methods and equipment should be considered part of the specification.
[0023] It should be noted that in the following figures, like numerals and letters represent like items, so that once an item is defined in one figure, it does not require further explanation in subsequent figures.
[0024] Embodiments of the present disclosure may be applied to electronic devices, such as terminal devices, computer systems, servers, and the like, that can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices, such as terminal devices, computer systems, servers, and the like, include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, networked personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments that include any one of the above systems.
[0025] Electronic devices such as terminal devices, computer systems, and servers may be described in the general context of computer system-executable instructions (e.g., program modules) executed by a computer system. Generally, program modules may include routines, programs, target programs, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer systems / servers may be implemented in distributed cloud computing environments in which tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located on storage media of local or remote computing systems, including storage devices.
[0026] Summary of the application Current voice interaction technology is usually unable to process multi-channel voice signals simultaneously, but can only process one-channel voice signals simultaneously, which means it cannot meet the requirements of multi-user, personalized voice recognition. Therefore, the technical solution disclosed herein needs to apply voice interaction technology to the scenario of recognizing multi-channel voice.
[0027] Currently, speech recognition models require processing of all data for speech signals, resulting in low speech recognition efficiency and large dialogue delays. This makes it impossible to meet the requirements for high efficiency and individualization for multi-user speech dialogue systems, especially in multi-channel speech recognition scenarios.
[0028] Exemplary System FIG. 1 illustrates an exemplary system architecture 100 to which a voice interaction method or voice interaction device according to an embodiment of the present disclosure can be applied.
[0029] 1, system architecture 100 may include terminal equipment 101, network 102, and server 103. Network 102 is a medium for providing a communication link between terminal equipment 101 and server 103. Here, network 102 may include various connection types, such as, but not limited to, wired, wireless communication links, or fiber optic cables.
[0030] A user can use terminal device 101 to interact with server 103 via network 102 to send and receive messages, etc. Various communication client applications may be installed on terminal device 101, such as a voice recognition application, a multimedia application, a search application, a web browser application, a shopping application, an instant messenger, etc.
[0031] The terminal device 101 may be an electronic device, including, but not limited to, mobile terminals such as in-vehicle terminals, mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants, mobile information terminals), PADs (tablet computers), and PMPs (Portable Media Players, portable multimedia players), as well as fixed terminals such as digital TVs, desktop computers, and smart home appliances.
[0032] The server 103 may be a device that can provide various service functions, such as a background speech recognition server that recognizes audio signals uploaded from the terminal device 101. The background speech recognition server can process the received audio to obtain instructions for controlling a voice interaction device and feed the instructions back to the terminal device 101.
[0033] In addition, the voice dialogue method provided in the embodiments of the present disclosure may be executed by the server 103 or by the terminal device 101, and correspondingly, the voice dialogue device may be installed in the server 103 or in the terminal device 101.
[0034] It should be understood that the number of terminal devices 101, networks 102, and servers 103 in Fig. 1 is merely exemplary. Any number of terminal devices, networks, and / or servers may be arranged depending on implementation requirements, and the present application is not limited thereto. Also, if there is no need to obtain audio signals remotely, the above system architecture may not include the network 102 and may include only the server or terminal devices. For example, if the terminal device 101 and the server 103 are connected in a wired manner, the network 102 may be omitted.
[0035] Exemplary Methods 2 is a schematic flowchart of a voice interaction method provided in an exemplary embodiment of the present disclosure. The method of this embodiment can be applied to an electronic device (the terminal device 101 or the server 103 shown in FIG. 1). As shown in FIG. 2, the method includes the following steps:
[0036] In step 201, at least one channel of audio signal is acquired.
[0037] In this embodiment, the electronic device can acquire at least one channel of audio signals locally or remotely. For example, when this embodiment is applied to an in-vehicle voice recognition scenario, the at least one channel of audio signals can be at least one voice signal of a passenger in the vehicle collected by at least one microphone mounted on the vehicle.
[0038] In step 202, a preset speech recognition model is utilized to recognize at least one channel of audio signal, and a first class recognition result is obtained by the speech recognition model.
[0039] In this embodiment, the electronic device can recognize at least one channel of audio signals using a preset speech recognition model, and during recognition, the preset speech recognition model can obtain a first-class recognition result. Here, the preset speech recognition model may be a model obtained by prior training using a large number of speech signal samples. The preset speech recognition model is used to recognize the at least one channel of input audio signals and obtain at least one phrase recognition result.
[0040] Typically, a preset speech recognition model may include multiple sub-models, such as a phonetics sub-model, a linguistic sub-model, a decoding network sub-model, etc. Furthermore, the phonetics sub-model is used to perform syllable segmentation on an audio signal, the linguistic sub-model is used to convert each syllable into a word, and the decoding network sub-model is used to select an optimal combination from multiple words to obtain a sentence.
[0041] In the above step 202, in the process of recognizing at least one channel of audio signal using a preset voice recognition model, the electronic device will usually first search the cache for recognition data corresponding to the current processing stage. If the corresponding recognition data does not exist in the cache, the electronic device will execute the above step 202 to obtain the recognition data, and the recognition data will be used as the first-class recognition result.
[0042] In step 203, recognition data for at least one channel of stored audio signals is determined from a cache.
[0043] In this embodiment, the electronic device can determine stored recognition data for at least one channel of audio signals from the cache. Typically, in the recognition process of the above-mentioned voice recognition model, the electronic device will first search the cache for whether recognition data corresponding to the current processing stage exists, and if so, extract the recognition data.
[0044] In step 204, a second class of recognition results is generated based on the stored recognition data.
[0045] In this embodiment, the electronic device can generate a second class of recognition result based on the stored recognition data extracted in step 203. For example, the stored recognition data can be the second class of recognition result, or the second class of recognition result can be obtained after processing the stored recognition data in a certain manner, where the certain processing includes proportional scaling, normalization, etc. of the recognition data.
[0046] The first class recognition results and the second class recognition results are usually intermediate results obtained during speech recognition model processing, such as probability scores of syllables and probability scores of words.
[0047] In step 205, the speech recognition model is utilized to process the first class recognition results and the second class recognition results to obtain phrase recognition results respectively corresponding to the at least one channel audio signal.
[0048] In this embodiment, the electronic device uses a speech recognition model to process the first-class recognition results and the second-class recognition results to obtain phrase recognition results corresponding to at least one channel of audio signals. Usually, the first-class recognition results and the second-class recognition results are intermediate results obtained by processing the speech recognition model, so the first-class recognition results and the second-class recognition results need to be further processed using the speech recognition model.
[0049] For example, the first-class and second-class recognition results may include a probability score for each syllable and a probability score for each word obtained by recognizing an audio signal, and the speech recognition model may use a path search algorithm (e.g., a Viterbi algorithm) to determine an optimal path from the recognized words corresponding to one audio signal, and obtain a sentence according to the optimal path as a phrase recognition result. Here, a one-channel audio signal can correspond to one phrase recognition result, and a multi-channel audio signal can correspond to a multi-channel phrase recognition result.
[0050] In step 206, semantic analysis is performed on each of the phrase recognition results to obtain at least one analysis result.
[0051] In this embodiment, the electronic device performs semantic analysis on each result of at least one phrase recognition result to obtain at least one analysis result. Here, each analysis result of the at least one analysis result corresponds to one audio signal. Here, the at least one analysis result may be structured data. For example, if the phrase recognition result is "set the air conditioner temperature to 25 degrees," the corresponding analysis result is "domain=car control, intent=air conditioner temperature setting, slot position=<temperature value=25>."
[0052] As a method for analyzing the meaning of the word recognition results, for example, a rule engine, a neural network engine, or the like can be used.
[0053] In step 207, instructions are generated based on at least one analysis result to control the voice interaction device to perform the corresponding function.
[0054] In this embodiment, the electronic device can generate an instruction to control the voice interaction device to execute a corresponding function based on at least one analysis result. Here, the above voice interaction device may be the above electronic device for executing the voice interaction method of the present disclosure, or may be an electronic device communicatively connected to the above electronic device. For example, if the voice interaction device is an in-vehicle air conditioner, and the analysis result is "Domain=Car control, Intention=Air conditioner temperature setting, Slot position=<Temperature value=25>", an instruction can be generated to control the in-vehicle air conditioner to set a predetermined preset temperature, where the predetermined preset temperature is 25°C.
[0055] The method provided in the embodiments of the present disclosure recognizes at least one channel of audio signal using a preset speech recognition model, and during the recognition, extracts stored recognition data from a cache to generate part of the recognition result, and another part of the recognition result is generated by the speech recognition model, thereby effectively reusing the stored recognition data and eliminating the need to process the entire amount of data in the speech recognition model, and further improving the processing efficiency for at least one channel of audio signal, and meeting the requirements of low resource consumption and low processing delay for electronic devices in multi-channel speech interaction scenarios.
[0056] In some optional embodiments, the electronic device can also store the recognition data obtained during the recognition process of the preset speech recognition model in a cache. Specifically, if the recognition data corresponding to a certain recognition step does not exist in the cache, the speech recognition model needs to perform the recognition step and store the obtained recognition data in the cache, which makes it easy to reuse the recognition data later.
[0057] In this embodiment, the recognition data obtained in the recognition process of the speech recognition model is stored in a cache, so that the recognition data can be reused and the recognition data in the cache is continuously updated. In addition, by using more stored recognition data in the model recognition process, the efficiency of speech recognition is further improved.
[0058] In some alternative embodiments, the specific execution process of the above step 201 is as follows:
[0059] First, an initial audio signal collected by an audio collecting device is received.
[0060] Here, the number of the audio collecting devices may be one or more, and each may be used to collect at least one channel of initial audio signal. The initial audio signal may be a signal obtained by the audio collecting device collecting the voice of at least one user. For example, a vehicle may include multiple audio collecting devices, each installed around a seat in the vehicle, and each audio collecting device may be used to collect the voice of a passenger in the corresponding seat. In this case, the collected audio signal typically includes a mixed voice signal of multiple users.
[0061] Next, a sound source separation process is performed on the initial audio signal to obtain an audio signal of at least one channel.
[0062] Here, the above-mentioned sound source separation processing method may employ existing technology, such as a blind source separation (BSS) algorithm, to separate the voice signals of multiple users, with each resulting channel of audio signal corresponding to one user. In an in-vehicle voice interaction scenario, the sound source separation processing can further associate each resulting channel of audio signal with a corresponding audio collecting device. Since each audio collecting device is installed near a corresponding seat, each resulting channel of audio signal can be associated with a corresponding seat. The sound source separation technology can separate the voice signals of multiple users and establish a one-to-one correspondence with different audio collecting devices. The specific implementation process can refer to conventional technical methods, and a detailed description will be omitted in this embodiment.
[0063] In this embodiment, the voices of multiple users can be separated by performing sound source separation on the initial audio signal, thereby associating each subsequent voice recognition result with the corresponding user, thereby improving the accuracy of multiple user voice interactions.
[0064] In some alternative embodiments, as shown in FIG. 3, step 202 may include the following substeps:
[0065] In step 2021, a speech recognition instance corresponding to each channel's audio signal is determined.
[0066] Here, the speech recognition instances may be constructed by code, each of which corresponds to one channel of audio signal, and each of which is used to recognize the corresponding one channel of audio signal.
[0067] In step 2022, each of the determined speech recognition instances is executed in parallel.
[0068] For example, a multi-threading method can be adopted to realize the parallel execution of each speech recognition instance, or each speech recognition instance can be executed on a different CPU to realize the parallel execution.
[0069] In step 2023, each speech recognition instance recognizes the corresponding audio signal by utilizing the preset speech recognition model.
[0070] Specifically, each speech recognition instance can call the preset speech recognition model in parallel or individually to recognize the corresponding audio signal, thereby realizing parallel recognition of audio signals. Typically, when recognizing at least one channel of audio signal, a preset speech recognition model is first loaded into memory, and each speech recognition instance can share the preset speech recognition model. In addition, when recognizing an audio signal using each speech recognition instance, the above cache can be shared, thereby improving the recognition efficiency of each speech recognition instance.
[0071] In this embodiment, by constructing a speech recognition instance corresponding to each audio signal and running each speech recognition instance in parallel, it is possible to realize simultaneous recognition of the speech of multiple users. Furthermore, each speech recognition instance jointly uses one speech recognition model to recognize the speech signal and jointly uses the same cache to store and retrieve recognition data, thereby realizing parallel speech recognition of at least one channel of audio signal and sharing the resources required for recognition. This improves the efficiency of speech recognition in multi-user speech dialogue scenarios. Since recognized data is stored in a shared cache, the stored recognition data can be directly retrieved in subsequent recognition processes without the need for repeated recognition, and further saves memory resources.
[0072] In some alternative embodiments, as shown in FIG. 4, step 206 may include the following substeps:
[0073] In step 2061, a semantic analysis instance (example) corresponding to each obtained phrase recognition result is determined.
[0074] Here, the analysis instances may be constructed by code, each analysis instance corresponding to one phrase recognition result of one channel of audio signal, and the analysis instances are used for structured analysis of the phrase recognition results.
[0075] In step 2062, each of the determined semantic analysis instances is executed in parallel.
[0076] For example, a multi-thread method can be adopted to realize the parallel execution of each semantic analysis instance, or each semantic analysis instance can be executed on a different CPU to realize the parallel execution.
[0077] In step 2063, each semantic analysis instance performs semantic analysis on the corresponding phrase recognition result.
[0078] Specifically, each semantic analysis instance can call modules such as a rule engine and a neural network engine that have been pre-installed in parallel to realize parallel analysis of the phrase recognition results.
[0079] This embodiment realizes simultaneous recognition and analysis of the speech of multiple users by constructing a semantic analysis instance corresponding to each phrase recognition result and running each semantic analysis instance in parallel, thereby constructing multiple links that allow simultaneous speech dialogue, and each semantic analysis instance jointly uses one semantic resource set, thereby improving the efficiency of speech recognition in multi-user speech dialogue scenarios.
[0080] Further, referring to FIG. 5, a schematic flowchart of another embodiment of a voice interaction method is shown. In this embodiment, as shown in FIG. 5, in addition to the embodiment shown in FIG. 2 above, step 202 may include the following steps:
[0081] In step 2024, a phonetic sub-model included in the speech recognition model is used to determine a set of syllables corresponding to each of the at least one channel of audio signals and a first probability score corresponding to each of the syllables in the set of syllables.
[0082] Here, the phonetic sub-model is used for syllable segmentation of the input audio signal. By way of example, the phonetic sub-model may include, but is not limited to, a Hidden Markov Model (HMM), a Gaussian Mixture Model (GMM), etc. The first probability score is used to characterize the probability that a syllable will be segmented correctly.
[0083] In step 2025, a set of words corresponding to each of the at least one channel of audio signals is determined using a language sub-model included in the speech recognition model.
[0084] Here, the language sub-model is used to determine a word set according to the above syllable set, and by way of example, the language sub-model may include, but is not limited to, an n-gram language model, a neural network language model, etc.
[0085] In step 2026, for each word in the word set, it is determined whether a second probability score corresponding to the word exists in the cache.
[0086] If not present, the language sub-model is utilized to determine a second probability score corresponding to the word, where the second probability score is used to characterize the probability of the recognized word occurring. For example, computing the probability of "air conditioner" occurring after "tsukeru" (put on) results in a second probability score corresponding to the word "air conditioner."
[0087] When the probability score of a word needs to be determined, the electronic device first searches whether the second probability score of the current word exists in the cache, and if not, uses the language sub-model to calculate the second probability score of the word. In this embodiment, because the language sub-model requires a large amount of data processing, in order to save processing costs, the second probability score generated by the language sub-model can be pre-stored using a cache, and the second probability score can be directly obtained from the cache when used.
[0088] In step 2027, the first probability score and a second probability score computed by the language submodel. Based on this, the recognition result of the first class is determined.
[0089] For example, each first probability score and a second probability score computed by the language submodel. can be determined as the first class recognition result.
[0090] Because the data processing volume of the language sub-model is large, the method provided in the embodiment corresponding to Figure 5 above uses a cache dedicated to storing the second probability scores generated by the language sub-model, so that the cache plays a more focused role, i.e., the cache is applied to processes with large data processing volumes and frequent data access, fully playing the role of using the cache to save computing resources, reducing redundant data in the cache and improving the efficiency of speech recognition.
[0091] Further, referring to FIG. 6, a schematic flowchart of another embodiment of the voice interaction method is shown. In this embodiment, as shown in FIG. 6, in addition to the embodiment shown in FIG. 5 above, step 203 may further include the following steps:
[0092] In step 2031, for each word in the word set, it is determined whether a second probability score corresponding to the word exists in the cache.
[0093] If so, the second probability score in the cache is determined as the second probability score for the word.
[0094] For example, if it is necessary to calculate the probability that "air conditioner" appears after "turn on", before the calculation, the cache is searched to see if the second probability score corresponding to "air conditioner" that was calculated has been previously stored. If it has been stored, it can be directly retrieved from the cache and used, thereby avoiding repeated calculation. If it has not been stored, it cannot be directly retrieved from the cache and must be calculated again.
[0095] In step 2032, First probability score and A recognition result for a second class is determined based on a second probability score determined from the cache.
[0096] For example, Each first probability score and A second probability score determined from the cache may be determined as the recognition result for the second class.
[0097] In the method provided in the embodiment corresponding to Figure 6 above, when determining the second probability score corresponding to a word, the second probability score is first retrieved from the cache, and the retrieved second probability score is determined as the second probability score of the word, thereby more narrowing down the calculation amount of the language sub-model and reducing the memory resources occupied by the recognition process of the language sub-model, thereby further improving the efficiency of speech recognition.
[0098] In some alternative embodiments, based on the example corresponding to FIG. 5 or FIG. 6 above, the above step 205 may be performed as follows:
[0099] First, a target path of a word set is determined in a decoding network included in a speech recognition model according to the first probability score and the second probability score included in the recognition result of the first class and the recognition result of the second class, respectively.
[0100] Here, the decoding network is a network constructed based on the above word set, and based on this network, an optimal path of the word combination is searched for within the network according to the first probability score and the second probability score, and this path is the target path.
[0101] The method of determining the optimal path according to the probability scores corresponding to syllables and the probability scores corresponding to words is a conventional technique, and a detailed description thereof will be omitted here.
[0102] Then, based on the target path, a phrase recognition result corresponding to each of the at least one channel of the audio signal is generated.
[0103] Specifically, a sentence formed by combining words corresponding to the target path can be determined as the phrase recognition result.
[0104] In this embodiment, the first probability score, the second probability score obtained using language sub-model computing, and the second probability score extracted from the cache are used to search for a target path in the decoding network to generate a phrase recognition result, thereby making full use of the second probability score stored in the cache during decoding, thereby improving the efficiency of generating phrase recognition results.
[0105] 7 shows a schematic diagram of an application scene of the voice interaction method of this embodiment. In the application scene of FIG. 7, the voice interaction method is applied to an in-vehicle voice interaction system.
[0106] 7, the multi-channel audio signals correspond to one dialogue chain included in a driver's seat dialogue chain 701, a passenger's seat dialogue chain 702, and another dialogue chain 703. Here, the driver's seat dialogue chain 701 is used for a driver to dialogue with the in-vehicle voice dialogue system, the passenger's seat dialogue chain 702 is used for a passenger in the passenger seat to dialogue with the in-vehicle voice dialogue system, and the other dialogue chain 703 is used for a passenger in another seat to dialogue with the in-vehicle voice dialogue system.
[0107] In addition, the decoding resource 704 includes a speech recognition model 7041 and a cache 7042, and the semantic resource 705 includes a rule engine 7051 for analyzing phrase recognition results and a neural network engine 7052. As can be seen from Figure 7, in the driver's seat dialogue chain 701, the electronic device generates a speech recognition instance A for the voice signal from the driver's seat and a speech recognition instance B for the voice from the passenger seat, and each speech recognition instance shares one set of decoding resource 704 and runs in parallel to obtain phrase recognition result C and phrase recognition result D.
[0108] Then, the electronic device constructs a word sense instance E and a word sense instance F, and the word sense instance E and the word sense instance F jointly use one word sense resource set to analyze the phrase recognition result C and the phrase recognition result D, respectively, to obtain a structured analysis result G and an analysis result H.
[0109] Then, the electronic device generates commands I, J, etc. based on the analysis results G and H, for example, command I is used to turn on the air conditioner, and command J is used to close the car windows. The in-vehicle voice dialogue device executes corresponding functions K and H based on the commands I and J. Similarly, the execution process of the other dialogue chains 703 is similar to the above-mentioned driver's seat dialogue chain 701 and passenger seat dialogue chain 702, so detailed description thereof will be omitted here.
[0110] Exemplary Apparatus 8 is a schematic structural diagram of a voice interaction device provided in an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device. As shown in FIG. 8, the voice interaction device includes: an acquisition module 801, a recognition module 802, a determination module 803, a first generation module 804, a processing module 805, an analysis module 806, and a second generation module 807.
[0111] Here, the acquisition module 801 is used to acquire at least one channel of audio signal, the recognition module 802 is used to recognize the at least one channel of audio signal using a preset speech recognition model and obtain a first class of recognition result using the speech recognition model, the determination module 803 is used to determine stored recognition data for the at least one channel of audio signal from the cache, the first generation module 804 is used to generate a second class of recognition result based on the stored recognition data, the processing module 805 is used to process the first class of recognition result and the second class of recognition result using the speech recognition model to obtain phrase recognition results respectively corresponding to the at least one channel of audio signal, the analysis module 806 is used to perform semantic analysis on each phrase recognition result respectively to obtain at least one analysis result, and the second generation module 807 is used to generate an instruction to control the voice interaction equipment to perform a corresponding function based on the at least one analysis result.
[0112] In this embodiment, the acquisition module 801 can acquire at least one channel of audio signals locally or remotely. For example, when this embodiment is applied to an in-vehicle voice recognition scenario, the at least one channel of audio signals can be at least one voice signal of a passenger in the vehicle collected by at least one microphone mounted on the vehicle.
[0113] In this embodiment, the recognition module 802 recognizes at least one channel of audio signals using a preset speech recognition model to obtain a first-class recognition result. Here, the speech recognition model may be a model obtained by pre-training using a large number of speech signal samples. The speech recognition model is used to recognize the input audio signals and obtain a phrase recognition result.
[0114] Typically, a speech recognition model may include multiple sub-models, such as a phonetics sub-model (used to perform syllable segmentation on an audio signal), a linguistic sub-model (used to convert each syllable into a word), and a decoding network (used to select the best combination of words to obtain a sentence).
[0115] During the recognition process of the above speech recognition model, the recognition module 802 usually first searches for recognition data corresponding to the current processing stage from the cache. If the corresponding recognition data does not exist in the cache, it uses the above speech recognition model for recognition, and the obtained recognition data is the first-class recognition result.
[0116] In this embodiment, the determination module 803 can determine stored recognition data for at least one channel of audio signals from the cache. Generally, in the process of recognition using the above-mentioned speech recognition model, the determination module 803 will first search the cache for recognition data corresponding to the current processing stage, and if the corresponding recognition data exists in the cache, extract the recognition data.
[0117] In this embodiment, the first generation module 804 can generate a second class of recognition result based on the extracted stored recognition data. For example, the stored recognition data can be the second class of recognition result, or the second class of recognition result can be obtained after processing the stored recognition data (for example, scaling the data proportionally, normalizing, etc.).
[0118] The first class recognition results and the second class recognition results are usually intermediate results obtained during speech recognition model processing, such as syllable probability scores and word probability scores.
[0119] In this embodiment, the processing module 805 uses a speech recognition model to process the first-class recognition results and the second-class recognition results to obtain phrase recognition results corresponding to at least one channel of audio signals. Typically, the first-class recognition results and the second-class recognition results are intermediate results obtained by processing the speech recognition model, so the speech recognition model needs to further process the first-class recognition results and the second-class recognition results. For example, the first-class recognition results and the second-class recognition results may include a probability score for each syllable and a probability score for each word obtained after recognizing the audio signal. For one audio signal, the speech recognition model uses a path search algorithm (e.g., the Viterbi algorithm) to determine an optimal path from the recognized words corresponding to the audio signal, and the resulting sentence is the phrase recognition result.
[0120] In this embodiment, the analysis module 806 performs semantic analysis on each phrase recognition result to obtain at least one analysis result. Here, among the at least one analysis result, each analysis result corresponds to one audio signal. In general, the analysis result may be structured data. For example, the phrase recognition result is "set the air conditioner temperature to 25 degrees," and the analysis result is "domain=car control, intent=air conditioner temperature setting, slot position=<temperature value=25>."
[0121] Note that a conventional technique may be used to analyze words and phrases, such as a rule engine or a neural network engine.
[0122] In this embodiment, the second generation module 807 can generate an instruction to control a voice interaction device to execute a corresponding function based on at least one analysis result. Here, the voice interaction device may be an electronic device in which the voice interaction device is installed, or an electronic device communicatively connected to the electronic device. For example, if the voice interaction device is an in-vehicle air conditioner, and the analysis result is "Domain = Car control, Intention = Air conditioner temperature setting, Slot position = <Temperature value = 25>", a command to control the in-vehicle air conditioner to set it to 25°C can be generated.
[0123] Please refer to FIG. 9, which is a schematic structural diagram of a voice interaction device provided in another exemplary embodiment of the present disclosure.
[0124] In some alternative embodiments, the device further includes a storage module 808 for caching recognition data obtained during the recognition process by the speech recognition model.
[0125] In some optional embodiments, the acquisition module 801 includes a receiving unit 8011 for receiving an initial audio signal collected by an audio collection device, and a processing unit 8012 for performing sound source separation processing on the initial audio signal to obtain at least one channel of audio signal.
[0126] In some alternative embodiments, the recognition module 802 includes a first determination unit 8021 for determining speech recognition instances corresponding to at least one channel of audio signals, a first execution unit 8022 for executing each of the determined speech recognition instances in parallel, and a recognition unit 8023 for each speech recognition instance to recognize the corresponding audio signal using a speech recognition model.
[0127] In some alternative embodiments, the analysis module 806 includes a second determination unit 8061 for determining a semantic analysis instance corresponding to each obtained phrase recognition result, a second execution unit 8062 for executing each determined semantic analysis instance in parallel, and an analysis unit 8063 for performing semantic analysis on each corresponding phrase recognition result using each semantic analysis instance.
[0128] In some alternative embodiments, the recognition module 802 includes a third determination unit 8024 for determining a set of syllables corresponding to each of the at least one channel audio signals and a first probability score corresponding to each of the syllables in the set of syllables using a phonetic sub-model included in the speech recognition model; a fourth determination unit 8025 for determining a set of words corresponding to each of the at least one channel audio signals using a linguistic sub-model included in the speech recognition model; and a fourth determination unit 8026 for determining, for each word in the set of words, whether a second probability score corresponding to the word exists in the cache, and if not, using the linguistic sub-model to find the second probability score corresponding to the word. Computing a fifth decision unit 8026 for determining the first probability score; and a second probability score computed by the language submodel. and a sixth determining unit 8027 for determining the first-class recognition result based on
[0129] In some alternative embodiments, the determination module 803 includes a seventh determination unit 8031 for determining, for a word in the word set, whether a second probability score corresponding to the word exists in the cache, and if so, determining the second probability score in the cache as the second probability score of the word; First probability score and and an eighth determining unit 8032 for determining a recognition result of the second class based on the second probability score determined from the cache.
[0130] In some alternative embodiments, the processing module 805 includes a ninth determination unit 8051 for determining a target path of a word set within a decoding network included in the speech recognition model according to the first probability score and the second probability score included in the first class recognition result and the second class recognition result, respectively, and a generation unit 8052 for generating phrase recognition results corresponding to at least one channel of audio signals based on the target path.
[0131] The voice interaction device provided in the above embodiments of the present disclosure recognizes at least one channel of audio signal using a preset voice recognition model, and during recognition, extracts stored recognition data from a cache to generate part of the recognition result, and another part of the recognition result is generated by the voice recognition model, thereby effectively reusing the stored recognition data without needing to process the entire amount of data in the voice recognition model, improving the processing efficiency of at least one channel of audio signal, and helping to still meet the requirements of low resource consumption and low processing delay even in multi-channel voice interaction scenarios.
[0132] Exemplary Electronic Devices An electronic device according to an embodiment of the present disclosure will be described below with reference to Fig. 10. The electronic device may be either or both of the terminal device 101 and the server 103 shown in Fig. 1, or may be a standalone device independent of them, and the standalone device can communicate with the terminal device 101 and the server 103 and receive input signals collected therefrom.
[0133] FIG. 10 illustrates a block diagram of an electronic device according to an embodiment of the present disclosure.
[0134] As shown in FIG. 10, the electronic device 1000 includes at least one processor 1001 and at least one memory 1002 .
[0135] Here, any one of the at least one processor 1001 may be a central processing unit (CPU) or other type of processing device having data processing capabilities and / or instruction execution capabilities, and can control other components within the electronic device 1000 to perform desired functions.
[0136] The memory 1002 may include one or more computer program products stored in various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored in the computer-readable storage medium, and the processor 1001 may execute the program instructions to implement the voice interaction method and / or other desired functions in each embodiment of the present disclosure. Various contents, such as recognition data, may also be stored in the computer-readable storage medium.
[0137] In one example, the electronic device 1000 may further include input devices 1003 and output devices 1004 connected to each other via a bus system and / or other form of connection (not shown).
[0138] For example, if the electronic device is the terminal device 101 or the server 103, the input device 1003 may be a device such as a microphone for inputting an audio signal. If the electronic device is a standalone device, the input device 1003 may be a communication network connector for receiving an audio signal input from the terminal device 101 or the server 103.
[0139] The output device 1004 can output various information, such as instructions for the voice interaction device to execute corresponding functions, to the outside. The output device 1004 may further include a display, a speaker, a printer, a communication network, and a remote output device connected thereto.
[0140] 10 illustrates only some of the components related to the present disclosure in the electronic device 1000, and omits components such as buses and input / output interfaces, etc. In addition, the electronic device 1000 may further include any other appropriate components according to specific application situations.
[0141] Exemplary Computer Program Products and Computer-Readable Storage Media In addition to the methods and apparatus described above, embodiments of the present disclosure may also be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform steps in the voice interaction methods according to various embodiments of the present disclosure described in the "Exemplary Methods" section above of this specification.
[0142] The computer program product may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and traditional procedural programming languages such as "C" or similar programming languages, and the program code may run entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or a server.
[0143] Additionally, an embodiment of the present disclosure may be a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, cause the processor to perform steps in the voice interaction method according to various embodiments of the present disclosure as described herein above in the "Exemplary Method" section.
[0144] The computer-readable storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a readable signal medium or a readable storage medium. The computer-readable storage medium may include, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a compact disk (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0145] Although the basic principles of the present disclosure have been described above in conjunction with specific embodiments, it should be understood that the benefits, advantages, and effects mentioned in the present disclosure are merely illustrative, not limiting, and that it should not be considered that each embodiment must have these benefits, advantages, and effects. Furthermore, the specific details of the above disclosure are merely illustrative and serve to facilitate understanding, rather than limiting, and the above details do not necessarily limit the present disclosure to be realized with the above specific details.
[0146] The embodiments in this specification are described in a sequential manner, with the differences between each embodiment being mainly described, and the same or similar parts between the embodiments may be referred to. The system embodiments are basically corresponding to the method embodiments, and therefore have been described relatively briefly. However, for related points, please refer to the description of the method embodiments.
[0147] Block diagrams of devices, apparatus, instruments, and systems according to the present disclosure are merely illustrative examples and are not intended to require or imply that they be necessarily connected, arranged, or configured in the manner shown in the block diagrams. One skilled in the art will recognize that these devices, apparatus, instruments, and systems may be connected, arranged, or configured in any manner. Terms such as "including," "comprising," and "having" are open vocabulary terms and may be used interchangeably to mean "including, but not limited to." As used herein, the terms "or" and "and" refer to the term "and / or" and may be used interchangeably therewith, unless the context clearly indicates otherwise. As used herein, the word "for example" means and may be used interchangeably with the phrase "for example, but not limited to."
[0148] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order used for the steps of the methods is for illustrative purposes only, and the steps of the methods of the present disclosure are not limited to the order specifically described above unless otherwise specified. Also, in some embodiments, the present disclosure may be embodied as a program recorded on a recording medium, which program includes machine-readable instructions for implementing the methods of the present disclosure. Thus, the present disclosure also includes a recording medium having a program stored thereon for executing the methods of the present disclosure.
[0149] It should be noted that in the devices, apparatuses, and methods of the present disclosure, each component or step may be disassembled and / or recombined, and such disassembly and / or recombination should be considered as an equivalent solution of the present disclosure.
[0150] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0151] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. While several exemplary aspects and embodiments have been described above, those skilled in the art will recognize certain variations, modifications, variations, additions, and subsets thereof.
Claims
1. obtaining at least one channel of audio signals, the at least one channel of audio signals being obtained by performing sound source separation on an initial audio signal, the at least one channel of audio signals each corresponding to one user; Recognizing the at least one channel audio signal using one of the preset speech recognition models, and determining a set of syllables corresponding to each of the at least one channel audio signal, a first probability score corresponding to each of the syllables in the set of syllables, and a set of words corresponding to each of the at least one channel audio signal; determining whether a stored second probability score corresponding to a word of the set of words of the at least one channel audio signal exists in a cache, the cache being the same cache commonly used in recognizing the at least one channel audio signal using the speech recognition model; if the second probability score corresponding to the word exists, generating a second class recognition result based on the first probability score and the second probability score determined from the cache; if the second probability score corresponding to the word does not exist in the cache, computing a second probability score corresponding to the word using the speech recognition model, determining a first class recognition result based on the first probability score and the second probability score computed by the speech recognition model; and storing the second probability score computed by the speech recognition model in the cache; utilizing the speech recognition model to process the first class recognition results and the second class recognition results to obtain at least one phrase recognition result corresponding to the at least one channel audio signal; performing semantic analysis on the phrase recognition results to obtain at least one analysis result; and generating an instruction for controlling the voice interaction device to perform a corresponding function based on the at least one analysis result. Voice interaction method.
2. The step of acquiring at least one channel of audio signal comprises: receiving an initial audio signal collected by an audio collecting device; and performing a sound source separation process on the initial audio signal to obtain the at least one channel audio signal. The method of claim 1.
3. The step of recognizing the at least one channel audio signal using one preset speech recognition model, and determining a set of syllables corresponding to each of the at least one channel audio signal, a first probability score corresponding to each of the syllables in the set of syllables, and a set of words corresponding to each of the at least one channel audio signal includes: determining a speech recognition instance corresponding to each of the at least one channel of audio signals; executing each of the determined speech recognition instances in parallel; each speech recognition instance utilizing the speech recognition model to recognize a corresponding audio signal; The method of claim 1.
4. The step of performing semantic analysis on the phrase recognition result to obtain at least one analysis result includes: determining a semantic analysis instance corresponding to each of the obtained phrase recognition results; executing each of the determined semantic analysis instances in parallel; and performing semantic analysis on the corresponding phrase recognition results using each semantic analysis instance to obtain the at least one analysis result. The method of claim 3.
5. The step of recognizing the at least one channel audio signal using one preset speech recognition model, and determining a set of syllables corresponding to each of the at least one channel audio signal, a first probability score corresponding to each of the syllables in the set of syllables, and a set of words corresponding to each of the at least one channel audio signal includes: determining a set of syllables corresponding to each of the at least one channel of audio signals and a first probability score corresponding to each of the syllables in the set of syllables using a phonetic sub-model included in the speech recognition model; determining a set of words corresponding to each of the at least one channel audio signals by using a language sub-model included in the speech recognition model; The method of claim 1.
6. If the second probability score corresponding to the word does not exist in the cache, the language sub-model included in the speech recognition model is used to compute the second probability score corresponding to the word. The method of claim 5.
7. the step of processing the first class recognition result and the second class recognition result using the speech recognition model to obtain at least one phrase recognition result corresponding to the at least one channel audio signal includes: determining a target path for the set of words within a decoding network included in the speech recognition model according to a first probability score and a second probability score included in the recognition results of the first class and the second class, respectively; generating at least one phrase recognition result corresponding to the at least one channel of audio signal based on the target path; The method of claim 6.
8. A voice dialogue device, an acquisition module for acquiring at least one channel of audio signals, the at least one channel of audio signals being obtained by performing sound source separation on an initial audio signal, the at least one channel of audio signals each corresponding to one user; a recognition module that recognizes the at least one channel audio signal using one of the preset speech recognition models, and determines a set of syllables corresponding to each of the at least one channel audio signal, a first probability score corresponding to each of the syllables in the set of syllables, and a set of words corresponding to each of the at least one channel audio signal; a determination module for determining whether a stored second probability score corresponding to a word in the set of words of the at least one channel audio signal is present in a cache, the cache being the same cache used in recognizing the at least one channel audio signal using the speech recognition model; a first generation module for generating a recognition result of a second class based on the first probability score and the second probability score determined from the cache if the second probability score corresponding to the word exists; If the cache does not contain a second probability score corresponding to the word, the recognition module uses the speech recognition model to compute a second probability score corresponding to the word, determines a recognition result of a first class based on the first probability score and the second probability score computed by the speech recognition model, and stores the second probability score computed by the speech recognition model in the cache; The voice dialogue device further comprises: a processing module for processing the first class recognition results and the second class recognition results using the speech recognition model to obtain at least one phrase recognition result corresponding to the at least one channel audio signal; an analysis module for performing semantic analysis on the phrase recognition result to obtain at least one analysis result; a second generating module for generating an instruction for controlling the voice interaction device to perform a corresponding function based on the at least one analysis result; Voice dialogue device.
9. A computer-readable storage medium having computer program instructions stored thereon, comprising: The computer program instructions, when executed, implement the method of any one of claims 1 to 7. A computer-readable storage medium.
10. a processor; a memory for storing instructions executable by the processor; The processor is adapted to read the executable instructions from the memory and execute the instructions to implement the method of any one of claims 1 to 7. electronic equipment.
Citation Information
Patent Citations
Computer program for operating computer as voice recognition device and sentence classification device, computer program for operating computer so as to realize method of generating hierarchized language model, and storage medium
JP2004198597A
Speech recognition apparatus, speech recognition method, and speech recognition program
JP2007093789A