A method for processing voice services, an electronic device, and a computer-readable storage medium
By deploying ASR and NLU modules in the end-side device, using corpus matching technology to query semantic status information, the problems of performance overhead and delay in voice service processing are solved, and more efficient voice service processing is achieved.
Patent Information
- Application Number
- CN202010702674.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-07-21
AI Technical Summary
In the existing voice service processing technology, the frequent execution of speech recognition and semantic understanding leads to large system performance overhead and high latency, affecting user experience.
By deploying ASR and NLU modules in the end-side device, semantic status information is queried locally using corpus matching technology to avoid duplicate execution of speech recognition and semantic understanding.
It reduces system performance overhead, reduces power consumption, improves voice service processing efficiency, shortens response delay, and improves user experience.
Smart Images

Figure CN114038462B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice service processing, and particularly to a voice service processing method, an electronic device, and a computer-readable storage medium.
Background Art
[0002] Artificial Intelligence (AI) is a new technical science that studies, develops, and applies theories, methods, techniques, and application systems for simulating, extending, and expanding human intelligence. Among them, Automatic Speech Recognition (ASR) technology is one of the important technologies in the field of artificial intelligence. Current speech recognition systems generally include: an ASR module, a Natural-language understanding (NLU) module, a Dialog Management (DM) module, a Natural-language generation (NLG) module, and a Text To Speech (TTS) module. Among them, the ASR module is used to convert the input voice signal into text information. The NLU module is used to convert the input text information into semantic information that can be understood by a machine. The DM module provides corresponding services based on the state of the conversation and the semantic information. The NLG module is used to generate natural language text according to the service information. The TTS module is used to convert natural language text into speech.
[0003] In the voice service processing flow of related technologies, after the voice is input into the ASR module, the text result is recognized by the ASR module, and the text result is input into the NLU module. The NLU module obtains the corresponding intent and slot of the text result, so that the DM module, the NLG module, and the TTS module can access relevant services or execute relevant actions, and display the execution result. However, in related technologies, voice recognition and semantic understanding are required for each execution of a voice service, resulting in a large performance overhead and affecting the system power consumption. In addition, in related technologies, multiple rounds of end-cloud interaction are also required, resulting in a relatively large delay and affecting the user experience.
Summary of the Invention
[0004] In view of this, the present invention provides a voice service processing method, an electronic device, and a computer-readable storage medium, which improve the processing efficiency of voice services through corpus matching, and avoid performing voice recognition and semantic understanding processes during voice service processing, thereby reducing system performance overhead and system power consumption.
[0005] On the one hand, an embodiment of the present invention provides a method for processing voice services, including:
[0006] Obtain the input first corpus;
[0007] Query a second corpus that matches the first corpus from the obtained multiple service records, and obtain the semantic status information corresponding to the second corpus;
[0008] Provide corresponding services according to the semantic status information corresponding to the second corpus.
[0009] In an alternative implementation, the first corpus includes original voice information or text information.
[0010] In an alternative implementation, before querying a second corpus that matches the first corpus from the obtained multiple service records and obtaining the semantic status information corresponding to the second corpus, it further includes:
[0011] Obtain the corpus, semantic status information, and execution results corresponding to multiple voice services;
[0012] Use the corpus, semantic status information, and execution results corresponding to the voice service as a service record to generate multiple service records.
[0013] In an alternative implementation, after using the corpus, semantic status information, and execution results corresponding to the voice service as a service record to generate multiple service records, it further includes:
[0014] According to the positive and negative feedback learning algorithm, perform positive and negative feedback learning on multiple service records with the same semantic status information, and determine the processing results of multiple service records with the same semantic status information. The processing results include adopting multiple service records or deleting multiple service records.
[0015] In an alternative implementation, the step of performing positive and negative feedback learning on multiple service records with the same semantic status information according to the positive and negative feedback learning algorithm and determining the processing results of multiple service records with the same semantic status information, where the processing results include adopting multiple service records or deleting multiple service records, includes:
[0016] Count the number of successful or failed times of multiple service records with the same semantic status information;
[0017] If it is determined that the number of successful times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as adopting multiple service records;
[0018] If it is determined that the number of failures is greater than or equal to a preset number, the processing results of multiple business records with the same semantic status information are determined to delete the multiple business records.
[0019] In an alternative implementation, the semantic status information includes an intention and slots, or an intention, slots, and context information.
[0020] In an alternative implementation, querying for a second corpus that matches the first corpus among the multiple acquired business records includes:
[0021] Calculating the voice similarity between the second corpus and the first corpus among the multiple business records;
[0022] Querying for the second corpus in the multiple business records where the voice similarity is greater than or equal to a preset threshold.
[0023] In an alternative implementation, querying for a second corpus that matches the first corpus among the multiple acquired business records includes:
[0024] Querying for the second corpus in the multiple business records, where the second corpus contains the first corpus.
[0025] In a second aspect of an alternative implementation, an embodiment of the present invention provides an electronic device, the device includes:
[0026] Obtaining an input first corpus;
[0027] Querying for a second corpus that matches the first corpus among the multiple acquired business records, and obtaining the semantic status information corresponding to the second corpus;
[0028] Providing a corresponding service according to the semantic status information corresponding to the second corpus.
[0029] In an alternative implementation, the first corpus includes original voice information or text information.
[0030] In an alternative implementation, when the instruction is executed by the device, the device specifically performs the following steps:
[0031] Obtaining the corpus, semantic status information, and execution results corresponding to multiple voice services;
[0032] Taking the corpus, semantic status information, and execution results corresponding to the voice service as a business record to generate multiple business records.
[0033] In an alternative implementation, when the instruction is executed by the device, the device specifically performs the following steps:
[0034] According to the positive and negative feedback learning algorithm, perform positive and negative feedback learning on multiple said business records with the same semantic state information, and determine the processing results of multiple said business records with the same semantic state information, where the processing results include adopting multiple said business records or deleting multiple said business records.
[0035] In an alternative implementation manner, when the instruction is executed by the device, the device specifically executes the following steps:
[0036] Count the number of successful times or the number of failed times of multiple said business records with the same semantic state information;
[0037] If it is determined that the number of successful times is greater than or equal to a preset number, determine the processing result of multiple said business records with the same semantic state information as adopting multiple said business records;
[0038] If it is determined that the number of failed times is greater than or equal to a preset number, determine the processing result of multiple said business records with the same semantic state information as deleting multiple said business records.
[0039] In an alternative implementation manner, the semantic state information includes intent and slot, or intent, slot, and context information.
[0040] In an alternative implementation manner, when the instruction is executed by the device, the device specifically executes the following steps:
[0041] Calculate the voice similarity between the second corpus and the first corpus in the multiple business records;
[0042] Query the second corpus in the multiple business records where the voice similarity is greater than or equal to a preset threshold.
[0043] In an alternative implementation manner, when the instruction is executed by the device, the device specifically executes the following steps:
[0044] Query the second corpus in the multiple business records, where the second corpus contains the first corpus.
[0045] In a second aspect, an embodiment of the present invention provides an electronic device, including: one or more processors; a memory; multiple application programs; and one or more computer programs, where the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, when the instructions are executed by the device, the device executes the voice service processing method in any possible implementation of any aspect above.
[0046] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium for program code to be executed by a device, and the program code includes instructions for executing the method in the first aspect or any possible implementation manner of the first aspect.
[0047] The technical solution provided by the embodiment of the present invention is applied to the fields of artificial intelligence and speech recognition technology. By obtaining the input first corpus, querying the second corpus that matches the first corpus in the obtained multiple business records, and obtaining the semantic status information corresponding to the second corpus, and providing corresponding services according to the semantic status information corresponding to the second corpus, it avoids the problems of needing to perform speech recognition and semantic understanding in the process of speech service processing, thereby reducing the system performance overhead and system functions, and improving the processing efficiency of speech services.
BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.
[0049] Figure 1 It is a flowchart of a speech service processing method in the related art;
[0050] Figure 2 It is an architecture diagram of a speech service processing system provided by an embodiment of the present invention;
[0051] Figure 3 It is a flowchart of a speech service processing method provided by an embodiment of the present invention;
[0052] Figure 4 It is a schematic diagram of a speech service processing method provided by an embodiment of the present invention;
[0053] Figure 5 It is a working flowchart of the record processing module 120 provided by an embodiment of the present invention;
[0054] Figure 6 It is a schematic block diagram of an electronic device provided by an embodiment of the present invention;
[0055] Figure 7 It is a structural diagram of an electronic device provided by an embodiment of the present invention.
DETAILED DESCRIPTION
[0056] In order to better understand the technical solutions of the present invention, the embodiments of the present invention will be described in detail below with reference to the drawings.
[0057] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0058] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0059] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects. For ease of understanding, examples are given for reference to the background art of the present invention and the description of related concepts of the embodiments of the present invention.
[0060] (1) Artificial Intelligence
[0061] Artificial Intelligence (AI for short) is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.
[0062] Artificial Intelligence is a branch of computer science that can understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc.
[0063] (2) Speech Recognition Technology
[0064] Speech recognition technology, also known as Automatic Speech Recognition (ASR for short), is a technology that converts human speech into text. Speech recognition technology is an important branch in the field of artificial intelligence technology and a multi-disciplinary field that is closely connected with many disciplines such as acoustics, phonetics, linguistics, digital signal processing theory, information theory, computer science, etc. The goal of ASR technology is to enable a computer to "dictate" the continuous speech spoken by different people, which is a technology for realizing the conversion from "sound" to "text".
[0065] (3) Voice Services
[0066] Voice services include transactions that obtain service responses by inputting voice or text. For example, when the voice input is "What's the weather like today", the service returns "Weather query results".
[0067] After the above description of related concepts, the following briefly introduces the voice service processing methods in related technologies.
[0068] Figure 1 For the flowchart of the voice service processing method in related technologies, as Figure 1 shown, the voice recognition system in related technologies usually includes a speech recognition module (ASR), a semantic understanding module (NLU), a dialogue management module (DM), a text-to-speech module (TTS), and a display result / action execution module. Specifically, in the process of voice service processing in related technologies, after the voice service processing system receives the input voice, it uses the speech recognition module to recognize the text information corresponding to the voice, inputs the text information into the semantic understanding module, obtains the intent and slots output by the semantic understanding module, and inputs the intent and slots into the dialogue management module, so that the dialogue management module interacts with the service management according to the intent and slots to obtain the service corresponding to the intent and slots. For example, the services corresponding to the intent and slots may include services such as opening navigation, opening video, opening music, and weather query. Further, the service execution result of the dialogue management module is displayed through the display result / action execution module, or the service corresponding to the intent and slots is executed through the action execution module. Optionally, a natural language text corresponding to the information of the service can also be generated through the text-to-speech module, so that the display result module displays the natural language text corresponding to the information of the service.
[0069] However, in related technologies, each time a voice service is executed, the ASR module and the NLU module need to execute corresponding steps. That is to say, each time a voice service is executed, the voice needs to be recognized to generate corresponding text information, and the text information is input into the semantic understanding module to obtain the intent and slots output by the semantic understanding module, resulting in a large performance overhead and affecting the system power consumption. In addition, in the voice service processing system of related technologies, the ASR module and the NLU module are usually deployed on cloud devices, while other modules are deployed on terminal devices, resulting in multiple rounds of end-cloud interactions during the voice service processing, resulting in a large delay and affecting the user experience.
[0070] In view of the problems in the related art, an embodiment of the present invention provides a voice service processing method, which is applied to the fields of artificial intelligence and speech recognition technology. By obtaining the input first corpus, querying the second corpus that matches the first corpus from the obtained multiple service records, and obtaining the semantic status information corresponding to the second corpus, and providing corresponding services according to the semantic status information corresponding to the second corpus, it avoids the problems of speech recognition and semantic understanding required in the process of voice service processing, thereby reducing the system performance overhead and system functions, and improving the processing efficiency of voice services.
[0071] After the introduction of the above related technologies, the voice service processing method of the present invention will be introduced in detail below.
[0072] Figure 2 The architecture diagram of a voice service processing system provided by an embodiment of the present invention is shown in Figure 2 As shown, the system may include an end-side device 110. For example, the end-side device may be a mobile phone, a tablet computer, a television, a desktop computer, a wearable device (such as a watch), a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc., which are devices with display functions. The specific type of the end-side device is not limited in the embodiments of the present application.
[0073] In an embodiment of the present invention, specifically, the electronic device 110 includes: a receiving module 111, a matching module 112, a storage module 113, an ASR module 114, an NLU module 115, a DM module 116, an NLG module 117, an execution module 118, a display module 119, and a record processing module 120. It should be noted that generally, modules such as the ASR module and the NLU module in the related art are usually deployed on the cloud side, and the difference between the voice service processing system of the embodiment of the present invention and the voice service processing system of the related art is that the ASR module 114, the NLU module 115, etc. in the embodiment of the present invention are all deployed on the end-side device.
[0074] In an embodiment of the present invention, the receiving module 111 is used to obtain the input first corpus. The first corpus includes original voice information or text information, where the text information includes original text information or text information converted from original voice information. For example, the receiving module 111 may include input devices such as a microphone or a keyboard. When the receiving module 111 includes a microphone, the obtained first corpus may include original voice information, and when the receiving module 111 includes a keyboard, the obtained first corpus may include original text information.
[0075] The matching module 112 is used to query a second corpus that matches the first corpus from the obtained multiple business records, and obtain semantic status information corresponding to the second corpus. In an embodiment of the present invention, the semantic status information includes an intention and slots, or an intention, slots, and context information. In an embodiment of the present invention, if no second corpus that matches the first corpus is queried from the obtained multiple business records, the first corpus is input into the NLU model to obtain the semantic status information corresponding to the first corpus output by the NLU model.
[0076] In an optional implementation manner, the matching module 112 is specifically configured to calculate the voice similarity between the second corpus and the first corpus in the multiple business records; and query the second corpus in the multiple business records whose voice similarity is greater than or equal to a preset threshold.
[0077] In another possible implementation manner, the matching module 112 is specifically configured to query the second corpus in the multiple business records, and the second corpus contains the first corpus.
[0078] The storage module 113 is used to store multiple business records. Among them, the storage module 113 may include a local database. In an embodiment of the present invention, by storing multiple business records in the local database, when the user executes the same business again, the corresponding semantic status information can be quickly and accurately found through the second corpus that matches the first corpus, and the corresponding business can be directly executed, so that the execution time of repeated services can be significantly shortened while reducing the process of local-cloud interaction and reducing the response delay.
[0079] The ASR module 114 is used to identify the first corpus and generate text information. For example, when the receiving module 111 includes a microphone, the obtained first corpus may include original voice information, and the original voice information can be converted into text information by the ASR module 114 for subsequent voice service processing.
[0080] The ULN module 115 is used to convert the text information input into the ASR module 114 into semantic status information that can be understood by a machine. Among them, the semantic status information includes an intention and slots, or an intention, slots, and context information.
[0081] The DM module 116 is used to provide corresponding services according to the semantic status information corresponding to the second corpus.
[0082] The NLG module 117 is used to generate natural language texts according to the service information.
[0083] The execution module 118 is used to execute the service, and the display module 119 is used to display the execution result of the corresponding service provided by the DM module 116. In a possible implementation manner, the display module 119 is further used to display the natural language text generated by the NLG module 117 according to the service information.
[0084] It should be noted that before storing multiple service records, the record processing module 120 is used to obtain the corpus, semantic status information, and execution results corresponding to multiple voice services, and use the corpus, semantic status information, and execution results corresponding to the voice service as a service record to generate multiple service records.
[0085] The record processing module 120 is further used to perform positive and negative feedback learning on multiple service records with the same semantic status information according to the positive and negative feedback learning algorithm, and determine the processing results of multiple service records with the same semantic status information. The processing results include adopting multiple service records or deleting multiple service records. In an optional implementation manner, the record processing module 120 is specifically used to count the number of successful times or failure times of multiple service records with the same semantic status information; if it is determined that the number of successful times is greater than or equal to the preset number of times, the processing result of multiple service records with the same semantic status information is determined to adopt multiple service records; if it is determined that the number of failure times is greater than or equal to the preset number of times, the processing result of multiple service records with the same semantic status information is determined to delete multiple service records.
[0086] The following introduces the actual application of the above system through two solutions:
[0087] In an optional solution in the embodiments of the present invention, as Figure 2 shown, when the user executes the same service, the first corpus obtained by the receiving module 111 is input into the matching module 112. The matching module 112 queries the second corpus that matches the first corpus from the obtained multiple service records, obtains the semantic status information corresponding to the second corpus, and inputs the semantic status information into the DM module 116, so that the DM module 116 provides corresponding services according to the semantic status information corresponding to the second corpus.
[0088] In another optional solution, as Figure 2As shown, when the user performs different services, the first corpus obtained by the receiving module 111 is input into the ASR module 114, so that the ASR module 114 recognizes the first corpus to generate text information, and the text information is input into the NLU module 115, so that the NLU module 115 generates corresponding semantic state information according to the first corpus, and the semantic state information is input into the DM module 116, so that the DM module 116 provides corresponding services according to the semantic state information corresponding to the second corpus.
[0089] In the embodiment of the present invention, through the above system 110, when the user is performing the same service, the second corpus matching the first corpus is queried from multiple obtained service records, and the semantic state information corresponding to the second corpus is obtained, and corresponding services are provided according to the semantic state information corresponding to the second corpus. When the user is performing different services, the first corpus is input into the NLU model, and the semantic state information corresponding to the first corpus output by the NLU model is obtained; corresponding services are provided according to the semantic state information corresponding to the first corpus, which solves the problem in the related art that the steps of speech recognition and semantic understanding need to be performed in the process of speech service processing, resulting in large performance overhead and affecting the system power consumption, and at the same time improves the processing efficiency of speech services. As an alternative solution, when the user is performing the same type of service, the function of corpus matching may also be realized through the above system 110, thereby improving the processing efficiency of speech services.
[0090] In addition, in the present invention, by generally deploying modules such as the ASR module and the NLU module on the end-side device, the problem that speech services require multiple rounds of end-cloud interaction, resulting in relatively large latency and affecting the user experience, is solved. The following combines Figure 3 and Figure 4 , including steps 102 to 112, to describe the process of the speech service processing method in detail.
[0091] Figure 3 is a flowchart of a speech service processing method provided by an embodiment of the present invention, Figure 4 is a schematic diagram of a speech service processing method provided by an embodiment of the present invention. As Figure 3 and Figure 4 shown, the method includes:
[0092] Step 102, obtain the input first corpus.
[0093] In the embodiment of the present invention, each step is executed by the end-side device.
[0094] In the embodiment of the present invention, the first corpus includes original speech information or text information, where the text information includes original text information or text information converted from original speech information.
[0095] For example, as Figure 2 shown, the first corpus may include the original voice information obtained by the receiving module 111, or the original text information obtained by the receiving module 111, or the text information generated after the input original voice information is recognized by the ASR module 114. The input manner of the first corpus in the present invention is not limited, and the first corpus with different input manners can be obtained according to requirements.
[0096] Step 104: Determine whether a second corpus matching the first corpus is retrieved from the obtained multiple service records. If so, execute Step 106; if not, execute Step 110.
[0097] In the embodiment of the present invention, before executing Step 104, it further includes:
[0098] Step 103a: Obtain the corpus, semantic status information, and execution result corresponding to multiple voice services.
[0099] In the embodiment of the present invention, a voice service includes a transaction that obtains a service response by inputting voice or text. The semantic status information includes an intention and slots, or an intention, slots, and context information, where the context information includes the associated information generated by multiple rounds of interaction of the voice service. The execution result is used to indicate the execution result of the voice service. Specifically, the process of obtaining the execution result can be determined by the execution module 118 in Figure 2 and the service processing information generated during the execution of the voice service, where the service processing information includes the execution result, and the execution result includes success or failure. It should be noted that in addition to the execution result, the service processing information may further include other information, which is not limited in the present invention.
[0100] In the voice service processing flow, the record processing module 120 obtains the corpus, semantic status information, and execution result of each voice service. Specifically, Figure 5 is the working flow chart of the record processing module 120 provided in the embodiment of the present invention. As Figure 5 shown, the obtaining process of the record processing module 120 specifically includes: obtaining the corpus corresponding to each voice service through the ASR module, obtaining the intention and slots corresponding to each corpus through the NLU module, obtaining the context information corresponding to each voice service through the DM module, and obtaining the execution result corresponding to each voice service through the execution module. In practical applications, for example, the record processing module 120 obtains the corpus, semantic status information, and execution result corresponding to 5 voice services through the above obtaining process, as shown in Table 1 below:
[0101] Table 1
[0102] Corpus Intention Slot Context Execution Result Call Zhang San Make a call Zhang San Associated Information of Multi-round Interaction Success Set an alarm for 8 o'clock Create an alarm Time - 8 o'clock Associated Information of Multi-round Interaction Success Play Great Compassion Mantra Play music Great Compassion Mantra —— Success Open the note Open an app Sys.app - Note —— Success Open the calendar Open an app Sys.app - Calendar —— Success
[0103] Step 103b: Use the corpus, semantic status information, and execution result corresponding to the voice service as a service record to generate multiple service records.
[0104] In the embodiment of the present invention, as shown in Table 1 above, using the corpus, semantic status information, and execution result corresponding to the voice service in Table 1 as a service record can generate 5 service records. In addition, a service record generated in the present invention can be stored in the end-side database in the form of a list.
[0105] Among them, the method for determining whether the execution result is successful may include: after the execution module executes the corresponding service according to the voice status information, determine whether the execution result is successful according to information such as human-computer interaction and result usage. For example, taking the service of making a call as an example, by determining whether the call duration is less than a preset threshold, if it is determined that the call duration is less than the preset threshold, it indicates that the execution result of this service is a failure. For example, taking the service of playing music as an example, by determining whether the music playing duration is less than a preset threshold, if it is determined that the music playing duration is less than the preset threshold, it indicates that the execution result of this service is a failure. For example, taking the service of creating an alarm as an example, by determining whether the set alarm is used, if it is determined that the set alarm is used, it indicates that the execution result of this service is successful. That is to say, if the service is interrupted during execution, it indicates that the execution result of this service is a failure, and vice versa, it indicates that the execution result of this service is successful.
[0106] In the embodiment of the present invention, the storage module 113 is used to store multiple service records. Among them, the storage module 113 may include an end-side database. By storing multiple service records with successful execution results in the end-side database, it is convenient for the user to execute the same voice service again, and it can perform subsequent step 106 to quickly and accurately find the corresponding semantic status information through the corpus, and directly execute the corresponding service according to the semantic status information, thereby greatly shortening the execution time of repeated services and improving the efficiency of voice service processing. In addition, by storing multiple service records with successful execution results in the end-side database, the services that have occurred can execute the voice service processing flow on the end-side device when there is no network connection.
[0107] Step 103c: According to the positive and negative feedback learning algorithm, perform positive and negative feedback learning on multiple service records with the same semantic status information, and determine the processing results of multiple service records with the same semantic status information. The processing results include adopting multiple service records or deleting multiple service records.
[0108] In an embodiment of the present invention, positive and negative feedback learning is performed on the execution results corresponding to multiple business records with the same semantic status information. The number of successful executions N and the number of failed executions M of the execution results corresponding to multiple business records are counted. Modeling calculations can be performed on N and M, and whether the business record can be applied to subsequent services is judged based on the processing result output by the model.
[0109] For example, in an alternative solution, the execution process of step 103c may specifically include: counting the number of successful executions or the number of failed executions of multiple business records with the same semantic status information; if it is determined that the number of successful executions is greater than or equal to a preset number, the processing result of multiple business records with the same semantic status information is determined to adopt multiple business records; if it is determined that the number of failed executions is greater than or equal to a preset number, the processing result of multiple business records with the same semantic status information is determined to delete multiple business records.
[0110] In an embodiment of the present invention, the preset number can be set as needed. For example, the preset number includes 3 times. In practical applications, for example, multiple business records with the same semantic status information include multiple business records of "calling Zhang San". Through the number of these business records, if it is determined that the number of successful executions of multiple business records is greater than 3 times, then this business record is adopted. It can be understood that the service corresponding to this business record is the service required by the user. If it is determined that the number of failed executions of multiple business records is greater than 3 times, then this business record is deleted. It can be understood that the service corresponding to this business record is not the service required by the user. Taking making a call as an example, if the business record of "calling Zhang San" fails 3 times, it indicates that the service corresponding to this business record is not the service required by the user, so this business record needs to be deleted.
[0111] By executing the above step 103, the corpus, semantic status information, and execution result corresponding to the voice service with a successful execution result can be used as a business record to generate multiple business records, so that the generated business records are all adoptable business records.
[0112] It should be noted that the method of judging the number of successful executions or the number of failed executions of multiple business records adopted in the above alternative solution is only a way of positive and negative feedback learning. In addition, other ways may also be included. For example, by judging the user perception experience of multiple business records with the same semantic status information, the processing result of multiple business records with the same semantic status information is determined. The present invention does not limit this and only gives an example for illustration.
[0113] After determining multiple applicable business records through the above steps 103a to 103c, during the execution of step 104, if it is determined that a second corpus matching the first corpus is retrieved from the multiple retrieved business records, it indicates that the voice service processing system has processed a business identical to the first corpus, and subsequent step 106 can be executed to obtain the semantic status information corresponding to the second corpus; if it is determined that no second corpus matching the first corpus is retrieved from the multiple retrieved business records, it indicates that the voice service processing system has not processed a business identical to the first corpus, and subsequent step 110 needs to be executed to input the first corpus into the NLU model to obtain the semantic status information corresponding to the first corpus output by the NLU model. That is to say, through the above steps 103a to 103c, the process of determining multiple applicable business records in a self-learning manner can be achieved, so as to determine a second corpus matching the first corpus in subsequent steps, further improving the processing efficiency of voice services.
[0114] Regarding the above step 104, it should be noted that when the first corpus is the original voice information, the retrieved second corpus includes voice information; when the first corpus is text information, the retrieved second corpus includes text information.
[0115] Step 106, obtain the semantic status information corresponding to the second corpus.
[0116] In the embodiments of the present invention, after retrieving a second corpus matching the first corpus, the semantic status information corresponding to the second corpus can be obtained from multiple business records. Specifically, as shown in Table 1 above, there is a corresponding relationship between the corpus and the semantic status information. The semantic status information includes an intention and slots, or an intention, slots, and context information. Among them, the context information may include associated information of multi-round interactions, and the multi-round interactions may include linear multi-round interactions and non-linear multi-round interactions. Specifically, the linear multi-round interaction is a way to actively initiate a follow-up question by the system when a required slot is missing to obtain the missing slot; the non-linear multi-round interaction is a question that requires referring to the previous context to obtain the complete intention of the user. For example, when executing the corpus "Call Zhang San" or "Set an alarm at 8 o'clock tomorrow" in Table 1 above, the context information needs to be obtained through the DM module 116. When executing the corpus "Open the note" or "Open the calendar" in Table 1 above, when the complete intention and slots can be obtained, there is no need to combine the context information, that is, the corresponding service can be executed. Taking the example where the user inputs the corpus "Set an alarm at 8 o'clock tomorrow" at 1 am, the system cannot determine whether the user's "alarm at 8 o'clock tomorrow" refers to the alarm 7 hours later or the alarm at 8 o'clock the next day. Therefore, the system actively initiates a follow-up question to obtain the missing slot, that is, the complete slot is determined through multi-round interactions.
[0117] In an embodiment of the present invention, when the second corpus is voice information, as an alternative solution, the specific execution process of step 106 may include:
[0118] Step 1061: Calculate the voice similarity between the second corpus and the first corpus in the multiple service records.
[0119] In an embodiment of the present invention, in an alternative manner, the text information and voiceprint information of the second corpus and the first corpus can be obtained respectively. By matching the similarity of the text information of the first corpus and the second corpus, and based on the principle of voice waveform matching, the similarity of the voiceprint information of the first corpus and the second corpus is matched, so as to calculate the voice similarity between the second corpus and the first corpus in the multiple service records.
[0120] Among them, the method for calculating the similarity of text information may include: determining the similarity of text information by judging the number of words with the same number in the text information.
[0121] The method for calculating the similarity of text information may include: inputting the voiceprint information of the first corpus into a preset voiceprint feature set to query and match the voiceprint features, inputting the voiceprint information of the second corpus into the preset voiceprint feature set to query and match the voiceprint features, and determining the similarity of the voiceprint information of the first corpus and the second corpus by judging the voiceprint features of the two corpora.
[0122] The method for calculating the voice similarity between the second corpus and the first corpus in the multiple service records may include: calculating the voice similarity according to the preset weights of the similarity of the obtained text information and the similarity of the voiceprint information, as well as the similarity of the text information and the similarity of the voiceprint information.
[0123] Step 1062: Query the second corpus in the multiple service records whose voice similarity is greater than or equal to a preset threshold.
[0124] In an embodiment of the present invention, for example, the preset threshold includes 95%. For example, when querying the second corpus in the multiple service records and the voice similarity between the second corpus and the first corpus is 100%, it indicates that the first corpus and the second corpus are completely matched.
[0125] In an embodiment of the present invention, when the second corpus is text information, as another alternative solution, the specific execution process of step 106 may include: querying the second corpus in the multiple service records, where the second corpus contains the first corpus.
[0126] In the embodiments of the present invention, there are two cases where the second corpus contains the first corpus. In an alternative solution, the second corpus is exactly the same as the first corpus. For example, if the input first corpus is "Call Zhang San", the retrieved matching second corpus is "Call Zhang San". In another alternative solution, the prefixes of the first corpus and the second corpus are the same. For example, if the input first corpus is "Call Zhang San", the retrieved matching second corpus is "I want to call Zhang San".
[0127] It should be noted that, compared with the related art, step 106 can obtain the semantic state information corresponding to the second corpus from multiple service records without performing the steps of speech recognition and semantic understanding, thereby reducing the system performance overhead and system power consumption.
[0128] Step 108: Provide a corresponding service according to the semantic state information corresponding to the second corpus.
[0129] In the embodiments of the present invention, the DM module provides a corresponding service according to the semantic state information corresponding to the second corpus. Specifically, after selecting a corresponding service, the execution module 118 executes the corresponding service, and the display module 119 displays the execution result of the service.
[0130] For example, the language state information corresponding to the second corpus includes: the intention is to make a call, and the slot is Zhang San. Then, the DM module 116 performs business logic processing according to the intention and the slot to generate an operation instruction. The operation instruction may include querying a number, a broadcast word, and executing a call. The execution module 118 completes querying the number, the broadcast word, and executing the call according to the call instruction, thereby completing the service process.
[0131] Step 110: Input the first corpus into the NLU model to obtain the semantic state information corresponding to the first corpus output by the NLU model.
[0132] In the embodiments of the present invention, the NLU module includes a method for obtaining the semantic state information corresponding to the first corpus. One is the NLU model, where the NLU model is a pre-trained semantic understanding module. Inputting the first corpus into the NLU model can obtain the semantic state information corresponding to the output first corpus. The second is the rule engine, that is, the semantic state information corresponding to the first corpus is obtained through the rule engine in the NLU module. The present invention does not limit the method for obtaining the semantic state information corresponding to the first corpus.
[0133] Step 112: Provide a corresponding service according to the semantic state information corresponding to the first corpus.
[0134] In the embodiments of the present invention, the execution process of this step can refer to the above step 108.
[0135] In an embodiment of the present invention, after step 112, the following is further included:
[0136] Step 113: Obtain the corpus, semantic status information, and execution result corresponding to the service, and use the corpus, semantic status information, and execution result corresponding to the service as a service record, and continue to execute step 103c.
[0137] In an embodiment of the present invention, by executing step 113, the service corresponding to the second corpus that fails to match successfully can be used as a service record, that is, the service record is enriched through the self-learning process, and the subsequent voice service processing efficiency is improved.
[0138] The voice service processing method provided by the embodiment of the present invention obtains the input first corpus, queries the second corpus that matches the first corpus from the obtained multiple service records, and obtains the semantic status information corresponding to the second corpus. According to the semantic status information corresponding to the second corpus, the corresponding service is provided. That is to say, if the second corpus that matches the first corpus can be queried from the obtained multiple service records, and the semantic status information corresponding to the second corpus is obtained, it is not necessary to analyze the first corpus through the NLU module to obtain the output voice status information, and directly send the semantic status information to the DM module, so that the DM module provides the corresponding service according to the semantic status information corresponding to the second corpus. If the second corpus that matches the first corpus cannot be queried from the obtained multiple service records, the first corpus needs to be input into the NLU model to obtain the semantic status information corresponding to the first corpus output by the NLU model.
[0139] The following two embodiments are used to illustrate the two processing methods in the above voice service processing process:
[0140] Example 1: The user says "Call Zhang San" for the first time
[0141] Such as Figure 2As shown in the figure, the receiving module 111 outputs the first corpus "Call Zhang San" of the user input it obtains to the matching module 112. Since no second corpus corresponding to the first corpus is matched, the ASR module 114 is used to identify the first corpus to generate text information. And the text information is input into a pre-trained NLU model for training to obtain the semantic state information corresponding to the first corpus output by the NLU model. For example, the voice state information includes an intent and a slot. Among them, the intent is to make a call, and the slot is Zhang San. The intent and the slot are input into the DM module so that the DM module 116 provides corresponding services according to the semantic state information corresponding to the first corpus. For example, the DM module 116 performs business logic processing according to the intent and the slot to generate an operation instruction. Among them, the operation instruction may include querying a number, a broadcast word, and executing a call. The execution module 118 completes querying the number, the broadcast word, and executing the call according to the call instruction, thereby completing the business process. Further, after the execution of the service ends, the corpus, intent, slot, and execution result corresponding to the service are used as a service record and stored in the end-side database, and the above step 103c is continued to be executed.
[0142] Example 2: The user says "Call Zhang San" for the second time
[0143] As Figure 2 As shown in the figure, the receiving module 111 inputs the first corpus "Call Zhang San" of the input it obtains to the matching module 112. The matching module 112 obtains multiple service records from the storage module 113, queries out the second corpus that matches the first corpus among the multiple service records, and obtains the semantic state information corresponding to the second corpus. The semantic state information corresponding to the second corpus is input into the DM module so that the DM module 116 provides corresponding services according to the semantic state information corresponding to the second corpus. That is to say, in Example 2, it is not necessary to input the text information into a pre-trained NLU model for training to obtain the semantic state information corresponding to the first corpus output by the NLU model, thereby solving the problem in the related art that the steps of voice recognition and semantic understanding need to be performed in the process of voice service processing, resulting in large performance overhead and affecting the system power consumption, and at the same time improving the processing efficiency of voice services.
[0144] The embodiments of the present invention are applied to the fields of artificial intelligence and speech recognition technology. By obtaining the input first corpus, querying out the second corpus that matches the first corpus among the obtained multiple service records, and obtaining the semantic state information corresponding to the second corpus, and providing corresponding services according to the semantic state information corresponding to the second corpus, the problem of needing to perform voice recognition and semantic understanding in the process of voice service processing is avoided, thereby reducing the system performance overhead and system function and improving the processing efficiency of voice services.
[0145] Figure 6 It is a schematic block diagram of an electronic device 110 provided by an embodiment of the present invention. It should be understood that the electronic device 110 can execute Figure 3 and Figure 4 each step in the voice service processing method. To avoid repetition, it will not be described in detail here. As Figure 6 shown, the electronic device 110 includes: a processing unit 401 and an execution unit 402.
[0146] The processing unit 401 is used to obtain the input first corpus; query the second corpus that matches the first corpus in the obtained multiple service records, and obtain the semantic status information corresponding to the second corpus.
[0147] The execution unit 402 is used to provide corresponding services according to the semantic status information corresponding to the second corpus.
[0148] In an embodiment of the present invention, the first corpus includes original voice information or text information.
[0149] In an embodiment of the present invention, the processing unit 401 is further used to obtain the corpus, semantic status information, and execution results corresponding to multiple voice services; use the corpus, semantic status information, and execution results corresponding to the voice service as a service record to generate multiple service records.
[0150] In an embodiment of the present invention, the processing unit 401 is further used to perform positive and negative feedback learning on multiple service records with the same semantic status information according to the positive and negative feedback learning algorithm, and determine the processing results of multiple service records with the same semantic status information. The processing results include adopting multiple service records or deleting multiple service records.
[0151] In an embodiment of the present invention, the processing unit 401 is further used to count the number of successful or failed times of multiple service records with the same semantic status information; if it is determined that the number of successful times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as adopting multiple service records; if it is determined that the number of failed times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as deleting multiple service records.
[0152] In an embodiment of the present invention, the semantic status information includes an intention and a slot, or an intention, a slot, and context information.
[0153] In an embodiment of the present invention, the processing unit 401 is further used to calculate the voice similarity between the second corpus and the first corpus in the multiple service records; query the second corpus in the multiple service records whose voice similarity is greater than or equal to a preset threshold.
[0154] In an embodiment of the present invention, the processing unit 401 is further configured to query the second corpus from the multiple service records, where the second corpus includes the first corpus.
[0155] In an embodiment of the present invention, the processing unit 401 is further configured to, if no second corpus matching the first corpus is queried from the multiple service records obtained, input the first corpus into the NLU model to obtain semantic state information corresponding to the first corpus output by the NLU model.
[0156] The execution unit 402 is further configured to provide corresponding services according to the semantic state information corresponding to the first corpus.
[0157] It should be understood that the electronic device 110 here is embodied in the form of functional units. The term "unit" here can be implemented in software and / or hardware forms, and no specific limitation is made thereto. For example, a "unit" can be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, and a memory, a merged logic circuit, and / or other suitable components that support the described functions.
[0158] Therefore, the units of the various examples described in the embodiments of the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0159] An embodiment of the present invention further provides an electronic device, which can be a terminal device or a circuit device built in the terminal device. This device can be used to execute the functions / steps in the above method embodiments.
[0160] Figure 7 As shown in the structural schematic diagram of an electronic device provided in an embodiment of the present invention, Figure 7 as shown, the electronic device 900 includes a processor 910 and a transceiver 920. Optionally, the electronic device 900 may further include a memory 930. Among them, the processor 910, the transceiver 920, and the memory 930 can communicate with each other through an internal connection path to transmit control and / or data signals. The memory 930 is used to store a computer program, and the processor 910 is used to call and run the computer program from the memory 930.
[0161] Optionally, the electronic device 900 may further include an antenna 940 for transmitting the wireless signals output by the transceiver 920.
[0162] The above-mentioned processor 910 and the memory 930 may be integrated into a processing device, and more commonly, they are independent components. The processor 910 is used to execute the program code stored in the memory 930 to implement the above functions. Specifically, in implementation, the memory 930 may also be integrated in the processor 910, or independent of the processor 910. The processor 910 may correspond to Figure 6 the processing unit 401 in the electronic device 110.
[0163] In addition, to make the functions of the electronic device 900 more complete, the electronic device 900 may further include one or more of an input unit 960, a display unit 970, an audio circuit 980, a camera 990, and a sensor 901, etc. The audio circuit may further include a speaker 982, a microphone 984, etc.
[0164] Optionally, the above-mentioned electronic device 900 may further include a power supply 950 for supplying power to various devices or circuits in the terminal device.
[0165] It should be understood that Figure 7 the shown electronic device 900 can implement Figure 3 and Figure 4 each process of the method embodiments shown. The operations and / or functions of each unit in the electronic device 900 are respectively for implementing the corresponding processes in the above method embodiments. Specifically, reference may be made to the descriptions in the above method embodiments. To avoid repetition, the detailed descriptions are appropriately omitted here.
[0166] It should be understood that Figure 7 the processor 910 in the shown electronic device 900 may be a system on a chip (SOC). The processor 910 may include a central processing unit (CPU), and may further include other types of processors. The CPU may be called the main CPU. Each part of the processors cooperate to implement the previous method process, and each part of the processors may selectively execute a part of the software driver program.
[0167] In summary, each part of the processors or processing units inside the processor 910 may cooperate together to implement the previous method process, and the corresponding software programs of each part of the processors or processing units may be stored in the memory 930.
[0168] The present invention also provides a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute each step in the voice service processing method as described above Figure 3 and Figure 4 shown.
[0169] In the above embodiments, the involved processor 910 may include, for example, a central processing unit (CPU), a microprocessor, a microcontroller or a digital signal processor, and may also include a GPU, an NPU and an ISP. The processor may further include necessary hardware accelerators or logic processing hardware circuits, such as an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the technical solution of the present invention. In addition, the processor may have the function of operating one or more software programs, and the software programs may be stored in the memory.
[0170] The memory may be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or may also be any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.
[0171] In the embodiments of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the relationship between related objects and indicates that there can be three relationships. For example, A and / or B can represent the cases where A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the related objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0172] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0173] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0174] In several embodiments provided by the present invention, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0175] The above is only the specific implementation manner of the present invention. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention and should be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for processing voice services, characterized in that Including: Obtain the input first corpus; Obtain the corpus, semantic status information, and execution results corresponding to multiple voice services; Take the corpus, semantic status information, and execution results corresponding to the voice service as a service record to generate multiple service records; Count the number of successful or failed times of multiple service records with the same semantic status information; If it is determined that the number of successful times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as adopting multiple service records; If it is determined that the number of failed times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as deleting multiple service records; Query in the obtained multiple service records for a second corpus that matches the first corpus, and obtain the semantic status information corresponding to the second corpus; Provide corresponding services according to the semantic status information corresponding to the second corpus.
2. The method according to claim 1, wherein the first corpus includes original voice information or text information.
3. The method according to claim 1, wherein The semantic status information includes intent and slot, or intent, slot, and context information.
4. The method according to claim 1, wherein The querying in the obtained multiple service records for a second corpus that matches the first corpus includes: Calculate the voice similarity between the second corpus and the first corpus in the multiple service records; Query in the multiple service records for a second corpus whose voice similarity is greater than or equal to a preset threshold.
5. The method according to claim 1, wherein The querying in the obtained multiple service records for a second corpus that matches the first corpus includes: Query in the multiple service records for the second corpus that contains the first corpus.
6. An electronic device, characterized in that, Including: One or more processors; A memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to perform the following steps: Obtain the input first corpus; Obtain the corpus, semantic status information, and execution results corresponding to multiple voice services; Take the corpus, semantic status information, and execution results corresponding to the voice service as a service record to generate multiple service records; Count the number of successful or failed times of multiple service records with the same semantic status information; If it is determined that the number of successful times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as adopting multiple service records; If it is determined that the number of failed times is greater than or equal to a preset number, determine the processing result of multiple service records with the same semantic status information as deleting multiple service records; Query in the obtained multiple service records for a second corpus that matches the first corpus, and obtain the semantic status information corresponding to the second corpus; Provide corresponding services according to the semantic status information corresponding to the second corpus.
7. The device according to claim 6, wherein the first corpus includes original voice information or text information.
8. The device according to claim 6, characterized in that, The semantic state information includes an intention and slots, or an intention, slots, and context information.
9. The device according to claim 6, characterized in that, When the instruction is executed by the device, the device specifically performs the following steps: Calculate the acoustic similarity between the second corpus and the first corpus in the multiple business records; Query, from the multiple business records, the second corpus whose acoustic similarity is greater than or equal to a preset threshold.
10. The device according to claim 6, characterized in that, When the instruction is executed by the device, the device specifically performs the following steps: Query, from the multiple business records, the second corpus that contains the first corpus.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the voice service processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speech guidance method and device, electronic device, and storage medium
CN109325097A