Voice interaction method, terminal, platform, electronic equipment and program product
By using offline natural language processing model on the smart terminal to perform semantic analysis of the voice stream and switching to the cloud service platform for online analysis when the confidence is low, the problem of high concurrency pressure on the server side during the voice interaction process of smart devices is solved, and efficient and accurate voice interaction is achieved.
Patent Information
- Application Number
- CN202311813296.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
In the process of voice recognition or voice comprehension, existing smart devices usually only support online interaction, resulting in high concurrency pressure on the server and hindering the popularization of Internet of Things voice services.
By using offline natural language processing model on the smart terminal to perform semantic analysis of the voice stream, if the confidence is low, the voice stream is transmitted to the cloud service platform for online analysis, obtain the target device function information and feed it back to the terminal.
It realizes automatic switching of voice processing between the terminal and the cloud, improves interaction response speed and accuracy, reduces network transmission pressure and cloud resource requirements, and improves user experience and system operation efficiency.
Smart Images

Figure CN120220665A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to a voice interaction method, a terminal, a platform, an electronic device, and a program product. Background Art
[0002] Currently, the intelligent devices used in the market generally only support the online interaction method during voice recognition or voice understanding. That is, during each interaction, the voice request has to request the cloud service. In the case of a high number of device connections, the concurrent pressure on the server side is extremely high, which poses a huge obstacle to the popularization of the voice services required by the subsequent Internet of Things.
[0003] Therefore, how to better perform voice interaction has become an urgent problem in the industry. Summary of the Invention
[0004] Embodiments of the present application provide a voice interaction method, a terminal, a platform, an electronic device, and a program product to solve the technical problem of how to better perform voice interaction.
[0005] In a first aspect, an embodiment of the present application provides a voice interaction method applied to an intelligent terminal, including:
[0006] When a voice stream received by the intelligent terminal contains a voice end identifier, semantic parsing is performed on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic parsing result;
[0007] When the first confidence level of the first semantic parsing result is less than a first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform;
[0008] When the second semantic parsing result obtained by parsing the voice stream through the online natural language processing model of the cloud service platform contains the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
[0009] In one embodiment, after the step of the intelligent terminal receiving the device function information transmitted by the cloud service platform, the method further includes:
[0010] The intelligent terminal sends a device capability report to the cloud service platform; the device capability report includes at least one of the following: device name, product ID, transaction ID, service data; the service data includes the target device function information;
[0011] The intelligent terminal receives the device capability report feedback information sent by the cloud service platform; the device capability report feedback information includes at least one of the following: transaction ID, functional data; wherein, the functional data includes: the second semantic parsing result.
[0012] Update the offline natural language processing model of the intelligent terminal according to the second semantic parsing result and the corresponding speech stream in the device capability report feedback information.
[0013] In one embodiment, before the step that the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information, the method further includes:
[0014] In the case that the voice end identifier is not detected in the speech stream received by the intelligent terminal within the first speech duration, transmit the speech stream to the cloud service platform.
[0015] In one embodiment, the calculation method of the first confidence level of the first semantic parsing result specifically includes:
[0016] Input the audio acoustic feature vector of the speech stream into a preset encoder to obtain a first acoustic feature;
[0017] Input the first corpus feature vector set constructed based on the first corpus template set of the intelligent terminal into the preset encoder to obtain a second corpus feature vector set;
[0018] Determine the first confidence level based on the matching degree between the first acoustic feature and the corpus feature set.
[0019] In a second aspect, an embodiment of the present application provides a voice interaction method, which is applied to a cloud service platform and includes:
[0020] In the case that the speech stream received by the intelligent terminal does not contain a voice end identifier and the speech length of the speech stream is greater than a second preset threshold, or the confidence level of the first semantic parsing result obtained by the intelligent terminal through semantic parsing of the speech stream by the offline natural language processing model is less than a first preset threshold, receive the speech text corresponding to the speech stream sent by the intelligent terminal;
[0021] The cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic parsing result;
[0022] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than a third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
[0023] In one embodiment, after the step of sending the target device function information corresponding to the target corpus template to the intelligent terminal when the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is higher than a third preset threshold, the method further includes:
[0024] When receiving the device capability report sent by the intelligent terminal, write the second semantic parsing result into the function data of the device capability report feedback information;
[0025] Feed back the device capability report feedback information to the intelligent terminal for the intelligent terminal to update the offline natural language processing model according to the second semantic parsing result in the device capability report feedback information.
[0026] In one embodiment, after the step of the cloud service platform performing semantic parsing on the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic parsing result, it further includes:
[0027] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is less than or equal to the third preset threshold, obtain keywords from the second semantic parsing result through the AC automaton algorithm, and construct a target sentence pattern template according to the keywords;
[0028] Obtain the template similarity between the target sentence pattern template and each corpus template in the second corpus template set;
[0029] When at least one of the template similarities is greater than a fourth preset threshold, generate corpus template guiding information based on the target corpus template, and send the corpus template guiding information and the target device function information corresponding to the target corpus template to the intelligent terminal;
[0030] Wherein, the target corpus template is the corpus template with the highest template similarity with the target sentence pattern template in the second corpus template set, and the corpus template guiding information includes the target corpus template.
[0031] In one embodiment, after the step of obtaining the template similarity between the target sentence pattern template and each corpus template in the second corpus template set, it further includes:
[0032] In the case that there is no corpus template in the second corpus template set whose similarity to the target sentence pattern template is greater than the fourth preset threshold, unrecognized information is generated and the unrecognized information is sent to the intelligent terminal.
[0033] In a third aspect, an embodiment of the present application provides an intelligent terminal, including a memory, a transceiver, and a processor;
[0034] The memory is used for storing a computer program; the transceiver is used for transceiving data under the control of the processor; the processor is used for reading the computer program in the memory and performing the following operations:
[0035] In the case that the confidence level of the first semantic parsing result obtained by the intelligent terminal through the offline natural language processing model for semantic parsing of the speech stream is less than the first preset threshold, the speech text corresponding to the speech stream sent by the intelligent terminal is received;
[0036] In the case that the first confidence level of the first semantic parsing result is less than the first preset threshold, the intelligent terminal transmits the speech text corresponding to the speech stream to the cloud service platform;
[0037] In the case that the second semantic parsing result obtained by the cloud service platform through the online natural language processing model for parsing the speech stream includes the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
[0038] In a fourth aspect, an embodiment of the present application provides a cloud service platform, including a memory, a transceiver, and a processor;
[0039] The memory is used for storing a computer program; the transceiver is used for transceiving data under the control of the processor; the processor is used for reading the computer program in the memory and performing the following operations:
[0040] In the case that the speech stream received by the intelligent terminal does not include a speech end identifier and the speech length of the speech stream is greater than the second preset threshold, or in the case that the confidence level of the first semantic parsing result obtained by the intelligent terminal through the offline natural language processing model for semantic parsing of the speech stream is less than the first preset threshold, the speech text corresponding to the speech stream sent by the intelligent terminal is received;
[0041] The cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through the online natural language processing model to obtain a second semantic parsing result;
[0042] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than a third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
[0043] In a fifth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory storing a computer program. When the processor executes the program, the steps of the voice interaction method described in the first aspect or the second aspect are implemented.
[0044] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the voice interaction method described in the first aspect or the second aspect are implemented.
[0045] In the voice interaction method, terminal, platform, electronic device and program product provided by the embodiments of the present application, when it is detected that the voice stream received by the intelligent terminal contains a voice end identifier, it indicates that the length of the voice stream is not too long at this time, and the terminal usually has parsing capabilities. Therefore, the voice stream is preferentially semantically parsed by the offline natural language processing model of the intelligent terminal on the intelligent terminal side. However, if the first confidence level of the first semantic parsing result is less than the first preset threshold, it means that the detection result of the intelligent terminal through offline detection is not highly reliable. The voice stream needs to be transmitted to the cloud service platform after being appended with an unfinished identifier, and the voice stream is parsed by the online natural language processing model of the cloud service platform to obtain a second semantic parsing result with higher credibility. If the target device function information corresponding to the intelligent terminal is included in the second semantic parsing result, the target device function information can be fed back to the intelligent terminal so that the intelligent terminal can execute the device function according to the target device function information. The solution of the present application can automatically switch the processing of the voice stream between the intelligent terminal and the cloud service platform, so that for the processing of voice, simple voice commands can be recognized and parsed on the intelligent terminal side, and when the terminal-side recognition fails to meet the requirements, it automatically switches to the cloud service platform-side recognition. This can not only improve the interaction response speed, but also ensure the accuracy of the interaction, while reducing the network transmission pressure and cloud resource requirements, improving the user interaction experience and reducing the system operation cost at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 One of the schematic flowcharts of the voice interaction method provided by the embodiments of the present application;
[0048] Figure 2 Schematic flowchart of the voice processing provided by the embodiments of the present application;
[0049] Figure 3 Second of the schematic flowcharts of the voice interaction method described in the embodiments of the present application;
[0050] Figure 4 Flowchart of voice request processing in the embodiments of the present application;
[0051] Figure 5 Schematic diagram of the structure of the intelligent terminal according to the embodiments of the present application;
[0052] Figure 6 Schematic diagram of the structure of the cloud service platform according to the embodiments of the present application;
[0053] Figure 7 Schematic diagram of the structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0055] Figure 1 One of the schematic flowcharts of the voice interaction method provided by the embodiments of the present application, as Figure 1 shown, may include:
[0056] Step 110, when a voice end identifier is detected within the first voice duration of the voice stream received by the intelligent terminal, performing semantic parsing on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic parsing result
[0057] In the embodiments of the present application, the intelligent terminal may specifically refer to a terminal device such as a smart phone, a smart watch, a smart speaker, a smart TV, a smart home device, etc., which has communication capabilities and data processing capabilities.
[0058] In the embodiments of the present application, the voice end identifier generally refers to a specific signal or marker used to identify the end of a voice segment in voice processing. In speech recognition, the voice end identifier can help the system determine the start and end times of the voice input, so as to more accurately perform the processing of converting speech to text.
[0059] In an optional embodiment, the voice end identifier can be determined through voice activity detection (VAD). VAD can help identify the valid part of the voice signal, thereby effectively determining the semantic end identifier of the voice.
[0060] The specific form and usage method of the voice end identifier may vary due to different systems and applications. It may involve technical means such as acoustic features, energy thresholds, silence detection, and VAD algorithms to judge the start and end of the voice. The intelligent terminal usually parses and judges the voice stream according to pre-set rules or algorithms.
[0061] In the embodiments of the present application, the first voice duration can be a pre-set voice duration, which is the duration calculated from the start of the voice. Specifically, it can be 2 seconds, 3 seconds, etc.
[0062] In the embodiments of the present application, if the voice stream received by the intelligent terminal contains a voice end identifier, this usually means that there is a clear signal or marker in the voice stream indicating the end point of the voice. At this time, it means that the voice issued by the user has been completed, and within the first voice duration, the voice end identifier is detected, indicating that the voice duration of this voice stream will not be too long. Therefore, the voice text corresponding to the voice stream can be directly semantically parsed through the offline natural language processing model in the intelligent terminal.
[0063] In the embodiments of the present application, the offline natural language processing model refers to a model that performs natural language processing locally on the intelligent terminal device, rather than relying on the Internet or a remote server. The advantage of this model is that it can perform speech recognition, text analysis, and semantic understanding without a network connection, which helps to improve the response speed of the device and protect user privacy.
[0064] More specifically, the offline natural language processing model described in the embodiments of the present application can be updated and trained regularly according to the data transmitted by the cloud service platform or other data sources.
[0065] In the embodiments of the present application, after the intelligent terminal receives the voice stream, it will first use speech recognition technology to convert the received voice stream into text form, and then the natural language processing model will process the converted text, including lexical analysis, syntactic analysis, and semantic analysis and other links, and finally obtain the first semantic parsing result.
[0066] Step 120, when the first confidence level of the first semantic parsing result is less than a first preset threshold, the intelligent terminal transmits the speech text corresponding to the speech stream to the cloud service platform;
[0067] In the embodiment of the present application, after obtaining the first semantic parsing result, the confidence level of the first semantic parsing result can be further verified by combining the corpus feature vector set constructed by the first corpus template set of the intelligent terminal.
[0068] In the embodiment of the present application, the first preset threshold can be a threshold set by the user according to needs. When the first confidence level of the first semantic parsing result is less than the first preset threshold, it indicates that the reliability of the first semantic parsing result is not high at this time, and it is necessary to request the cloud service platform to process it.
[0069] Therefore, the intelligent terminal can further mark the speech text corresponding to the speech stream with an unfinished identifier and transmit it to the corresponding cloud service platform, so that after receiving the speech text corresponding to the speech stream, the cloud service platform parses the speech stream through an online natural language processing model.
[0070] In the embodiment of the present application, the cloud service platform may specifically refer to a service providing platform for cloud computing technology, which provides various computing resources and services to users through the Internet. In the embodiment of the present application, the cloud service platform may specifically be a cloud service platform in communication connection with the intelligent terminal, and this cloud service platform can provide an online natural language processing model that is continuously updated to ensure the recognition accuracy of the online natural language processing model.
[0071] In the embodiment of the present application, after the online natural language processing model processes the speech stream, it will output a second semantic parsing result. At this time, the second semantic parsing result has a high accuracy, so there is no need to perform confidence verification again.
[0072] After obtaining the second semantic parsing result, it can be further matched with the device function information corresponding to the intelligent terminal. The device function information may specifically include device hot words and function hot words. The device hot words may include: <room>: Living room | Kitchen | Bedroom | Second bedroom | Washroom | Balcony; <device>: Socket | Lamp | Table lamp | Switch | Curtain; The functional verbs can include: <open>: Open | Turn on | On; <close>: Close|Shut|Shut down.
[0073] In the embodiment of the present application, when the second semantic parsing result contains at least one device hot word and at least one function hot word corresponding to the intelligent terminal, it can be determined that the second semantic parsing result contains the target device function information corresponding to the intelligent terminal.
[0074] The target device function information described in the embodiment of the present application may include at least one device hot word and at least one function hot word.
[0075] The cloud server will further feedback the target device function information to the corresponding intelligent terminal.
[0076] Step 130, when the second semantic parsing result obtained by the cloud service platform through the online natural language processing model for parsing the speech stream contains the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
[0077] In the embodiment of the present application, after the intelligent terminal receives the target device function information transmitted by the cloud service platform, it can further implement the corresponding device function according to at least one device hot word and at least one function hot word in the target device function information.
[0078] For example, when the feedback target device function information is <room>: Living room <device>: Curtain; <open>: When it is turned on, the corresponding function is executed to open the curtain of the object.
[0079] In another optional embodiment, when the intelligent terminal detects network fluctuations or detects that the intelligent terminal is in an offline state, it still executes according to the first semantic parsing result. However, before executing the instruction, it can initiate a secondary confirmation to the user.
[0080] Thus, it can effectively ensure that the voice command can be normally recognized in the state of network fluctuations or device offline.
[0081] In the embodiment of the present application, when it is detected that the voice stream received by the intelligent terminal contains a voice end identifier, it indicates that the length of the voice stream is not too long at this time, and the terminal usually has parsing capabilities. Therefore, the voice stream is preferentially semantically parsed by the offline natural language processing model of the intelligent terminal on the intelligent terminal side. However, if the first confidence level of the first semantic parsing result is less than the first preset threshold, it indicates that the detection result of the intelligent terminal through offline detection is not highly reliable. The voice stream needs to be transmitted to the cloud service platform after being appended with an unfinished identifier, and the voice stream is parsed by the online natural language processing model of the cloud service platform to obtain a second semantic parsing result with higher credibility. And if the target device function information corresponding to the intelligent terminal is included in the second semantic parsing result, the target device function information can be fed back to the intelligent terminal so that the intelligent terminal can execute the device function according to the target device function information. The solution of the present application can automatically switch the processing of the voice stream between the intelligent terminal and the cloud service platform, so that for the processing of voice, simple voice commands can be recognized and parsed on the intelligent terminal side, and when the end-side recognition fails to meet the requirements, it automatically switches to the cloud service platform side for recognition. This can not only improve the interaction response speed, but also ensure the accuracy of the interaction, while reducing the network transmission pressure and cloud resource requirements, improving the user interaction experience and reducing the operation cost of the system at the same time.
[0082] Optionally, after the step of the intelligent terminal receiving the device function information transmitted by the cloud service platform, the method further includes:
[0083] The intelligent terminal sends a device capability report to the cloud service platform; the device capability report includes at least one of the following: device name, product ID, transaction ID, service data; the service data includes the target device function information;
[0084] The intelligent terminal receives the device capability report feedback information sent by the cloud service platform; the device capability report feedback information includes at least one of the following: transaction ID, function data; wherein, the function data includes: the second semantic parsing result.
[0085] Update the offline natural language processing model of the intelligent terminal according to the second semantic parsing result and the corresponding speech stream in the device capability report feedback information.
[0086] In the embodiment of the present application, after the step of the terminal receiving the device function information transmitted by the cloud service platform, the terminal will further attempt to update the offline natural language processing model, so that when encountering a speech stream that could not be recognized before, it can be normally recognized.
[0087] Therefore, the intelligent terminal can publish a message with the topic of device capability report xxx / [di] / deviceCapabilityReport, and the intelligent terminal can also subscribe in advance to the message of xxx / [di] / deviceCapabilityReportResp published by the corresponding cloud service platform.
[0088] The cloud service platform can subscribe in advance to the information of xxx / [di] / deviceCapabilityReport that the corresponding intelligent terminal may publish. Therefore, after the intelligent terminal publishes the device capability report, the cloud service platform can receive the device capability report published by the intelligent terminal.
[0089] The device capability report includes at least one of the following: device name deviceName, product ID productiid, transaction ID productiid, service data service[]; the service data will include the target device function information.
[0090] Specifically, it can be:
[0091]
[0092] After the cloud service platform receives the device capability report sent by the cloud service platform, it will determine the corresponding second semantic parsing result according to the target device function information in the report, and write the second semantic parsing result into the function data in the device capability report feedback information.
[0093] In an optional embodiment, the device capability report feedback information may specifically include:
[0094]
[0095] In the embodiment of the present application, the second semantic parsing result that the server hopes to synchronize to the terminal will be recorded in functionTem[], so that after the terminal device receives this information, it can update the offline natural language processing model according to the second semantic parsing result.
[0096] In the embodiments of the present application, the intelligent terminal reports the target device function information to the cloud service platform by itself, enabling the cloud service platform to write the corresponding second semantic parsing result into the function data according to the reported data in a timely manner, and synchronizing the feedback information of the device capability report to the intelligent terminal, so as to facilitate the model update of the offline natural language processing model in the offline terminal, thereby effectively ensuring the recognition accuracy of the offline natural language processing model.
[0097] Optionally, before the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information, the method further includes:
[0098] When the voice stream received by the intelligent terminal does not detect a voice end flag within the first voice duration, the voice stream is transmitted to the cloud service platform.
[0099] In the embodiments of the present application, when the voice end flag is not detected within the first voice duration, it indicates that the data volume involved in the current voice stream is large, and the offline natural language processing model of the intelligent terminal may not be able to effectively and accurately process it at this time. Therefore, in order to improve the processing efficiency, the currently cached voice stream can be transmitted to the cloud service platform for recognition at this time, and the voice stream before detecting the voice end flag and the previously cached voice stream are uploaded to the cloud service platform as a coherent voice stream.
[0100] After detecting the voice end flag of the voice stream, temporarily stop uploading the voice stream to the cloud service platform, but continue to determine whether the voice end flag can be detected in the new voice stream within the first voice duration.
[0101] In the embodiments of the present application, after receiving the offline voice, the cloud service platform will first perform text conversion on it, and then input it into the online natural language processing model in the cloud service platform for voice stream parsing to obtain a second semantic parsing result. If the second semantic parsing result contains the target device function information corresponding to the intelligent terminal, the target device function information will be sent to the intelligent terminal.
[0102] The intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
[0103] In the embodiments of the present application, when the received voice stream does not detect a voice end flag within the first voice duration, it indicates that the voice stream that cannot be effectively processed by the local intelligent terminal is transmitted to the cloud service platform for processing, which can effectively ensure the recognition accuracy.
[0104] Optionally, the method for calculating the first confidence level of the first semantic parsing result specifically includes:
[0105] Input the audio acoustic feature vector of the speech stream into a preset encoder to obtain a first acoustic feature;
[0106] Input the first corpus feature vector set constructed based on the first corpus template set of the intelligent terminal into the preset encoder to obtain a second corpus feature vector set;
[0107] Determine the first confidence level based on the matching degree between the first acoustic feature and the corpus feature set.
[0108] In the embodiments of the present application, the audio acoustic features of the speech stream may specifically include: Mel Frequency Cepstral Coefficients (MFCC), Audio Power Spectral Density (PSD), etc. These features can reflect the key information of the audio signal.
[0109] In the embodiments of the present application, after the audio acoustic feature vector of the speech stream is input into the preset encoder and processed by the encoder, the first acoustic feature corresponding to the speech stream can be obtained.
[0110] In the embodiments of the present application, the first corpus template set of the intelligent terminal may specifically include all corpus data received by the intelligent terminal from "function data" functionTem[] in history.
[0111] The first corpus template machine may specifically further include the corpus data carried in "function data" functionTem[] of the standard interface.
[0112] In the embodiments of the present application, after obtaining the first corpus template set, the audio feature vector of each corpus template in the first corpus template set can be further extracted, so as to construct a second corpus feature vector set according to the audio feature vectors of each corpus template.
[0113] The first acoustic feature can be further compared with each audio feature vector in the second corpus feature vector set to obtain the matching degree between the first acoustic feature and each audio feature vector.
[0114] In an optional embodiment, the highest value in the matching degree can be used as the first confidence level.
[0115] In another optional embodiment, the average value of each confidence level can also be used as the first confidence level.
[0116] In the embodiments of the present application, by comparing and analyzing the audio acoustic features of the voice stream and each corpus feature in the first corpus feature vector set constructed from the first corpus template set of the intelligent terminal, the first confidence level of the first semantic parsing result corresponding to the voice stream can be effectively determined.
[0117] In an alternative embodiment, Figure 2 is a schematic diagram of the voice processing flow provided by the embodiments of the present application, as Figure 2 shown, including:
[0118] After the intelligent terminal is activated, the intelligent terminal can receive a voice stream;
[0119] During the process of receiving the voice stream, it is determined whether vad is detected. If not detected and the time of the voice stream is not greater than 2 seconds, the voice stream continues to be received; if not detected and the time of the voice stream is greater than 2 seconds, the voice stream is transmitted to the cloud service platform for cloud voice recognition;
[0120] If vad is detected during the process of receiving the voice stream, the voice recognition result is processed by the offline natural language processing model in the intelligent terminal.
[0121] If the confidence level of the output result of the offline NLP model is not less than 0.5, the offline result is sent down.
[0122] If the confidence level of the output result of the offline NLP model is less than 0.5, after the voice stream is recognized as text, the recognized voice text is uploaded to the cloud service platform for cloud processing.
[0123] Then, the online NLP model recognizes the voice text and executes the corresponding skill processing logic.
[0124] Figure 3 is the second schematic diagram of the voice interaction method flow described in the embodiments of the present application, as Figure 3 shown, including:
[0125] Step 310, in the case where the confidence level of the first semantic parsing result of the semantic parsing of the voice stream by the offline natural language processing model in the intelligent terminal is less than the first preset threshold, receive the voice text corresponding to the voice stream sent by the intelligent terminal;
[0126] In the embodiments of the present application, the cloud service platform may specifically refer to a service providing platform for cloud computing technology, which provides various computing resources and services to users through the Internet. In the embodiments of the present application, the cloud service platform may specifically be a cloud service platform in communication connection with the intelligent terminal, and this cloud service platform can provide an online natural language processing model that is continuously updated to ensure the recognition accuracy of the online natural language processing model.
[0127] In an embodiment of the present application, when the confidence level of the first semantic parsing result obtained by the intelligent terminal through the offline natural language processing model for semantic parsing of the speech stream is less than the first preset threshold, it indicates that the reliability of the first semantic parsing result is not high at this time, and it is necessary to request the cloud service platform to process it.
[0128] Therefore, the intelligent terminal can further mark the speech text corresponding to the speech stream with an incomplete identifier and transmit it to the corresponding cloud service platform, so that after receiving the speech text corresponding to the speech stream, the cloud service platform parses the speech stream through the online natural language processing model.
[0129] Step 320, the cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through the online natural language processing model to obtain a second semantic parsing result;
[0130] The online natural language processing model described in the embodiments of the present application can specifically be a model that can be continuously updated in a timely manner by obtaining the latest data source. Therefore, the online natural language processing model can provide a high recognition accuracy.
[0131] In an embodiment of the present application, the speech text corresponding to the speech stream can be input into the online natural language processing model, and a second semantic parsing result will be output. The second semantic parsing result contains information such as the meaning, significance, and purpose of the text. This information may be used in subsequent business logics or human-computer interactions.
[0132] Step 330, when the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than the third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
[0133] In an embodiment of the present application, the second corpus template set of the cloud service platform can specifically be the corpus templates already stored in the cloud service platform. This second corpus template set can be obtained by synchronizing relevant corpus templates to the cloud service platform after the user completes relevant data settings through intelligent terminals such as mobile phones, or can be continuously updated.
[0134] More specifically, in an embodiment of the present application, the second corpus template set can also be a set of corpus templates corresponding to the category of the intelligent terminal that sends the speech stream to the cloud service platform.
[0135] In an embodiment of the present application, the second semantic parsing result can be further compared with the matching degree of each corpus template in the second corpus template set of the cloud service platform to obtain the matching degree of each corpus template and the second semantic parsing result.
[0136] When there is a corpus template with a matching degree greater than the third preset threshold among the matching degrees of each corpus template, it is determined that there is a corresponding corpus template in the second semantic parsing result that can be directly called at this time. Therefore, the corpus template with the largest matching degree can be further used as the target corpus template.
[0137] In an embodiment of the present application, target device function information often corresponds to the target corpus template, that is, it can include device hot words and function hot words. The device hot words can include: <room>:Living room | Kitchen | Bedroom | Second bedroom | Washroom | Balcony; <device>: Socket | Lamp | Table lamp | Switch | Curtain; The functional verbs can include: <open>: Open | Turn on | On; <close>: Close | Shut | Shut down.
[0138] The cloud service platform can further send the target device function information corresponding to the target corpus template to the corresponding intelligent terminal, so that the intelligent terminal can execute the corresponding device function according to the target device function information.
[0139] Optionally, when the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is higher than the third preset threshold, after the step of sending the target device function information corresponding to the target corpus template to the intelligent terminal, the method further includes:
[0140] When receiving the device capability report sent by the intelligent terminal, write the second semantic parsing result into the function data of the device capability report feedback information;
[0141] Feed back the device capability report feedback information to the intelligent terminal for the intelligent terminal to update the offline natural language processing model according to the second semantic parsing result in the device capability report feedback information.
[0142] In the embodiments of the present application, the intelligent terminal can publish a message with the topic of device capability report xxx / [di] / deviceCapabilityReport, and the intelligent terminal can also subscribe in advance to the message of xxx / [di] / deviceCapabilityReportResp published by the corresponding cloud service platform.
[0143] The cloud service platform can subscribe in advance to the information of xxx / [di] / deviceCapabilityReport that the corresponding intelligent terminal may publish. Therefore, after the intelligent terminal publishes the device capability report, the cloud service platform can receive the device capability report published by the intelligent terminal.
[0144] The device capability report includes at least one of the following: device name deviceName, product ID productiid, transaction ID productiid, service data service[]; the service data will include the target device function information.
[0145] Specifically, it can be:
[0146]
[0147] After the cloud service platform receives the device capability report sent by the cloud service platform, it will determine the corresponding second semantic parsing result according to the target device function information in the report, and write the second semantic parsing result into the function data in the device capability report feedback information.
[0148] In an optionally implemented embodiment, the device capability report feedback information may specifically include:
[0149]
[0150] In the embodiment of the present application, the second semantic parsing result that the server hopes to synchronize to the terminal is recorded in functionTem[], so that after receiving this information, the terminal device can update the offline natural language processing model according to the second semantic parsing result.
[0151] In the embodiment of the present application, the intelligent terminal reports the target device function information to the cloud service platform by itself, so that the cloud service platform can write the corresponding second semantic parsing result into the function data according to the reported data even if, and synchronize it to the intelligent terminal through the device capability report feedback information, so as to update the model of the offline natural language processing model in the offline terminal, thereby effectively ensuring the recognition accuracy of the offline natural language processing model.
[0152] Optionally, after the step of the cloud service platform performing semantic parsing on the speech text corresponding to the speech stream through the online natural language processing model to obtain the second semantic parsing result, it further includes:
[0153] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is less than or equal to the third preset threshold, keywords are obtained from the second semantic parsing result through the AC automaton algorithm, and a target sentence pattern template is constructed according to the keywords;
[0154] Obtain the template similarity between the target sentence pattern template and each corpus template in the second corpus template set;
[0155] When at least one of the template similarities is greater than the fourth preset threshold, corpus template guiding information is generated based on the target corpus template, and the corpus template guiding information and the target device function information corresponding to the target corpus template are sent to the intelligent terminal;
[0156] Wherein, the target corpus template is the corpus template with the highest template similarity to the target sentence pattern template in the second corpus template set, and the corpus template guiding information includes the target corpus template.
[0157] In the embodiment of the present application, when the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is less than or equal to the third preset threshold, it indicates that it is difficult to match the corresponding corpus template through the second semantic parsing result at this time, and it is difficult to determine the corresponding target device function information.
[0158] Therefore, the AC automaton method can be further used to extract the key slots of the sentence pattern, splice them according to the grammatical structure, and then compare them with the template set to obtain the closest semantic expression. And in the form of prompts, the user is guided to express in the correct sentence
[0159] In the embodiment of the present application, the AC automaton algorithm can effectively extract keywords from the second semantic parsing result, and then construct the target sentence pattern template only according to the extracted keywords.
[0160] In the embodiment of the present application, the corresponding sample feature matrix representation can be further constructed according to the constructed target sentence pattern template; specifically, it can be according to For the j-th sample representation matrix of the i-th category Use the activation function squash for non-linear transformation, and then calculate the weighted similarity of the transformed template sentence pattern to judge the similarity between the sentence pattern template generated by the AC keyword and each corpus template in the second corpus template set.
[0161] In an optional embodiment, the average of k samples in the template set is taken to represent the class vector
[0162] Where s represents the generated sentence pattern The similarity between the template and the selected template in the template set, Y j Represents the number of overlapping keywords between the generated template and the template set, and T represents the number of keywords included in the generated template.
[0163] In the embodiment of the present application, when the template similarity of at least one of the templates is greater than the fourth preset threshold, it indicates that there may be a close sentence pattern template in the second semantic parsing result. At this time, it may be due to the user's unclear expression.
[0164] Therefore, the corpus template with the highest template similarity to the target sentence pattern template in the second corpus template set can be used as the target corpus template, and then the corpus template guidance information is generated according to the target corpus, and the corpus template guidance information is sent to the intelligent terminal; thus helping the user to perform voice control correctly.
[0165] In an optional embodiment, when the cloud service platform sends the corpus template guidance information to the intelligent terminal, it will also send the target device function information corresponding to the target corpus template, so that the intelligent terminal can directly execute the target device function information without the user having to repeat it again, which can effectively improve the user's voice control efficiency.
[0166] For example, User: It's too dark. Don't save power and turn on the kitchen light.
[0167] The intelligent terminal will control the kitchen light to turn on and feedback to the user: Okay, from now on, you can directly say "Turn on the kitchen light" (corpus template guiding information), thus helping the user with voice control.
[0168] In the embodiment of the present application, when it is impossible to directly and effectively match a corpus template with a relatively high credibility in the second corpus template set, the keywords can be further extracted to construct a target sentence pattern template, and then, through an approximate fusion matching method, the template similarity between the target sentence pattern template and each corpus template in the second corpus template set is compared, so as to effectively screen out the target sentence pattern template that the user really wants to match. Then, the corpus template guiding information and the target device function information corresponding to the target corpus template are sent to the intelligent terminal, and the user can be effectively helped to give effective feedback through the corpus template guiding information.
[0169] Optionally, after the step of obtaining the template similarity between the target sentence pattern template and each corpus template in the second corpus template set, the following is further included:
[0170] In the case that there is no corpus template in the second corpus template set whose similarity to the target sentence pattern template is greater than the fourth preset threshold, an unrecognized information is generated and sent to the intelligent terminal.
[0171] In the embodiment of the present application, in the case that there is no corpus template in the second corpus template set whose similarity to the target sentence pattern template is greater than the fourth preset threshold, it means that an effective corpus template cannot be effectively matched at this time and the instruction cannot be processed. At this time, only the unrecognized information can be sent to the intelligent terminal.
[0172] Optionally, before the step that the cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic parsing result, the following is further included:
[0173] In the case that the speech stream received by the intelligent terminal does not detect a speech end flag within the first speech duration, the speech stream transmitted by the intelligent terminal is received;
[0174] The cloud service platform performs text recognition on the speech stream to obtain the speech text corresponding to the speech stream.
[0175] In the embodiment of the present application, in the case that the speech stream received by the intelligent terminal does not detect a speech end flag within the first speech duration, it means that the speech duration of the speech stream is relatively long and the data volume of the involved speech stream is also large. At this time, the computing power of the intelligent terminal may not be able to accurately process it. Therefore, the speech stream can be transmitted to the cloud service platform.
[0176] The cloud service platform performs text recognition to obtain the speech text corresponding to the speech stream, and then the cloud service platform performs semantic analysis on the speech stream and the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic analysis result, and performs subsequent processing based on the second semantic analysis result.
[0177] In the embodiment of the present application, by forwarding the voice stream with a longer voice duration to the cloud service platform for processing, it is possible to effectively avoid the smart terminal from facing a large data pressure and effectively ensure the hierarchical efficiency of the processing.
[0178] Figure 4 This is a flow chart of voice request processing in an embodiment of the present application, such as Figure 4 As shown, it includes: after receiving the voice request, the intelligent terminal can perform semantic analysis on it through the offline NLP model to obtain a first semantic analysis result, and further determine whether the confidence of the first semantic analysis result is less than 0.5. If it is not less than 0.5, the terminal side executes the control logic;
[0179] If it is less than 0.5, the cloud NLP will perform the analysis. After the cloud NLP obtains the second semantic analysis result, it will further match the second semantic analysis result with the cloud corpus template to determine whether the matching degree is less than 0.7;
[0180] If the matching degree is not less than 0.7, it means that the corpus template is matched normally, and the execution parameters can be sent down, and the generalized template can be sent down to the smart terminal;
[0181] If the matching degree is less than 0.7, it means that the existing corpus template is not hit. The AC automaton can be used to extract keywords, and the keywords are spliced into corpus template guidance information according to grammatical rules to guide the user to speak, and the execution parameters corresponding to the corpus template guidance information are issued, and the corpus template guidance information is issued.
[0182] In the embodiment of the present application, by establishing a set of terminal-cloud automatic switching rules, the intelligent voice device can perform terminal-side recognition and analysis for simple instructions during interaction, and automatically switch to cloud-side recognition when the terminal-side recognition cannot meet the requirements. This can not only improve the interactive response speed, but also ensure the accuracy of the interaction, while reducing the network transmission pressure and cloud resource requirements, improving the user interaction experience while reducing the system's operating costs.
[0183] In an optional embodiment, a voice interaction system may also be included, which may include four parts: middleware, cloud service platform, third-party applications and smart terminals. Middleware and third-party applications are usually set in smart terminals.
[0184] The main modules of the intelligent terminal include: a voice information processing module, a natural language processing module, an offline voice processing middleware, an edge gateway, and a skills module.
[0185] Voice information processing module: This module is used to convert the user's voice file into text information. During the conversion process, it will combine the hot words in various fields uploaded by the user for priority matching.
[0186] Natural language processing module: This module filters sensitive words based on the text information spoken by the user, then preferentially matches question-and-answer pairs, and finally performs NLP parsing, and transmits the parsed skill domains and parsing results to the intelligent terminal.
[0187] Offline voice processing middleware: This module is responsible for functions such as offline NLP parsing, offline NLP engine update, and offline hot word synchronization.
[0188] Edge gateway: Realize security authentication, key update, and network connection functions between the gateway and the attached devices.
[0189] Skills module: This module is used to feedback follow-up questions and subsequent response logics according to the device-reported attribute information and in combination with the NLP parsing results.
[0190] The intelligent terminal and cloud service platform provided by the embodiments of the present application are described below. The intelligent terminal and cloud service platform described below can be mutually corresponding and referred to the voice interaction method described above.
[0191] The intelligent terminal involved in the embodiments of the present application may be a device that provides voice and / or data connectivity to users, a handheld device with a wireless connection function, or other processing devices connected to a wireless modem, etc. In different systems, the name of the terminal device may also be different. For example, in a 5G system, the terminal device may be called a user equipment (UE).
[0192] Figure 5 For the structural schematic diagram of the intelligent terminal according to the embodiments of the present application, refer to Figure 5 , the embodiments of the present application also provide an intelligent terminal, which may include: a memory 510, a transceiver 520, and a processor 530;
[0193] The memory 510 is used to store computer programs; the transceiver 520 is used to transmit and receive data under the control of the processor 530; the processor 530 is used to read the computer programs in the memory 510 and perform the following operations:
[0194] When a voice end identifier is detected within the first voice duration of the voice stream received by the intelligent terminal, semantic parsing is performed on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic parsing result;
[0195] When the first confidence level of the first semantic parsing result is less than a first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform;
[0196] When the second semantic parsing result obtained by parsing the voice stream through the online natural language processing model by the cloud service platform includes the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
[0197] Among them, in Figure 5 The bus architecture can include any number of interconnected buses and bridges. Specifically, various circuits represented by one or more processors represented by the processor 530 and the memory represented by the memory 510 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art. Therefore, they will not be further described herein. The bus interface provides an interface. The transceiver 520 can be multiple components, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on the transmission medium. For different user devices, the user interface 540 can also be an interface capable of externally connecting or internally connecting required devices.
[0198] The processor 530 is responsible for managing the bus architecture and general processing, and the memory 510 can store the data used by the processor 530 when performing operations.
[0199] The processor 530 is used to execute any of the methods provided in the embodiments of the present application according to the obtained executable instructions by calling the computer program stored in the memory 510. The processor and the memory can also be physically separated.
[0200] Optionally, the processor 530 is further used to perform the following operations:
[0201] The intelligent terminal sends a device capability report to the cloud service platform; the device capability report includes at least one of the following: device name, product ID, transaction ID, service data; the service data includes the target device function information;
[0202] The intelligent terminal receives the device capability report feedback information sent by the cloud service platform; the device capability report feedback information includes at least one of the following: transaction ID, functional data; wherein, the functional data includes: the second semantic parsing result.
[0203] Update the offline natural language processing model of the intelligent terminal according to the second semantic parsing result and the corresponding speech stream in the device capability report feedback information.
[0204] Optionally, the processor 530 is further configured to perform the following operations:
[0205] In the case that the speech end identifier is not detected in the speech stream received by the intelligent terminal within the first speech duration, transmit the speech stream to the cloud service platform.
[0206] Optionally, the processor 530 is further configured to perform the following operations:
[0207] Input the audio acoustic feature vector of the speech stream into a preset encoder to obtain a first acoustic feature;
[0208] Input the first corpus feature vector set constructed based on the first corpus template set of the intelligent terminal into the preset encoder to obtain a second corpus feature vector set;
[0209] Determine the first confidence level based on the matching degree between the first acoustic feature and the corpus feature set.
[0210] Figure 6 For the structural schematic diagram of the cloud service platform according to the embodiment of the present application, refer to Figure 6 , the embodiment of the present application further provides a cloud service platform, which may include: a memory 610, a transceiver 620, and a processor 630;
[0211] The memory 610 is used to store computer programs; the transceiver 620 is used to transmit and receive data under the control of the processor 630; the processor 630 is used to read the computer programs in the memory 610 and perform the following operations:
[0212] In the case that the confidence level of the first semantic parsing result of the speech stream by the intelligent terminal through the offline natural language processing model is less than the first preset threshold, receive the speech text corresponding to the speech stream sent by the intelligent terminal;
[0213] The cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic parsing result;
[0214] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than a third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
[0215] Wherein, in Figure 6 Among them, the bus architecture can include any number of interconnected buses and bridges. Specifically, various circuits represented by one or more processors represented by the processor 630 and the memory represented by the memory 610 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art. Therefore, they will not be further described herein. The bus interface provides an interface. The transceiver 620 can be multiple components, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on the transmission medium. The processor 630 is responsible for managing the bus architecture and general processing, and the memory 610 can store the data used by the processor 630 when executing operations.
[0216] Optionally, the processor 630 is further configured to perform the following operations:
[0217] When receiving the device capability report sent by the intelligent terminal, write the second semantic parsing result into the function data of the device capability report feedback information;
[0218] Feedback the device capability report feedback information to the intelligent terminal for the intelligent terminal to update the offline natural language processing model according to the second semantic parsing result in the device capability report feedback information.
[0219] Optionally, the processor 630 is further configured to perform the following operations:
[0220] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is less than or equal to the third preset threshold, obtain keywords from the second semantic parsing result through the AC automaton algorithm, and construct a target sentence pattern template according to the keywords;
[0221] Obtain the template similarity between the target sentence pattern template and each corpus template in the second corpus template set;
[0222] When at least one of the template similarities is greater than a fourth preset threshold, generate corpus template guidance information based on the target corpus template, and send the corpus template guidance information and the target device function information corresponding to the target corpus template to the intelligent terminal;
[0223] Among them, the target corpus template is the corpus template with the highest template similarity to the target sentence pattern template in the second corpus template set, and the corpus template guiding information includes the target corpus template.
[0224] Optionally, the processor 630 is further configured to perform the following operations:
[0225] In the case that there is no corpus template in the second corpus template set whose similarity to the target sentence pattern template is greater than the fourth preset threshold, generate unrecognized information and send the unrecognized information to the intelligent terminal.
[0226] Optionally, the processor 630 is further configured to perform the following operations:
[0227] In the case that no voice end flag is detected within the first voice duration of the voice stream received by the intelligent terminal, receive the voice stream transmitted by the intelligent terminal;
[0228] The cloud service platform performs text recognition on the voice stream to obtain the voice text corresponding to the voice stream.
[0229] It should be noted here that the terminal and network device provided in the embodiments of the present application can implement all the method steps implemented in the above method embodiments and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments in this embodiment will not be specifically described herein again.
[0230] Figure 7 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application, as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete communication with each other through the communication bus 740. The processor 710 may call a computer program in the memory 730 to execute the steps of the voice interaction method, for example, including:
[0231] In the case that a voice end flag is detected within the first voice duration of the voice stream received by the intelligent terminal, perform semantic parsing on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic parsing result;
[0232] In the case that the first confidence level of the first semantic parsing result is less than the first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform;
[0233] When the second semantic parsing result obtained by parsing the speech stream through the online natural language processing model on the cloud service platform includes the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information;
[0234] Or when the confidence level of the first semantic parsing result obtained by the intelligent terminal through the offline natural language processing model for semantic parsing of the speech stream is less than the first preset threshold, the cloud service platform receives the speech text corresponding to the speech stream sent by the intelligent terminal;
[0235] The cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through the online natural language processing model to obtain a second semantic parsing result;
[0236] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than the third preset threshold, the cloud service platform sends the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
[0237] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0238] On the other hand, the embodiment of the present application further provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer executes the steps of the voice interaction method, for example, including:
[0239] When a voice end flag is detected within the first voice duration of the voice stream received by the intelligent terminal, semantic analysis is performed on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic analysis result;
[0240] When the first confidence level of the first semantic analysis result is less than a first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform;
[0241] When the second semantic analysis result obtained by the cloud service platform through the online natural language processing model for parsing the voice stream contains the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information;
[0242] Or when the confidence level of the first semantic analysis result of the semantic analysis of the voice stream by the intelligent terminal through the offline natural language processing model is less than the first preset threshold, receive the voice text corresponding to the voice stream sent by the intelligent terminal;
[0243] The cloud service platform performs semantic analysis on the voice text corresponding to the voice stream through the online natural language processing model to obtain a second semantic analysis result;
[0244] When the matching degree between the second semantic analysis result and the target corpus template in the second corpus template set of the cloud service platform is greater than a third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic analysis result in the second corpus template set.
[0245] On the other hand, an embodiment of the present application further provides a processor-readable storage medium, and the processor-readable storage medium stores a computer program, and the computer program is used to enable the processor to execute the steps of the voice interaction method, for example, including:
[0246] When a voice end flag is detected within the first voice duration of the voice stream received by the intelligent terminal, semantic analysis is performed on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic analysis result;
[0247] When the first confidence level of the first semantic analysis result is less than a first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform;
[0248] When the second semantic parsing result obtained by parsing the speech stream through the online natural language processing model on the cloud service platform includes the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information;
[0249] Or when the confidence level of the first semantic parsing result of the speech stream parsed by the intelligent terminal through the offline natural language processing model is less than the first preset threshold, receive the speech text corresponding to the speech stream sent by the intelligent terminal;
[0250] The cloud service platform parses the speech text corresponding to the speech stream through the online natural language processing model to obtain a second semantic parsing result;
[0251] When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than the third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
[0252] The processor-readable storage medium can be any available medium or data storage device accessible by the processor, including but not limited to magnetic memory (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memory (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid state drives (SSD)), etc.
[0253] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0254] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0255] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.< / close> < / open> < / device> < / room> < / open> < / device> < / room> < / close> < / open> < / device> < / room>
Claims
1. A voice interaction method, applied to an intelligent terminal, characterized in that, Including: When a voice end flag is detected within a first voice duration of a voice stream received by the intelligent terminal, semantic parsing is performed on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic parsing result; When a first confidence level of the first semantic parsing result is less than a first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform; When a second semantic parsing result obtained by parsing the voice stream through an online natural language processing model by the cloud service platform includes target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
2. The voice interaction method according to claim 1, wherein After the step where the intelligent terminal receives the device function information transmitted by the cloud service platform, the method further includes: The intelligent terminal sends a device capability report to the cloud service platform; the device capability report includes at least one of the following: device name, product ID, transaction ID, service data; the service data includes the target device function information; The intelligent terminal receives device capability report feedback information sent by the cloud service platform; the device capability report feedback information includes at least one of the following: transaction ID, function data; wherein, the function data includes: the second semantic parsing result; Update the offline natural language processing model of the intelligent terminal according to the second semantic parsing result and the corresponding voice stream in the device capability report feedback information.
3. The voice interaction method according to claim 1, characterized in that Before the step where the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information, the method further includes: When the voice stream received by the intelligent terminal does not detect a voice end flag within the first voice duration, transmit the voice stream to the cloud service platform.
4. The voice interaction method according to claim 2, wherein The calculation method of the first confidence level of the first semantic parsing result specifically includes: Input the audio acoustic feature vector of the voice stream into a preset encoder to obtain a first acoustic feature; Input a first corpus feature vector set constructed based on a first corpus template set of the intelligent terminal into the preset encoder to obtain a second corpus feature vector set; Determine the first confidence level based on the matching degree between the first acoustic feature and the corpus feature set.
5. A voice interaction method is applied to a cloud service platform, characterized in that, Including: When the confidence level of the first semantic parsing result obtained by the intelligent terminal performing semantic parsing on the voice stream through the offline natural language processing model is less than the first preset threshold, receive the voice text corresponding to the voice stream sent by the intelligent terminal; The cloud service platform performs semantic parsing on the voice text corresponding to the voice stream through an online natural language processing model to obtain a second semantic parsing result; When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than a third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
6. The voice interaction method according to claim 5, wherein, After the step of sending the target device function information corresponding to the target corpus template to the intelligent terminal when the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is higher than a third preset threshold, the method further includes: When receiving the device capability report sent by the intelligent terminal, write the second semantic parsing result into the function data of the device capability report feedback information; Feed back the device capability report feedback information to the intelligent terminal for the intelligent terminal to update the offline natural language processing model according to the second semantic parsing result in the device capability report feedback information.
7. The voice interaction method according to claim 5, wherein After the step that the cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic parsing result, it further includes: When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is less than or equal to the third preset threshold, obtain keywords from the second semantic parsing result through the AC automaton algorithm, and construct a target sentence pattern template according to the keywords; Obtain the template similarity between the target sentence pattern template and each corpus template in the second corpus template set; When at least one of the template similarities is greater than a fourth preset threshold, generate corpus template guiding information based on the target corpus template, and send the corpus template guiding information and the target device function information corresponding to the target corpus template to the intelligent terminal; Wherein, the target corpus template is the corpus template with the highest template similarity with the target sentence pattern template in the second corpus template set, and the corpus template guiding information includes the target corpus template.
8. The voice interaction method according to claim 7, wherein After the step of obtaining the template similarity between the target sentence pattern template and each corpus template in the second corpus template set, it further includes: When there is no corpus template in the second corpus template set whose similarity with the target sentence pattern template is greater than the fourth preset threshold, generate unrecognized information and send the unrecognized information to the intelligent terminal.
9. The voice interaction method according to claim 5, wherein Before the step that the cloud service platform performs semantic parsing on the speech text corresponding to the speech stream through an online natural language processing model to obtain a second semantic parsing result, it further includes: When no speech end flag is detected within a first speech duration of the speech stream received by the intelligent terminal, receive the speech stream transmitted by the intelligent terminal; The cloud service platform performs text recognition on the speech stream to obtain the speech text corresponding to the speech stream.
10. An intelligent terminal, characterized in that, Including a memory, a transceiver, and a processor; A memory for storing a computer program; a transceiver for transmitting and receiving data under the control of the processor; a processor for reading the computer program in the memory and performing the following operations: When a voice end identifier is detected within a first voice duration of the voice stream received by the intelligent terminal, semantic parsing is performed on the voice text corresponding to the voice stream through the offline natural language processing model of the intelligent terminal to obtain a first semantic parsing result; When the first confidence level of the first semantic parsing result is less than a first preset threshold, the intelligent terminal transmits the voice text corresponding to the voice stream to the cloud service platform; When the second semantic parsing result obtained by parsing the voice stream through the online natural language processing model of the cloud service platform contains the target device function information corresponding to the intelligent terminal, the intelligent terminal receives the target device function information transmitted by the cloud service platform and implements the device function according to the target device function information.
11. A cloud service platform, characterized in that, It includes a memory, a transceiver, and a processor; A memory for storing a computer program; A transceiver for transmitting and receiving data under the control of the processor; a processor for reading the computer program in the memory and performing the following operations: When the confidence level of the first semantic parsing result of the semantic parsing of the voice stream by the intelligent terminal through the offline natural language processing model is less than a first preset threshold, receive the voice text corresponding to the voice stream sent by the intelligent terminal; The cloud service platform performs semantic parsing on the voice text corresponding to the voice stream through the online natural language processing model to obtain a second semantic parsing result; When the matching degree between the second semantic parsing result and the target corpus template in the second corpus template set of the cloud service platform is greater than a third preset threshold, send the target device function information corresponding to the target corpus template to the intelligent terminal; wherein, the target corpus template is the corpus template with the highest matching degree with the second semantic parsing result in the second corpus template set.
12. An electronic device, comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the voice interaction method according to any one of claims 1 to 4, or implements the steps of the voice interaction method according to any one of claims 5 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the voice interaction method according to any one of claims 1 to 4, or implements the steps of the voice interaction method according to any one of claims 5 to 9.